Skip to content
Ali Akbari
Menu

Case study

TFLMarkets

A live prop-trading and tournament platform at TFL, the company I co-founded. Trading is simulated; the wallet, payments and identity checks are real. I designed the backend and wrote most of it with one other engineer, built most of its five web apps and set up production.
EmploymentImplementedenvironment:Production× 10 claims
Role
Architecture, backend, frontend and production (co-founder, with one other engineer)
Period
Apr 2026 – Sep 2026
Updated
11 October 2026
  • Go
  • Rust
  • Kafka
  • Avro
  • gRPC
  • WebSockets
  • PostgreSQL
  • PgBouncer
  • TimescaleDB
  • QuestDB
  • Redis
  • Nuxt
  • WebGL2
  • Docker Swarm
  • Prometheus
  • k6
Jump to a section
  1. Overview
  2. Architecture
  3. Event backbone
  4. Web applications and the WebGL chart
  5. Production operations
  6. Incident review
  7. Load test
  8. How the work was done
  9. My role
  10. Limitations
  11. Evidence index

Overview

TFLMarkets lets traders take simulated prop-firm challenges and enter trading tournaments on live market prices. Orders never reach a real market, but the fees, the wallet, payouts and identity checks involve real money and real people. It runs in production.

Between April and September 2026 I designed the platform and built it with one other engineer. Of about 2,200 backend commits, about 1,600 are mine, and about 70% of the frontend commits. Most of the code was written by AI coding agents from specifications I wrote; I reviewed, tested, integrated and deployed it. Five services were brought in from earlier company code and extended.

Architecture

Visualization from source

Service inventory and the two main data paths

Drawn from the source code and configuration; nothing here is running.

Services
25
20 in Go, one each in Rust, C#, Python, TypeScript
Avro event schemas
180
22 event domains
Grafana dashboards
36
~520 alert rules
Go test functions
~6,800
Rust tests ~1,130
Trading· 3 services
  • trading-engine-serviceRust · built in this project

    In-memory order/position engine: one actor per account, tick-driven P/L, SL/TP, margin call/stop-out, prop drawdown rules, sharded ownership with HA failover, crash recovery by replay.

  • websocket-gatewayGo · built in this project

    Browser WebSocket termination: ticket auth, per-channel subscriptions, tick throttling, relays trading ops to the engine over gRPC, fans out account/price/notification events.

  • risk-serviceGo · built in this project

    Pattern detection and risk flags on trading events; backward drawdown replay.

Market data· 4 services
  • broker-feed-bridgeC# · built in this project

    One container per upstream broker terminal account; streams ticks over gRPC. A pool runs hot-standby.

  • price-feed-serviceGo · built in this project

    Owns the bridge pool (encrypted credentials, per-account Swarm services), active-account selector with health-based failover, authoritative tick stream.

  • market-data-serviceGo · built in this project

    Tick-to-candle aggregation, candle history API, symbol/session catalogue, multi-source failover and deep-history backfill.

  • history-backfill-sourceTypeScript · built in this project

    History-only source serving bounded M1 bid/ask candle windows from a public datafeed to the market-data backfill worker; never a live source.

Accounts and identity· 4 services
  • auth-serviceGo · imported and extended

    User-facing auth flows, sessions, WS ticket endpoint, KYC orchestration.

  • authentication-serviceGo · imported and extended

    Token authority: issue/validate tokens, WS tickets.

  • kyc-ml-servicePython · imported and extended

    Face detection/matching for KYC.

  • account-serviceGo · built in this project

    Trading account lifecycle (demo, broker, prop phases, funded, tournament) and account events.

Money· 4 services
  • wallet-serviceGo · built in this project

    Internal ledger/balances.

  • payment-serviceGo · built in this project

    Deposits through external payment gateways (card and crypto), webhooks.

  • plan-serviceGo · built in this project

    Product catalogue: challenge plans, coupons, tournament templates, broker account offers.

  • referral-serviceGo · built in this project

    Referral attribution and commissions.

Engagement: prop, broker, tournament· 5 services
  • prop-serviceGo · built in this project

    Prop challenge lifecycle: phases, rule evaluation from trading events, funded accounts.

  • broker-serviceGo · built in this project

    Simulated broker account products.

  • tournament-serviceGo · built in this project

    Trading competitions, leaderboards.

  • engagement-serviceGo · built in this project

    Gamification.

  • analytics-serviceGo · built in this project

    Trading statistics from trading events (time-series).

Platform and ops· 5 services
  • admin-serviceGo · built in this project

    Admin gateway: RBAC, audit log, aggregates other services' admin APIs.

  • notification-serviceGo · imported and extended

    Email + in-app notifications from domain events.

  • support-serviceGo · imported and extended

    Support tickets.

  • content-serviceGo · built in this project

    Bilingual blog/content backend.

  • backup-agentshell · built in this project

    Scheduled encrypted off-site backups.

Live prices

  1. upstream broker terminals → broker-feed-bridge (pool, one per account)vendor terminal API
  2. broker-feed-bridge → price-feed-servicegRPC server-streaming, one stream per bridge; selector picks the active account, standbys only update health
  3. price-feed-service → tick store (QuestDB)write
  4. price-feed-service → Redislast-tick cache + pub/sub tick:{symbol}
  5. price-feed-service → trading-engine-servicegRPC Subscribe stream (bounded mailbox per account actor; slow actors shed, never block the stream)
  6. price-feed-service → Kafka market.ticks_rawKafka producer
  7. Kafka market.ticks_raw → market-data-serviceconsumer -> candle aggregation -> TimescaleDB hypertables; emits candle_sealed
  8. Redis tick:{symbol} → websocket-gatewaypub/sub, throttled per client
  9. websocket-gateway → browserWebSocket /ws/platform/prices
  10. history-backfill-source → market-data-serviceREST, backfill only (insert-if-absent; live candles win)

Orders and domain events

  1. browser → websocket-gatewayWebSocket op frame (place_order etc.)
  2. websocket-gateway → trading-engine-servicegRPC
  3. trading-engine-service → websocket-gatewayop_ack/op_result + account updates over Redis pub/sub account:{id}:update
  4. trading-engine-service → Kafka trading.* (order_placed, position_closed, breach_detected, margin_call, ...)outbox -> Kafka, Avro
  5. Kafka trading.* → prop-service, analytics-service, risk-service, tournament-serviceconsumer groups
  6. account-service → Kafka account.* (created, activated, status_changed, breached, ...)transactional outbox -> Kafka, Avro
  7. Kafka account.* → trading-engine-service (lifecycle consumer spawns account actors), notification-service, referral-service, othersconsumer groups
  8. Kafka (all domains) → notification-serviceconsumer -> email / in-app
  9. failed consumers → per-topic DLQretry then dead-letter
Compiled from the company repository (service directories, stack files and CI) in October 2026, with brand and provider names replaced by generic ones. Service names are the real ones; nothing here is running.

Twenty Go services, a Rust trading engine, a C# bridge to the price source, a Python service for identity-document checks and a small TypeScript history source. Most service-to-service calls are REST; the hot paths (prices into the engine, orders from the gateway) are gRPC. Every service owns its own PostgreSQL database behind PgBouncer.

TFLMarkets: price path and event pathPrices flow from the broker terminals through a C# bridge to the Go price-feed service, which streams them to the Rust trading engine over gRPC, stores ticks in QuestDB and publishes them on Redis Pub/Sub, where the WebSocket gateway fans them out to the web applications. Orders go from the web applications through the WebSocket gateway to the engine over gRPC. Domain services and the engine write events to an outbox in the same transaction as their state; publishers send them to Kafka with Avro schemas, and consumers such as risk, prop, tournament and notifications read them, with failed messages going to dead-letter queues. The market-data service builds candles into TimescaleDB. Every service reports to Prometheus, Loki and Tempo.gRPCticksWebSocketordersevents via outboxsame transactionpublisherraw ticksPrice sourcebroker terminalsPrice bridgeC# · gRPC streamPrice feedGo · failoverTrading engineRust · in memoryWeb appsNuxt · WebGL chartQuestDBticksRedis Pub/Subprices · account updatesWebSocket gatewayGo · ticketedDomain servicesGo · account · wallet · prop · …PostgreSQL + outboxone database per service · PgBouncerKafkaAvro · schema registryConsumersrisk · prop · tournaments · alerts · DLQMarket-data servicecandlesTimescaleDBcandle historyObservabilityPrometheus · Alertmanager · Grafana · Loki · Tempo · OpenTelemetry
Figure: simplified from the service inventory above. Prices never go through Kafka; business events always do. Production ran all of this on a single node.

Event backbone

Domain events (account created, order filled, breach detected and so on) travel through Kafka, encoded with Avro against 180 schemas in a Schema Registry. Each service writes its events to an outbox table in the same transaction as its own state change, and a publisher sends them on, so a crash can never leave a change without its event. Order placement is idempotent on a client order id, and consumers that cannot process a message send it to a per-topic dead-letter queue.

Live prices take a different path. The price-feed service streams ticks to the Rust engine over gRPC and publishes them on Redis Pub/Sub, where the WebSocket gateway I wrote picks them up and sends them to each client, throttled per connection. Market data is stored in QuestDB for ticks and TimescaleDB hypertables for candles.

Web applications and the WebGL chart

Five Nuxt applications share a monorepo with common design tokens, UI components and an API client: the trading terminal, the user dashboard (wallet, KYC), the admin back office, the marketing site and a bilingual blog.

The terminal's chart is built in-house, with no charting library: series, axes, crosshair and about 70 drawing tools. I wrote its WebGL2 series renderer:

  • All coordinate maths runs in 64-bit floats in JavaScript and only pixel positions are sent to the GPU, because FX prices lose precision in 32-bit floats.
  • Candles are drawn as quads grouped by colour: four draw calls by default, however many candles are on screen, from one reused vertex buffer.
  • It shares its pixel maths with the Canvas 2D renderer, so the two produce the same pixels. If WebGL is unavailable, a shader fails or the context is lost for more than two seconds, the chart switches to Canvas 2D.
Screenshots

The trading terminal and its WebGL chart

Images of the real product; the note below says where and how they were captured.

Trading terminal with an EURUSD one-hour candlestick chart drawn by the WebGL renderer, an order panel and open positions.
EURUSD, one-hour candles, with the order panel and positions
Trading terminal showing a gold (XAUUSD) candlestick chart drawn by the WebGL renderer.
XAUUSD chart
Captured on 11 October 2026 from the terminal application running locally against the repository's own mock backend: demo accounts and mock prices, no production data. The chart reported its WebGL renderer as active.

Production operations

Production runs on Docker Swarm, deployed by GitLab CI from a runner on the host: each service is rolled forward and checked for health. I set up the monitoring stack (Prometheus with about 520 alert rules including SLO burn-rate alerts, 36 generated Grafana dashboards, Loki for logs, Tempo and OpenTelemetry for traces, and Alertmanager) and wrote the runbooks: more than 60 operational runbooks and about 130 more for individual alerts.

Backups came late. From August 2026 the databases have encrypted offsite backups and point-in-time recovery with pgBackRest.

Incident review

In August 2026 three small symptoms were reported together: lock warnings in the PostgreSQL log, connection resets, and a dozen events that failed to publish. Nothing looked broken and no alert had fired. I investigated read-only against production first, and found three separate causes and a fourth, more serious problem nobody had reported.

  1. Advisory locks behind a transaction pooler. Outbox publishers elected a leader with PostgreSQL session advisory locks. Behind PgBouncer in transaction mode, the lock stays on a backend the application no longer controls, so leadership could be held by two replicas or by none. Production showed zero advisory locks held while four services believed they had one.
  2. Connection resets. Two were explained by a redeploy; 23 could not be attributed with the logging in place, and I reported them as unattributable instead of guessing.
  3. Transient Kafka outages. The single broker had fenced itself three times. All twelve affected events were shown to have been delivered and handled, from four independent facts.
  4. Unbounded retries. Events that could never succeed (a schema mismatch, a missing topic) were retried without limit and blocked the head of their queue. One row had over 5 million attempts. Eighteen risk events had been stuck for 49 days, because the only signal was a counter that fired once.

What I changed. A shared retry library that classifies failures (transient, permanent, ambiguous, missing topic), backs off and parks permanent failures for review, live on 11 of 13 event publishers; the two money services were held for a second review. A pooler-safe lock and a Redis lease elector, adopted by the first service when my work ended. Alert rules for work that is over budget, and a repair tool for the stuck events that reconciles them instead of replaying them.

Lessons. "Healthy" is not "working"; a counter is not a backlog; a retry column that nothing reads is not a limit; and pooling mode is part of your concurrency model.

Load test

Recorded results

k6 load test: 100 traders placing orders over WebSocket

Published run records, shown as they were recorded.

Orders placed
2,250
all succeeded, 0 errors
Order round trip, p95
60 ms
median 55 ms · max 86 ms
Price tick to client, p95
91 ms
Account update to client, p95
206 ms

Order round trip (WebSocket send to result)

median55 msp9058 msp9560 msmax86 mstarget p95 50 ms

Setup: 100 virtual users from one machine, each with two WebSockets (account and prices), placing a market order every 30 seconds for 10 minutes, against the production stack before launch, on demo accounts.

Threshold set before the runObservedResult
place_order_op_result_ms p(95)<50p95=60msMissed
ws_fanout_observed_ms{family:account} p(95)<200p95=206msMissed
ws_fanout_observed_ms{family:tick} p(95)<200p95=91msMet
place_order_success_rate rate>0.99100% (2250/2250)Met
op_error_total count<500000Met
ws_handshake_failures_total count<200000Met
http_req_duration{kind:ws_ticket} p(95)<1000p95=15msMet
The four runs before this one placed no orders at all

Four consecutive runs over ~32 hours had every order rejected (0 of 2,250 succeeded each time) while WebSockets, auth and tickets all worked. The failures were not load-related: they came from the path a brand-new account takes from creation to being tradable in the in-memory engine. Fixes landed between runs; the fifth run placed all 2,250 orders.

  1. 2026-05-25: Test-harness plumbing: per-IP WebSocket handshake rate-limit CIDR bypass at both the edge proxy and the WS gateway (both layers must agree or the gateway/edge still 429s), and an admin bulk-seed endpoint so synthetic users go through the normal outbox -> Kafka user.created path instead of direct SQL inserts that downstream services never observe.
  2. 2026-05-26 early: Path-scoped IP-binding bypass in the token validator for the seed endpoint, so in-cluster seeding with a minted admin token works (signature/expiry/admin/session checks still enforced).
  3. 2026-05-26 midday: Event-contract drift between the Go account service and the Rust trading engine: the engine consumer read a legacy `account_kind` field while the producer had moved to a canonical 11-value `kind`. Engine now prefers `kind` and falls back to the legacy field; producer again populates the legacy fields; a jsonb cast error in the account repository was fixed. The seed script was rewritten to create demo accounts through the real API so each one flows account.created -> activate -> outbox -> Kafka -> engine.
  4. 2026-05-26 evening: Avro encoding: made `kind`/`status` nullable for backward compatibility and fixed nullable timestamp-millis union handling in the goavro encoder (with a regression test).
  5. 2026-05-26 evening: Engine HA mode never flipped to Live: the readiness watcher only accepted Replaying, but HA mode skips the cold-start step that sets it, so the precondition now accepts Booting or Replaying. Also fixed UUID hex formatting in the seed script so bearer minting and user lookup agree, and added logging of every op_error frame in the k6 script.
  6. 2026-05-27 01:03 (+04:00), ~1 h before the successful run: Root cause of the remaining rejections: actors created by the live lifecycle consumer for brand-new accounts were never registered in the symbol->account tick-routing index (the cold-start and recovery paths did this; the live path did not). Without ticks the actor's last-tick cache stayed empty and every PlaceOrder was rejected as SESSION_CLOSED. Fix: register every new actor for every symbol in the live symbol book. A latent bug that only affects fresh accounts that have never traded, which is exactly what a load test with freshly seeded accounts creates.
  • Single load-generator machine: all 100 VUs came from one client, and per-IP rate limits had to be bypassed for it. It does not model 100 independent networks.
  • Ran against the production stack (before public launch) using seeded demo accounts, not staging as the script header recommends; no real-money accounts were involved.
  • Order latency is client-perceived: WS send -> op_result, including network, edge proxy and WS gateway, not engine-internal time.
  • WS fan-out latency compares client clock to server timestamps, so it includes any cross-host clock skew.
  • Two thresholds were missed: order p95 was 60 ms against a 50 ms target, and account-event fan-out p95 was 206 ms against a 200 ms target. Tick fan-out (91 ms) and all success/error gates passed.
  • Order latency was very tight (min 52, median 55, max 86 ms), which suggests a fixed network/hop floor rather than queueing under load.
  • Load was modest: about 3.3 orders/s (one order per VU every 30 s); this is a correctness-under-concurrency and latency test, not a throughput ceiling test.
From the k6 summary of the run on 26 May 2026 (22:05 UTC), committed to the company repository. The raw output stays private, so these figures are self-reported and cannot be reproduced from public material.

How the work was done

The design lives in 30 architecture specifications. Coding agents implemented from them, and I reviewed, tested and integrated the result. About two thirds of the commits carry an AI co-author tag. The repository has about 6,800 Go test functions, about 1,100 Rust tests and integration tests against real PostgreSQL with Testcontainers, run locally rather than in the deployment pipeline. My public example of the same way of working, with tests you can inspect, is BoundedCode.

My role

  • Co-founder and technical lead; designed the architecture and wrote the specifications.
  • Wrote most of the backend, including the WebSocket gateway and nearly all of the infrastructure, CI and monitoring; the other engineer wrote more of the prop and tournament services than I did.
  • Built most of the five frontends, including the WebGL renderer.
  • Ran production: deployments, incidents, backups.

Limitations

  • Everything here is self-reported; the code and raw data are private.
  • Production ran on a single node, and backups arrived in August 2026.
  • The load test is one run from one machine, and it missed two of its own targets.
  • The incident fixes were still being rolled out when my work ended.

Evidence index

Every claim this case study relies on, rendered from the evidence manifests.

Implementedenvironment:ProductionSelf-reported · private source

TFLMarkets backend architecture

Designed the TFLMarkets backend and wrote most of it with one other engineer: 20 Go services, a Rust order-execution engine and a Python KYC service, communicating over REST, gRPC and Kafka.

BackendEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Event backbone with outbox, idempotency and dead-letter handling

Built the event backbone on Kafka with Avro schemas in a Schema Registry and a transactional outbox in each service, plus idempotent order placement and dead-letter handling.

BackendEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Real-time price path to the trading terminal

Live prices flow from the price feed to the Rust engine over gRPC and reach traders through Redis Pub/Sub and a WebSocket gateway that I wrote, with ticketed connections, heartbeats and reconnection with backoff on the client.

Distributed systemsEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Time-series storage for market data

Market data is stored in TimescaleDB hypertables and QuestDB, which serve chart history to the trading terminal separately from the transactional PostgreSQL databases.

Data systemsEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Five Nuxt applications and a custom WebGL2 chart renderer

Built five Nuxt applications in a shared monorepo (trading terminal, user dashboard, admin back office, marketing site and blog) and wrote a custom WebGL2 candlestick renderer with a Canvas 2D fallback for the trading terminal.

WebEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Monitoring stack with SLO burn-rate alerts

Set up Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager for the platform, with generated dashboards and SLO burn-rate alert rules, and wrote runbooks for the alerts.

ObservabilityEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Production deployment, backups and recovery

Set up production on Docker Swarm with GitLab CI deployments, wrote more than 60 runbooks, and added encrypted offsite backups with point-in-time recovery.

SREEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Production incident investigation

Investigated a production reliability incident, traced it to unbounded outbox retries and session advisory locks behind PgBouncer transaction pooling, bounded the retries in 11 of 13 event publishers and began replacing the locks.

SREEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Order-placement load test with k6

Load-tested order placement with k6 at 100 concurrent users: 2,250 orders, all successful, 60 ms at p95 from sending the order over the WebSocket to receiving its result.

BackendEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · private source

Specification-first delivery with coding agents

Wrote 30 architecture specifications that coding agents implemented from; I owned the design, review, testing and integration.

VerificationEmployment · no public artifactsEvidence

← All work