Jump to a section
Overview
TFLMarkets lets traders take simulated prop-firm challenges and enter trading tournaments on live market prices. Orders never reach a real market, but the fees, the wallet, payouts and identity checks involve real money and real people. It runs in production.
Between April and September 2026 I designed the platform and built it with one other engineer. Of about 2,200 backend commits, about 1,600 are mine, and about 70% of the frontend commits. Most of the code was written by AI coding agents from specifications I wrote; I reviewed, tested, integrated and deployed it. Five services were brought in from earlier company code and extended.
Architecture
Service inventory and the two main data paths
Drawn from the source code and configuration; nothing here is running.
- Services
- 25
- 20 in Go, one each in Rust, C#, Python, TypeScript
- Avro event schemas
- 180
- 22 event domains
- Grafana dashboards
- 36
- ~520 alert rules
- Go test functions
- ~6,800
- Rust tests ~1,130
Trading· 3 services
trading-engine-serviceRust · built in this project
In-memory order/position engine: one actor per account, tick-driven P/L, SL/TP, margin call/stop-out, prop drawdown rules, sharded ownership with HA failover, crash recovery by replay.
websocket-gatewayGo · built in this project
Browser WebSocket termination: ticket auth, per-channel subscriptions, tick throttling, relays trading ops to the engine over gRPC, fans out account/price/notification events.
risk-serviceGo · built in this project
Pattern detection and risk flags on trading events; backward drawdown replay.
Market data· 4 services
broker-feed-bridgeC# · built in this project
One container per upstream broker terminal account; streams ticks over gRPC. A pool runs hot-standby.
price-feed-serviceGo · built in this project
Owns the bridge pool (encrypted credentials, per-account Swarm services), active-account selector with health-based failover, authoritative tick stream.
market-data-serviceGo · built in this project
Tick-to-candle aggregation, candle history API, symbol/session catalogue, multi-source failover and deep-history backfill.
history-backfill-sourceTypeScript · built in this project
History-only source serving bounded M1 bid/ask candle windows from a public datafeed to the market-data backfill worker; never a live source.
Accounts and identity· 4 services
auth-serviceGo · imported and extended
User-facing auth flows, sessions, WS ticket endpoint, KYC orchestration.
authentication-serviceGo · imported and extended
Token authority: issue/validate tokens, WS tickets.
kyc-ml-servicePython · imported and extended
Face detection/matching for KYC.
account-serviceGo · built in this project
Trading account lifecycle (demo, broker, prop phases, funded, tournament) and account events.
Money· 4 services
wallet-serviceGo · built in this project
Internal ledger/balances.
payment-serviceGo · built in this project
Deposits through external payment gateways (card and crypto), webhooks.
plan-serviceGo · built in this project
Product catalogue: challenge plans, coupons, tournament templates, broker account offers.
referral-serviceGo · built in this project
Referral attribution and commissions.
Engagement: prop, broker, tournament· 5 services
prop-serviceGo · built in this project
Prop challenge lifecycle: phases, rule evaluation from trading events, funded accounts.
broker-serviceGo · built in this project
Simulated broker account products.
tournament-serviceGo · built in this project
Trading competitions, leaderboards.
engagement-serviceGo · built in this project
Gamification.
analytics-serviceGo · built in this project
Trading statistics from trading events (time-series).
Platform and ops· 5 services
admin-serviceGo · built in this project
Admin gateway: RBAC, audit log, aggregates other services' admin APIs.
notification-serviceGo · imported and extended
Email + in-app notifications from domain events.
support-serviceGo · imported and extended
Support tickets.
content-serviceGo · built in this project
Bilingual blog/content backend.
backup-agentshell · built in this project
Scheduled encrypted off-site backups.
Live prices
- upstream broker terminals → broker-feed-bridge (pool, one per account)vendor terminal API
- broker-feed-bridge → price-feed-servicegRPC server-streaming, one stream per bridge; selector picks the active account, standbys only update health
- price-feed-service → tick store (QuestDB)write
- price-feed-service → Redislast-tick cache + pub/sub tick:{symbol}
- price-feed-service → trading-engine-servicegRPC Subscribe stream (bounded mailbox per account actor; slow actors shed, never block the stream)
- price-feed-service → Kafka market.ticks_rawKafka producer
- Kafka market.ticks_raw → market-data-serviceconsumer -> candle aggregation -> TimescaleDB hypertables; emits candle_sealed
- Redis tick:{symbol} → websocket-gatewaypub/sub, throttled per client
- websocket-gateway → browserWebSocket /ws/platform/prices
- history-backfill-source → market-data-serviceREST, backfill only (insert-if-absent; live candles win)
Orders and domain events
- browser → websocket-gatewayWebSocket op frame (place_order etc.)
- websocket-gateway → trading-engine-servicegRPC
- trading-engine-service → websocket-gatewayop_ack/op_result + account updates over Redis pub/sub account:{id}:update
- trading-engine-service → Kafka trading.* (order_placed, position_closed, breach_detected, margin_call, ...)outbox -> Kafka, Avro
- Kafka trading.* → prop-service, analytics-service, risk-service, tournament-serviceconsumer groups
- account-service → Kafka account.* (created, activated, status_changed, breached, ...)transactional outbox -> Kafka, Avro
- Kafka account.* → trading-engine-service (lifecycle consumer spawns account actors), notification-service, referral-service, othersconsumer groups
- Kafka (all domains) → notification-serviceconsumer -> email / in-app
- failed consumers → per-topic DLQretry then dead-letter
Twenty Go services, a Rust trading engine, a C# bridge to the price source, a Python service for identity-document checks and a small TypeScript history source. Most service-to-service calls are REST; the hot paths (prices into the engine, orders from the gateway) are gRPC. Every service owns its own PostgreSQL database behind PgBouncer.
Event backbone
Domain events (account created, order filled, breach detected and so on) travel through Kafka, encoded with Avro against 180 schemas in a Schema Registry. Each service writes its events to an outbox table in the same transaction as its own state change, and a publisher sends them on, so a crash can never leave a change without its event. Order placement is idempotent on a client order id, and consumers that cannot process a message send it to a per-topic dead-letter queue.
Live prices take a different path. The price-feed service streams ticks to the Rust engine over gRPC and publishes them on Redis Pub/Sub, where the WebSocket gateway I wrote picks them up and sends them to each client, throttled per connection. Market data is stored in QuestDB for ticks and TimescaleDB hypertables for candles.
Web applications and the WebGL chart
Five Nuxt applications share a monorepo with common design tokens, UI components and an API client: the trading terminal, the user dashboard (wallet, KYC), the admin back office, the marketing site and a bilingual blog.
The terminal's chart is built in-house, with no charting library: series, axes, crosshair and about 70 drawing tools. I wrote its WebGL2 series renderer:
- All coordinate maths runs in 64-bit floats in JavaScript and only pixel positions are sent to the GPU, because FX prices lose precision in 32-bit floats.
- Candles are drawn as quads grouped by colour: four draw calls by default, however many candles are on screen, from one reused vertex buffer.
- It shares its pixel maths with the Canvas 2D renderer, so the two produce the same pixels. If WebGL is unavailable, a shader fails or the context is lost for more than two seconds, the chart switches to Canvas 2D.
The trading terminal and its WebGL chart
Images of the real product; the note below says where and how they were captured.


Production operations
Production runs on Docker Swarm, deployed by GitLab CI from a runner on the host: each service is rolled forward and checked for health. I set up the monitoring stack (Prometheus with about 520 alert rules including SLO burn-rate alerts, 36 generated Grafana dashboards, Loki for logs, Tempo and OpenTelemetry for traces, and Alertmanager) and wrote the runbooks: more than 60 operational runbooks and about 130 more for individual alerts.
Backups came late. From August 2026 the databases have encrypted offsite backups and point-in-time recovery with pgBackRest.
Incident review
In August 2026 three small symptoms were reported together: lock warnings in the PostgreSQL log, connection resets, and a dozen events that failed to publish. Nothing looked broken and no alert had fired. I investigated read-only against production first, and found three separate causes and a fourth, more serious problem nobody had reported.
- Advisory locks behind a transaction pooler. Outbox publishers elected a leader with PostgreSQL session advisory locks. Behind PgBouncer in transaction mode, the lock stays on a backend the application no longer controls, so leadership could be held by two replicas or by none. Production showed zero advisory locks held while four services believed they had one.
- Connection resets. Two were explained by a redeploy; 23 could not be attributed with the logging in place, and I reported them as unattributable instead of guessing.
- Transient Kafka outages. The single broker had fenced itself three times. All twelve affected events were shown to have been delivered and handled, from four independent facts.
- Unbounded retries. Events that could never succeed (a schema mismatch, a missing topic) were retried without limit and blocked the head of their queue. One row had over 5 million attempts. Eighteen risk events had been stuck for 49 days, because the only signal was a counter that fired once.
What I changed. A shared retry library that classifies failures (transient, permanent, ambiguous, missing topic), backs off and parks permanent failures for review, live on 11 of 13 event publishers; the two money services were held for a second review. A pooler-safe lock and a Redis lease elector, adopted by the first service when my work ended. Alert rules for work that is over budget, and a repair tool for the stuck events that reconciles them instead of replaying them.
Lessons. "Healthy" is not "working"; a counter is not a backlog; a retry column that nothing reads is not a limit; and pooling mode is part of your concurrency model.
Load test
k6 load test: 100 traders placing orders over WebSocket
Published run records, shown as they were recorded.
- Orders placed
- 2,250
- all succeeded, 0 errors
- Order round trip, p95
- 60 ms
- median 55 ms · max 86 ms
- Price tick to client, p95
- 91 ms
- Account update to client, p95
- 206 ms
Order round trip (WebSocket send to result)
Setup: 100 virtual users from one machine, each with two WebSockets (account and prices), placing a market order every 30 seconds for 10 minutes, against the production stack before launch, on demo accounts.
| Threshold set before the run | Observed | Result |
|---|---|---|
| place_order_op_result_ms p(95)<50 | p95=60ms | Missed |
| ws_fanout_observed_ms{family:account} p(95)<200 | p95=206ms | Missed |
| ws_fanout_observed_ms{family:tick} p(95)<200 | p95=91ms | Met |
| place_order_success_rate rate>0.99 | 100% (2250/2250) | Met |
| op_error_total count<50000 | 0 | Met |
| ws_handshake_failures_total count<20000 | 0 | Met |
| http_req_duration{kind:ws_ticket} p(95)<1000 | p95=15ms | Met |
The four runs before this one placed no orders at all
Four consecutive runs over ~32 hours had every order rejected (0 of 2,250 succeeded each time) while WebSockets, auth and tickets all worked. The failures were not load-related: they came from the path a brand-new account takes from creation to being tradable in the in-memory engine. Fixes landed between runs; the fifth run placed all 2,250 orders.
- 2026-05-25: Test-harness plumbing: per-IP WebSocket handshake rate-limit CIDR bypass at both the edge proxy and the WS gateway (both layers must agree or the gateway/edge still 429s), and an admin bulk-seed endpoint so synthetic users go through the normal outbox -> Kafka user.created path instead of direct SQL inserts that downstream services never observe.
- 2026-05-26 early: Path-scoped IP-binding bypass in the token validator for the seed endpoint, so in-cluster seeding with a minted admin token works (signature/expiry/admin/session checks still enforced).
- 2026-05-26 midday: Event-contract drift between the Go account service and the Rust trading engine: the engine consumer read a legacy `account_kind` field while the producer had moved to a canonical 11-value `kind`. Engine now prefers `kind` and falls back to the legacy field; producer again populates the legacy fields; a jsonb cast error in the account repository was fixed. The seed script was rewritten to create demo accounts through the real API so each one flows account.created -> activate -> outbox -> Kafka -> engine.
- 2026-05-26 evening: Avro encoding: made `kind`/`status` nullable for backward compatibility and fixed nullable timestamp-millis union handling in the goavro encoder (with a regression test).
- 2026-05-26 evening: Engine HA mode never flipped to Live: the readiness watcher only accepted Replaying, but HA mode skips the cold-start step that sets it, so the precondition now accepts Booting or Replaying. Also fixed UUID hex formatting in the seed script so bearer minting and user lookup agree, and added logging of every op_error frame in the k6 script.
- 2026-05-27 01:03 (+04:00), ~1 h before the successful run: Root cause of the remaining rejections: actors created by the live lifecycle consumer for brand-new accounts were never registered in the symbol->account tick-routing index (the cold-start and recovery paths did this; the live path did not). Without ticks the actor's last-tick cache stayed empty and every PlaceOrder was rejected as SESSION_CLOSED. Fix: register every new actor for every symbol in the live symbol book. A latent bug that only affects fresh accounts that have never traded, which is exactly what a load test with freshly seeded accounts creates.
- Single load-generator machine: all 100 VUs came from one client, and per-IP rate limits had to be bypassed for it. It does not model 100 independent networks.
- Ran against the production stack (before public launch) using seeded demo accounts, not staging as the script header recommends; no real-money accounts were involved.
- Order latency is client-perceived: WS send -> op_result, including network, edge proxy and WS gateway, not engine-internal time.
- WS fan-out latency compares client clock to server timestamps, so it includes any cross-host clock skew.
- Two thresholds were missed: order p95 was 60 ms against a 50 ms target, and account-event fan-out p95 was 206 ms against a 200 ms target. Tick fan-out (91 ms) and all success/error gates passed.
- Order latency was very tight (min 52, median 55, max 86 ms), which suggests a fixed network/hop floor rather than queueing under load.
- Load was modest: about 3.3 orders/s (one order per VU every 30 s); this is a correctness-under-concurrency and latency test, not a throughput ceiling test.
How the work was done
The design lives in 30 architecture specifications. Coding agents implemented from them, and I reviewed, tested and integrated the result. About two thirds of the commits carry an AI co-author tag. The repository has about 6,800 Go test functions, about 1,100 Rust tests and integration tests against real PostgreSQL with Testcontainers, run locally rather than in the deployment pipeline. My public example of the same way of working, with tests you can inspect, is BoundedCode.
My role
- Co-founder and technical lead; designed the architecture and wrote the specifications.
- Wrote most of the backend, including the WebSocket gateway and nearly all of the infrastructure, CI and monitoring; the other engineer wrote more of the prop and tournament services than I did.
- Built most of the five frontends, including the WebGL renderer.
- Ran production: deployments, incidents, backups.
Limitations
- Everything here is self-reported; the code and raw data are private.
- Production ran on a single node, and backups arrived in August 2026.
- The load test is one run from one machine, and it missed two of its own targets.
- The incident fixes were still being rolled out when my work ended.
Evidence index
Every claim this case study relies on, rendered from the evidence manifests.
TFLMarkets backend architecture
Designed the TFLMarkets backend and wrote most of it with one other engineer: 20 Go services, a Rust order-execution engine and a Python KYC service, communicating over REST, gRPC and Kafka.
Event backbone with outbox, idempotency and dead-letter handling
Built the event backbone on Kafka with Avro schemas in a Schema Registry and a transactional outbox in each service, plus idempotent order placement and dead-letter handling.
Real-time price path to the trading terminal
Live prices flow from the price feed to the Rust engine over gRPC and reach traders through Redis Pub/Sub and a WebSocket gateway that I wrote, with ticketed connections, heartbeats and reconnection with backoff on the client.
Time-series storage for market data
Market data is stored in TimescaleDB hypertables and QuestDB, which serve chart history to the trading terminal separately from the transactional PostgreSQL databases.
Five Nuxt applications and a custom WebGL2 chart renderer
Built five Nuxt applications in a shared monorepo (trading terminal, user dashboard, admin back office, marketing site and blog) and wrote a custom WebGL2 candlestick renderer with a Canvas 2D fallback for the trading terminal.
Monitoring stack with SLO burn-rate alerts
Set up Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager for the platform, with generated dashboards and SLO burn-rate alert rules, and wrote runbooks for the alerts.
Production deployment, backups and recovery
Set up production on Docker Swarm with GitLab CI deployments, wrote more than 60 runbooks, and added encrypted offsite backups with point-in-time recovery.
Production incident investigation
Investigated a production reliability incident, traced it to unbounded outbox retries and session advisory locks behind PgBouncer transaction pooling, bounded the retries in 11 of 13 event publishers and began replacing the locks.
Order-placement load test with k6
Load-tested order placement with k6 at 100 concurrent users: 2,250 orders, all successful, 60 ms at p95 from sending the order over the WebSocket to receiving its result.
Specification-first delivery with coding agents
Wrote 30 architecture specifications that coding agents implemented from; I owned the design, review, testing and integration.