Jump to a section
Context
A trading platform used by real customers with real money. Customers watch live prices and charts, and place orders. They move funds through purchase, wallet and payout flows, and pass identity (KYC) checks. The platform is a set of backend services in Go, Rust and TypeScript, with several web applications on top.
Problem
Two kinds of work share one system and pull in different directions:
- Live data has to reach every connected customer quickly and keep flowing through partial failures.
- Money-moving workflows have to be exactly right. A retried message must not pay twice, and a crashed service must not lose a state change it has already acknowledged.
Scale
Not disclosed. The figures are proprietary and are published neither here nor on /benchmarks.
Architecture
Market data is ingested into Kafka using Avro schemas governed by a schema registry. Streaming services fan data out over WebSockets to the web applications, and a custom WebGL chart renderer draws live and historical series. Historical data is served from TimescaleDB and QuestDB. Transactional services expose REST and gRPC APIs. They publish their state changes through a transactional outbox, so the database write and the event can never disagree.
My responsibility
- Platform architecture and engineering standards, and support for production incident resolution.
- The real-time data path, end to end.
- Streaming and historical data services on TimescaleDB and QuestDB.
- APIs (REST, gRPC, WebSocket) with schema governance, and the event-processing patterns behind them.
- The observability stack and alerting.
- Deployments, CI/CD, monitoring, incident investigation, backups and recovery.
This is team work: I lead architecture and engineering standards, and mentor engineers through reviews and written runbooks.
Constraints
- Compliance-sensitive flows (payments, payouts, KYC) ship only through controlled, traceable releases.
- Live data and transactional workloads run side by side; neither may starve the other.
Key decisions
Failure modes
| Failure class | Mechanism that handles it |
|---|---|
| Duplicate delivery after a retry | Idempotent consumers |
| Service crash between a database write and its event | Transactional outbox |
| Message that can never be processed | Bounded retries, then a dead-letter queue |
| Incompatible event change | Schema governance at the registry |
| Failure that passes ordinary health checks | Alerts on customer-facing symptoms, investigated with metrics, logs and traces |
Reliability and operations
Alerts are tied to customer impact rather than to individual component metrics. Several failures were found this way that ordinary health checks had passed. The stack is Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager. Order placement and WebSocket fan-out are load-tested with k6. Backups and recovery are part of my operational responsibility, together with deployments and CI/CD.
Security
Money-moving and identity workflows ship through controlled, traceable releases, and every change is code reviewed. Further security details of a live financial system are deliberately not published.
How the work is done
Development is specification-first. Written architecture specifications guide AI coding agents, and every change is checked by automated integration, end-to-end and load tests before release. Runbooks are written for the operations they describe. My public example of the same discipline, with inspectable tests, is BoundedCode.
Results
Qualitative only: the platform runs live for real customers, and money-moving flows ship through controlled releases. No figures are published here.
Lessons
- A health check that tests only liveness gives false confidence. Alert on what customers experience.
- The outbox pattern is cheaper than debugging a lost event in a payment flow.
- Write the specification before the code, especially when an AI agent writes much of the code.
Evidence index
Every claim this case study relies on, rendered from the evidence manifests.
Real-time market-data path in production
Built and operate an end-to-end real-time data path for a live trading platform: ingestion into Kafka with Avro schemas, WebSocket streaming to several web applications, and a custom WebGL chart renderer for live and historical data.
Safe event processing for money-moving workflows
Delivered REST, gRPC and WebSocket APIs with schema governance, and event processing using a transactional outbox, idempotency, retries and dead-letter queues for purchase, wallet, payout and KYC workflows.
Streaming and historical time-series services
Designed streaming and historical data services on TimescaleDB and QuestDB that serve live subscriptions and time-series queries in production.
Observability stack with customer-impact alerting
Built and maintain a production observability stack — Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager — with alerts tied to customer impact, and used it to find the root cause of failures that passed ordinary health checks.
Production operations for a live trading platform
Own deployments, CI/CD, monitoring, incident investigation, backups and recovery for a live, customer-facing trading platform.
Specification-first, AI-assisted delivery with test gates
Run an AI-first engineering workflow in production: written architecture specifications guide coding agents, and every change is checked with automated integration, end-to-end and load tests before release.