Skip to content
Ali Akbari
Menu

Case study

Real-time trading platform

Backend and platform architecture for a live, customer-facing trading platform: a real-time market-data path, money-moving workflows built on safe event processing, time-series services, and the observability and operations work that keeps it running.
EmploymentImplementedenvironment:Production× 6 claims
Role
Platform architecture, backend and production operations
Period
2023 — present
Updated
8 October 2026
  • Go
  • Rust
  • TypeScript
  • Kafka
  • Avro
  • WebSockets
  • gRPC
  • TimescaleDB
  • QuestDB
  • PostgreSQL
  • Redis
  • Prometheus
  • Grafana
  • OpenTelemetry
Jump to a section
  1. Context
  2. Problem
  3. Scale
  4. Architecture
  5. My responsibility
  6. Constraints
  7. Key decisions
  8. Failure modes
  9. Reliability and operations
  10. Security
  11. How the work is done
  12. Results
  13. Lessons
  14. Evidence index

Context

A trading platform used by real customers with real money. Customers watch live prices and charts, and place orders. They move funds through purchase, wallet and payout flows, and pass identity (KYC) checks. The platform is a set of backend services in Go, Rust and TypeScript, with several web applications on top.

Problem

Two kinds of work share one system and pull in different directions:

  • Live data has to reach every connected customer quickly and keep flowing through partial failures.
  • Money-moving workflows have to be exactly right. A retried message must not pay twice, and a crashed service must not lose a state change it has already acknowledged.

Scale

Not disclosed. The figures are proprietary and are published neither here nor on /benchmarks.

Architecture

Real-time trading platform (sanitized)Market data is ingested into Kafka with Avro schemas, processed by streaming services, fanned out over WebSockets to web applications and stored in time-series databases for history. Transactional services for purchases, wallets, payouts and KYC write to PostgreSQL and publish events through a transactional outbox. Every service reports to a shared observability stack.outbox relayREST · gRPC · history queriesMarket dataexternal feedsIngestionnormalizeKafkaAvro · schema registryStreaming serviceslive stateWebSocket fan-outTransactional servicespurchase · wallet · payout · KYCPostgreSQLstate + outbox tableTime-series storesTimescaleDB · QuestDBWeb applicationsWebGL chartsObservabilityPrometheus · Grafana · Loki · Tempo · OpenTelemetry · Alertmanager
Figure — Generic redrawing with no employer-specific names; scale is not disclosed. Failed events retry with a bound, then go to a dead-letter queue (not drawn).

Market data is ingested into Kafka using Avro schemas governed by a schema registry. Streaming services fan data out over WebSockets to the web applications, and a custom WebGL chart renderer draws live and historical series. Historical data is served from TimescaleDB and QuestDB. Transactional services expose REST and gRPC APIs. They publish their state changes through a transactional outbox, so the database write and the event can never disagree.

My responsibility

  • Platform architecture and engineering standards, and support for production incident resolution.
  • The real-time data path, end to end.
  • Streaming and historical data services on TimescaleDB and QuestDB.
  • APIs (REST, gRPC, WebSocket) with schema governance, and the event-processing patterns behind them.
  • The observability stack and alerting.
  • Deployments, CI/CD, monitoring, incident investigation, backups and recovery.

This is team work: I lead architecture and engineering standards, and mentor engineers through reviews and written runbooks.

Constraints

  • Compliance-sensitive flows (payments, payouts, KYC) ship only through controlled, traceable releases.
  • Live data and transactional workloads run side by side; neither may starve the other.

Key decisions

Failure modes

Failure classMechanism that handles it
Duplicate delivery after a retryIdempotent consumers
Service crash between a database write and its eventTransactional outbox
Message that can never be processedBounded retries, then a dead-letter queue
Incompatible event changeSchema governance at the registry
Failure that passes ordinary health checksAlerts on customer-facing symptoms, investigated with metrics, logs and traces

Reliability and operations

Alerts are tied to customer impact rather than to individual component metrics. Several failures were found this way that ordinary health checks had passed. The stack is Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager. Order placement and WebSocket fan-out are load-tested with k6. Backups and recovery are part of my operational responsibility, together with deployments and CI/CD.

Security

Money-moving and identity workflows ship through controlled, traceable releases, and every change is code reviewed. Further security details of a live financial system are deliberately not published.

How the work is done

Development is specification-first. Written architecture specifications guide AI coding agents, and every change is checked by automated integration, end-to-end and load tests before release. Runbooks are written for the operations they describe. My public example of the same discipline, with inspectable tests, is BoundedCode.

Results

Qualitative only: the platform runs live for real customers, and money-moving flows ship through controlled releases. No figures are published here.

Lessons

  • A health check that tests only liveness gives false confidence. Alert on what customers experience.
  • The outbox pattern is cheaper than debugging a lost event in a payment flow.
  • Write the specification before the code, especially when an AI agent writes much of the code.

Evidence index

Every claim this case study relies on, rendered from the evidence manifests.

Implementedenvironment:ProductionSelf-reported · proprietary

Real-time market-data path in production

Built and operate an end-to-end real-time data path for a live trading platform: ingestion into Kafka with Avro schemas, WebSocket streaming to several web applications, and a custom WebGL chart renderer for live and historical data.

Distributed systemsEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · proprietary

Safe event processing for money-moving workflows

Delivered REST, gRPC and WebSocket APIs with schema governance, and event processing using a transactional outbox, idempotency, retries and dead-letter queues for purchase, wallet, payout and KYC workflows.

BackendEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · proprietary

Streaming and historical time-series services

Designed streaming and historical data services on TimescaleDB and QuestDB that serve live subscriptions and time-series queries in production.

Data systemsEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · proprietary

Observability stack with customer-impact alerting

Built and maintain a production observability stack — Prometheus, Grafana, Loki, Tempo, OpenTelemetry and Alertmanager — with alerts tied to customer impact, and used it to find the root cause of failures that passed ordinary health checks.

ObservabilityEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · proprietary

Production operations for a live trading platform

Own deployments, CI/CD, monitoring, incident investigation, backups and recovery for a live, customer-facing trading platform.

SREEmployment · no public artifactsEvidence
Implementedenvironment:ProductionSelf-reported · proprietary

Specification-first, AI-assisted delivery with test gates

Run an AI-first engineering workflow in production: written architecture specifications guide coding agents, and every change is checked with automated integration, end-to-end and load tests before release.

VerificationEmployment · no public artifactsEvidence

← All work