Skip to content
Ali Akbari
Menu

Case study

Soul

Architecture and foundation plan for a privacy-first mobile app for conversational personal support. Design only: the documents are complete and frozen, and nothing has been built.
Personal labDesignedenvironment:Lab× 1 claims
Role
Author of the architecture and foundation plan
Period
Sep 2026
Updated
11 October 2026
  • Go
  • Kubernetes (design)
  • PostgreSQL
  • Kafka
  • Expo / React Native
Jump to a section
  1. Overview
  2. Architecture in brief
  3. Decisions and trade-offs
  4. Review of the design
  5. Limitations
  6. Evidence index

Overview

Soul is a planned mobile and web app in which people talk through everyday personal challenges and receive supportive guidance and small practical tasks, grounded in approved books, professional resources and international guidelines. I wrote its architecture (version 1.0-rc, frozen on 28 September 2026), a review of how that design would age, and a foundation execution plan. The documents were written with heavy AI assistance under my direction.

Nothing has been implemented. This page is here to show how I approach a design, not to claim experience with the technologies in it.

Architecture in brief

  • Clients: one Expo / React Native codebase for iOS, Android and the web.
  • Backend: stateless Go services, as a modular monolith by default; any additional service needs a written decision.
  • State: one logical PostgreSQL database, Kafka as a durable backbone that is never the system of record, and a cache for ephemeral data only.
  • Security: an isolated key service, workload identity, sender-bound tokens, field-level and object-level encryption, and account deletion by destroying the user's keys.
  • Hosting: starts on a single server in Germany and grows by adding nodes to the same cluster.

Decisions and trade-offs

Design only

Key architecture decisions

Taken from design documents. Nothing has been implemented.

  1. 01Start on one node, grow the same cluster

    Production starts on a single Debian 13 server running a single-node RKE2 cluster in Germany. Growth comes from adding RKE2 nodes to the same cluster; the cluster is never rebuilt to grow.

    Why, and what it costs

    Why: Node growth must never change contracts (code, DSNs, endpoints, workload identities, schemas, encryption formats, client protocol). Services use only Kubernetes Service DNS, never node IPs or host networking, so the same architecture runs from Stage 1 to large scale.

    Trade-off: Stage 1 is explicitly not highly available, and Kubernetes plus the full security plane carry a large fixed footprint before any users. The longevity review estimates 11-15 GiB idle and roughly 30 components for one operator, mitigated by demand-driven activation (A9).

  2. 02One logical PostgreSQL database

    All user and business data lives in one logical PostgreSQL database ('soul') behind stable logical endpoints. The only exception is an isolated key-store cluster that holds just wrapped keys and key-lifecycle metadata, enforced by a CI schema allowlist.

    Why, and what it costs

    Why: Clear sources of truth. Replicas, failover, pooling and any future distribution stay invisible to the application.

    Trade-off: The documents admit the one-million-connection design point is likely beyond one primary. The review found that the planned horizontal exit (Citus) conflicted with other frozen decisions until A10 was added.

  3. 03User-local OLTP transactions (A10)

    Every user-scoped OLTP transaction is single-user and distribution-local. Every user-transactional table, including outbox rows, jobs, idempotency records and session state, leads with user_id. Two-phase commit stays disabled (max_prepared_transactions = 0).

    Why, and what it costs

    Why: Added after the longevity review so that sharding by user_id later is a bounded migration rather than a multi-layer refactor. The product has almost no cross-user data.

    Trade-off: Any atomic transaction that spans users needs a superseding ADR, and system-scoped jobs and events need separate tables or contracts.

  4. 04Kafka transports, Postgres decides

    Kafka is the durable asynchronous backbone but not the system of record. Durable job rows in PostgreSQL are authoritative, and the outbox-to-Kafka event is only a wake-up and distribution signal.

    Why, and what it costs

    Why: Losing Kafka never loses a job. Kafka (and Valkey) can be replaced by building a new cluster, cutting over and discarding the old one.

    Trade-off: Extra machinery and write amplification: outbox, job leases and fencing. The design never claims exactly-once delivery without a mechanism that proves it.

  5. 05Valkey is ephemeral

    Valkey is never authoritative. At Stage 1 it is a standalone, ephemeral instance, and at Stage 2 it is replaced by a fresh cluster rather than expanded in place.

    Why, and what it costs

    Why: Correctness must never depend on caches, retries or perfect delivery. 'Push is a hint, pull is the truth.'

    Trade-off: A Valkey outage or a reconnect storm falls back to the Sync path. The review flags that Sync is not yet defined as a cheap delta and must be before clients ship.

  6. 06Dedicated key service and crypto-erasure

    Sensitive fields and objects use envelope encryption with versioned AAD schemas. Keys are handled only by an isolated key-service (SPIFFE mTLS) backed by OpenBao. Account deletion zeroes the wrapped DEKs and is recorded in an off-site destruction ledger kept with two providers.

    Why, and what it costs

    Why: Deletion becomes provable and bounded: access is denied in about 1 s (10 s worst case), and erasure becomes irreversible across backups once backup retention has elapsed.

    Trade-off: Each interaction needs extra key-service RPCs. OpenBao unwrap load at high concurrency is not modelled yet, and key-custody ceremonies need rare expertise.

  7. 07Provider-independent synchronous safety gate

    Safety Gate A (deterministic rules plus a local classifier) runs synchronously on every message, with no dependency on an AI provider and a 40 ms p99 objective. On failure or timeout it switches to a scripted safe path. Gate B checks AI output chunk by chunk.

    Why, and what it costs

    Why: If the AI pipeline is unavailable, the safety check is never bypassed. A scripted response is a degraded but safe mode.

    Trade-off: Capacity has to be reserved for both gates, and a scripted fallback gives users a less helpful reply during outages.

  8. 08Invariants enforced by the system

    Rules such as maximum transaction duration, topology independence, import boundaries and certificate-approval policy are enforced by database limits, CI analyzers, conftest/Kyverno admission policies and lint rules.

    Why, and what it costs

    Why: Conventions erode as a team grows. Mechanical enforcement is what the review credits for the absence of single-node leaks into application code.

    Trade-off: A heavy upfront Foundation scope for one person before any product value ships. The review calls this the main schedule risk.

Paraphrased from the architecture documents (v1.0-rc, design frozen on 28 September 2026) and their longevity review. These are design decisions only; no code exists.

Review of the design

A separate review stress-tested the frozen design against growth. Its conclusion was "yes, with conditions": the contracts hold as the system grows, but the planned path for scaling the database out conflicted with four other decisions. The fix was cheap because no data exists yet, and it was folded into the design as an amendment. The main remaining risk is operational: one person would be running about 30 components.

Limitations

  • Design only; no prototype, no tests, no measurements.
  • The review was mine, not an external one.

Evidence index

Every claim this case study relies on, rendered from the evidence manifests.

Designedenvironment:LabSelf-reported · private source

Architecture design for a privacy-first mobile platform

Wrote the architecture and foundation plan for Soul, a mobile app for conversational personal support: Go services on Kubernetes, one logical PostgreSQL database, Kafka as a durable backbone, field-level encryption and crypto-erasure for deletion, followed by an adversarial review of how the design would age.

Distributed systemsPersonal lab · no public artifactsEvidence

← All work