Skip to content
Ali Akbari
Menu

Case study

MyAgent

A private prototype of a personal agent platform that runs on one workstation. Models only extract, classify and draft; every real-world action goes through a database-enforced ledger, needs a human approval on Telegram and runs at most once.
Personal labImplementedenvironment:Lab× 2 claims
Role
Sole developer (private prototype)
Period
Oct 2026
Updated
11 October 2026
  • Go
  • Temporal
  • PostgreSQL
  • Python
  • Playwright
  • llama.cpp
  • Docker Compose
  • Prometheus
  • Grafana
Jump to a section
  1. Overview
  2. Architecture
  3. The action ledger
  4. Choosing a local model
  5. Tests
  6. Limitations
  7. Evidence index

Overview

MyAgent runs around the clock on one workstation under Docker Compose and hosts two agents that share one safety core. One finds job postings, checks eligibility rules and fit, and drafts applications. The other audits small-company websites and drafts proposals that may only cite findings it has verified. Neither sends anything without a human pressing a button on Telegram.

It is a prototype: it runs end to end against local fake services in dry-run mode, which is the default. Live sending needs three separate switches, and I have no record of it being used. It was built in a few days with heavy use of AI coding agents, to a design I wrote.

Architecture

Go services built into one image, each running a different binary: a control plane (policy, approvals, the ledger), a reconciler, a Telegram gateway, Gmail sync and send, a browser executor, a discovery worker and a model router. Python services run Playwright and a read-only crawler. PostgreSQL is the system of record and Temporal runs the workflows. Each service gets only the credentials and database role it needs: the Telegram gateway has no database access, and the model router is the only service that holds model credentials.

The action ledger

Every external effect is a row in one table. A PostgreSQL trigger rejects any state change that is not in the allowed list, rows can never be deleted, and the content is frozen once the action is proposed: an edit creates a new action and supersedes the old one.

Visualization from source

Follow an action through the ledger

Drawn from the source code and configuration; nothing here is running.

  1. PROPOSED

PROPOSED · in progress · no external effect can have happened yet

The trigger allows 5 moves from here:

Colour key: amber = an effect may have happened, so retrying blindly is unsafe; green = sent and verified; red = failed after an attempt; grey = ended before any effect.

States and transitions copied from the migration that installs the guard trigger in MyAgent's private repository. Every move offered here is one the database accepts; any other is rejected with an error.
  • At most once. Execution has exactly one attempt. A crash or timeout while executing moves the action to UNKNOWN, never back to a retry.
  • Reconcile, don't retry. A reconciler for each kind of action checks what actually happened, for example whether the email is in the sent folder, and only allows a retry with proof that nothing was sent. After repeated inconclusive checks it asks a human.
  • Approvals bound to content. Each Telegram button carries a single-use token tied to a hash of the exact content. If the content changed, the approval is void.
  • Tamper-evident audit. Audit rows are append-only and hash-chained, and the daily digest carries the chain head.

Choosing a local model

The models run locally with llama.cpp on an 8 GB laptop GPU. I compared two models on the agent's own tasks, using its production prompts, JSON schemas and validators:

Recorded results

Two local models on the agent's own tasks

Published run records, shown as they were recorded.

TaskMetricQwen3 4BQwen3.6 35B-A3Bp50 latency (4B / 35B)
reply_classificationclass_agreement0.82 (n=50)0.967 (n=30)0.7 s / 2.4 s
job_extractnationality_restriction / min_years_within_1 / technology_f10.79 / 0.969 / 0.536 (n=100)0.933 / 1.00 / 0.644 (n=30)3.0 s / 7.5 s
fit_scorespearman_vs_reference0.842 (n=100)0.858 (n=30)2.2 s / 6.9 s
website_analysisjson_valid_rate / software_company / opportunity_type_jaccard0.42 / 1.00 / 0.397 (n=50)1.00 / 1.00 / 0.301 (n=30)39.9 s / 12.4 s

Peak GPU memory: 4988 MiB for the 4B model, 7560 MiB for the 35B MoE model with its expert layers on the CPU. Outcome recorded in the project docs: the 35B model handles the safety-relevant tasks; the 4B model only triages URLs.

  • The 35B figures come from the first 30 items of each set (time budget); the 4B figures from the full sets. Not like for like.
  • The 4B model's website-analysis output was truncated at its token limit, so most of its answers were invalid JSON.
  • fit_score is compared against one annotator's rubric score, which is an ordinal anchor, not ground truth.
From the bench reports of 1 October 2026 (local, not published). Hardware: RTX 4060 Laptop GPU with 8 GB, llama.cpp. The labelled evaluation set is hand-made and fictional; one run per model.

Tests

132 Go test functions (including the workflow under Temporal's test environment, database permission checks per role, and an end-to-end crash scenario) and 32 Python tests, all run locally.

Limitations

  • A prototype with a squashed history; not used for real sending.
  • Known gaps are documented in the project: some internal services accept any caller on their network, several services share a database role, and there is no TLS inside Docker.
  • The model comparison is small, single-run and not like for like (see the notes under the table).

Evidence index

Every claim this case study relies on, rendered from the evidence manifests.

Implementedenvironment:LabSelf-reported · private source

A database-enforced ledger for every external action

In MyAgent, every email, form submission or application is a row in one action ledger whose 16-state, 38-transition machine is enforced by a PostgreSQL trigger. A Temporal workflow drives each action, executes it at most once, and sends anything ambiguous to UNKNOWN for reconciliation instead of retrying blindly.

Distributed systemsPersonal lab · no public artifactsEvidence
Implementedenvironment:LabSelf-reported · private source

Choosing a local model with the agent's own tasks

Benchmarked two local models (Qwen3 4B and Qwen3.6 35B-A3B MoE) on an 8 GB laptop GPU against a hand-labelled evaluation set, using the agent's production prompts, schemas and validators, and chose the model per task from the results.

AI infrastructurePersonal lab · no public artifactsEvidence

← All work