Jump to a section
Overview
MyAgent runs around the clock on one workstation under Docker Compose and hosts two agents that share one safety core. One finds job postings, checks eligibility rules and fit, and drafts applications. The other audits small-company websites and drafts proposals that may only cite findings it has verified. Neither sends anything without a human pressing a button on Telegram.
It is a prototype: it runs end to end against local fake services in dry-run mode, which is the default. Live sending needs three separate switches, and I have no record of it being used. It was built in a few days with heavy use of AI coding agents, to a design I wrote.
Architecture
Go services built into one image, each running a different binary: a control plane (policy, approvals, the ledger), a reconciler, a Telegram gateway, Gmail sync and send, a browser executor, a discovery worker and a model router. Python services run Playwright and a read-only crawler. PostgreSQL is the system of record and Temporal runs the workflows. Each service gets only the credentials and database role it needs: the Telegram gateway has no database access, and the model router is the only service that holds model credentials.
The action ledger
Every external effect is a row in one table. A PostgreSQL trigger rejects any state change that is not in the allowed list, rows can never be deleted, and the content is frozen once the action is proposed: an edit creates a new action and supersedes the old one.
Follow an action through the ledger
Drawn from the source code and configuration; nothing here is running.
- PROPOSED
PROPOSED · in progress · no external effect can have happened yet
The trigger allows 5 moves from here:
Colour key: amber = an effect may have happened, so retrying blindly is unsafe; green = sent and verified; red = failed after an attempt; grey = ended before any effect.
- At most once. Execution has exactly one attempt. A crash or timeout while executing moves the action to UNKNOWN, never back to a retry.
- Reconcile, don't retry. A reconciler for each kind of action checks what actually happened, for example whether the email is in the sent folder, and only allows a retry with proof that nothing was sent. After repeated inconclusive checks it asks a human.
- Approvals bound to content. Each Telegram button carries a single-use token tied to a hash of the exact content. If the content changed, the approval is void.
- Tamper-evident audit. Audit rows are append-only and hash-chained, and the daily digest carries the chain head.
Choosing a local model
The models run locally with llama.cpp on an 8 GB laptop GPU. I compared two models on the agent's own tasks, using its production prompts, JSON schemas and validators:
Two local models on the agent's own tasks
Published run records, shown as they were recorded.
| Task | Metric | Qwen3 4B | Qwen3.6 35B-A3B | p50 latency (4B / 35B) |
|---|---|---|---|---|
| reply_classification | class_agreement | 0.82 (n=50) | 0.967 (n=30) | 0.7 s / 2.4 s |
| job_extract | nationality_restriction / min_years_within_1 / technology_f1 | 0.79 / 0.969 / 0.536 (n=100) | 0.933 / 1.00 / 0.644 (n=30) | 3.0 s / 7.5 s |
| fit_score | spearman_vs_reference | 0.842 (n=100) | 0.858 (n=30) | 2.2 s / 6.9 s |
| website_analysis | json_valid_rate / software_company / opportunity_type_jaccard | 0.42 / 1.00 / 0.397 (n=50) | 1.00 / 1.00 / 0.301 (n=30) | 39.9 s / 12.4 s |
Peak GPU memory: 4988 MiB for the 4B model, 7560 MiB for the 35B MoE model with its expert layers on the CPU. Outcome recorded in the project docs: the 35B model handles the safety-relevant tasks; the 4B model only triages URLs.
- The 35B figures come from the first 30 items of each set (time budget); the 4B figures from the full sets. Not like for like.
- The 4B model's website-analysis output was truncated at its token limit, so most of its answers were invalid JSON.
- fit_score is compared against one annotator's rubric score, which is an ordinal anchor, not ground truth.
Tests
132 Go test functions (including the workflow under Temporal's test environment, database permission checks per role, and an end-to-end crash scenario) and 32 Python tests, all run locally.
Limitations
- A prototype with a squashed history; not used for real sending.
- Known gaps are documented in the project: some internal services accept any caller on their network, several services share a database role, and there is no TLS inside Docker.
- The model comparison is small, single-run and not like for like (see the notes under the table).
Evidence index
Every claim this case study relies on, rendered from the evidence manifests.
A database-enforced ledger for every external action
In MyAgent, every email, form submission or application is a row in one action ledger whose 16-state, 38-transition machine is enforced by a PostgreSQL trigger. A Temporal workflow drives each action, executes it at most once, and sends anything ambiguous to UNKNOWN for reconciliation instead of retrying blindly.
Choosing a local model with the agent's own tasks
Benchmarked two local models (Qwen3 4B and Qwen3.6 35B-A3B MoE) on an 8 GB laptop GPU against a hand-labelled evaluation set, using the agent's production prompts, schemas and validators, and chose the model per task from the results.