Skip to content
Ali Akbari
Menu

Case study

BoundedCode

An AI software-engineering control plane in Go for long-running work on large repositories: bounded context, a network-less sandbox, a crash-safe task ledger, and a verification gate that refuses to call a change done just because the tests are green.
Personal labTested in CIenvironment:Lab× 8 claims
Role
Sole designer and maintainer
Period
Oct 2026 — public alpha
Source
akynte/boundedcodev0.1.0-alpha.3 · be1aa9f
Updated
8 October 2026
  • Go
  • SQLite
  • Docker
  • llama.cpp
  • OpenHands SDK
  • gitleaks
  • GitHub Actions
Jump to a section
  1. Problem
  2. Context
  3. Constraints
  4. Architecture
  5. Hard problems
  6. Decisions
  7. Rejected alternatives
  8. Failure model
  9. Security model
  10. Observability
  11. Tests and verification
  12. Benchmarks
  13. Incidents
  14. AI assistance and human verification
  15. Reproduce
  16. Limitations
  17. At 10x scale
  18. Source and evidence
  19. Evidence index

Problem

Coding agents working on large codebases fail in two predictable ways. They overflow context: pouring the repository into the prompt exceeds what a local model can hold and drives up the cost of frontier models. And they mistake green tests for done: an agent reports success because the suite passes, even when no test touches the code it changed.

BoundedCode is a Go control plane around existing tools — llama.cpp for local inference, the OpenHands agent SDK, a code-graph index and language servers — that is built around the opposite defaults.

Context

It is meant for one engineer running long tasks against real repositories on their own machine, with a local model by default and a cloud model only when they choose one. It is a public alpha: validated on a small held-out sample of real tasks, on one machine, with one model.

Constraints

  • One laptop: an 8 GB consumer GPU and 64 GB of RAM. The local model has to be useful within that.
  • The agent's code is untrusted. It may follow prompt-injected instructions found in the repository.
  • Tasks run for a long time and must survive Ctrl-C, crashes and reboots without losing work.
  • Nothing is pushed or merged on the user's behalf.

Architecture

BoundedCode task loopA task enters the Go control plane, which records everything in a SQLite ledger. The context planner builds a bounded pack for the agent, which runs in a network-less container and reaches the model only through the control plane's gateway over stdio. Edits land on a task branch. Verification runs in the sandbox; failures go back to the planner, and passing changes go through the behavioural-evidence check, which ends in task_verified only when a new test fails before and passes after.sandbox · no networkstatecommitspassedfailed → retry packTaskchat or CLIControl planeGo · budgets · policyTask ledgerSQLite · leasesContext plannerbounded packAgentterminal · edit toolsTask worktreeagent/<task> branchModel gatewaystdioLocal modelllama.cppVerificationtargeted → fullEvidence checkfail before · pass aftertask_verifiedtests_green (unverified)
Figure — BoundedCode's loop (redrawn from the project's README).

Each attempt starts from a small, task-specific context pack. The agent runs in a container with no network and has only terminal, file-edit and task-tracker tools; its model calls travel back over stdio to a gateway in the control plane. Work lands as commits on a task branch in its own git worktree. Verification runs targeted, then full, inside the sandbox, followed by the behavioural-evidence check. A SQLite ledger records every attempt so a task can resume anywhere.

Hard problems

Telling "verified" apart from "green". The obvious gate — run the tests — accepts a change that no test exercises. The gate instead re-runs the change's new or modified tests against the base commit: only a test that fails there and passes with the change counts. Tests that also pass on the base, or that do not even build on the base, are not credited.

Surviving a SIGKILL mid-turn. A killed runner leaves an active lease, a half-finished attempt and uncommitted edits. A second runner must not race the first, must take over once the lease is stale, and must resume from authoritative state rather than from the dead process's memory.

Containing an agent that has been turned against you. Text in the repository can instruct the agent to weaken verification or exfiltrate secrets. Configuration is therefore read from the base commit, protected paths are enforced, and context reads are confined so a planted symlink cannot pull a host secret into the model's context.

Decisions

  • ADR-0002 — embedded SQLite for control-plane state: one file, transactional, no service to run.
  • ADR-0003 — the agent runs in a container sandbox, not on the host.
  • ADR-0009 — frontier escalation is an optional, policy-triggered exception.

Rejected alternatives

Failure model

FailureBehaviourEvidence
Runner killed mid-taskLease blocks a concurrent runner; a stale lease is taken over; the attempt is resumed from the ledgerboundedcode-crash-resume
Change passes tests that do not exercise itClassified tests-green (unverified), not task-verifiedboundedcode-fail-before-pass-after
Agent rewrites verification configIgnored: config comes from the base commit; the path is protectedboundedcode-prompt-injection-containment
Git admin directory redirectedTask blocks with an integrity decision before host git runsboundedcode-task-isolation
Secret scanner missingFull gate errors instead of passing silentlyboundedcode-secret-handling
Out of attempts, tokens or timeTask is blocked, not failed; task resume continues it—

Security model

The agent and everything it runs are untrusted. The container gets no network, drops all capabilities, runs as a non-root user with no-new-privileges, and never mounts the home directory, ~/.ssh or the Docker socket. Secret paths are masked. Diffs are secret-scanned on the host. Escalation packets have API keys and host paths removed.

Limitations · from manifest boundedcode-sandbox-hardening

  • CI verifies argument construction only. The live container escape probes require Docker and are skipped in CI.
  • Isolation ultimately depends on the container runtime and host kernel.
  • Container escape resistance verified in CI: NOT VERIFIED (runtime probes run only locally).

Observability

Every task writes an audit log to the ledger: attempts, strategies, verification stages, escalations and recoveries (task.recovered events after a crash). The validation reports are built from those records.

Tests and verification

CI on every push to main runs gofmt, go vet, go test ./... and go test -race ./..., govulncheck, golangci-lint, license checks, cross-builds for six platforms and a gitleaks scan of the full history. The claims on this page cite the specific test functions that assert them. At the pinned release, all CI jobs succeeded.

Tests that need a real container engine are skipped in CI. The live sandbox escape probes, for example, run only locally. That limit is stated on the claims it affects.

Benchmarks

BoundedCode's validation results are published as reports with machine-readable run records in the repository. They are task-outcome evaluations, not latency or throughput benchmarks, so they are linked as documents rather than shown on /benchmarks. As the README states: on six held-out tasks not used during development, five ended task-verified and passed the datasets' hidden acceptance tests, all five using only the local model. That is a small sample, not a statistical evaluation. Earlier rounds scored 0 of 8 and then 1 of 8 on differently screened tasks, and in development runs the gate classified five changes as task-verified that then failed hidden tests; the reports explain each case. Fail-before/pass-after is evidence, not proof of correctness.

Incidents

No incident replays are published for this project.

AI assistance and human verification

From the project's README: "I designed the architecture, the threat model, the verification model and the evaluation protocol, and made the release and scope decisions. Implementation, test runs and first drafts of the reports were produced with heavy use of AI coding agents, under that design and review." The validation reports are published unedited, including results that did not support a release.

Reproduce

git clone https://github.com/akynte/boundedcode.git && cd boundedcode
git checkout v0.1.0-alpha.3
go test ./...          # what CI runs, minus Docker-gated tests
go test -race ./...

The repository README documents the full set-up for running tasks with a local model (Linux, Docker, an NVIDIA GPU for local inference, or a cloud model API instead).

Limitations

  • Validated on six held-out tasks, with one model, on one machine. It shows no proven advantage over the same model run without BoundedCode.
  • macOS and Windows are experimental, and some test packages fail there.
  • Frontier escalation is implemented and tested against a fake, but it has not been shown to improve outcomes.

At 10x scale

The single SQLite ledger and per-machine sandbox are deliberate for a single-user tool. Running many engineers' tasks would first need a shared task store with real leases, such as Postgres with row locks, and a scheduler for GPU time. The control plane's state boundaries were drawn so those parts could change without changing the verification gate.

Source and evidence

Repository akynte/boundedcode at release v0.1.0-alpha.3. Every claim above links to files at commit be1aa9f. See the evidence matrix at /proof.

Evidence index

Every claim this case study relies on, rendered from the evidence manifests.

Tested in CIenvironment:Lab

Fail-before / pass-after verification gate

A coding-agent change is classified TASK_VERIFIED only when a test added or changed by that change fails on the base commit and passes with it; code-only changes, tests that already pass on the base, and tests that do not build on the base are not credited.

Verificationboundedcode v0.1.0-alpha.32 test files, 6 artifactsEvidence
Tested in CIenvironment:Lab

Crash-safe task ledger with lease takeover

After the runner process is killed with SIGKILL mid-task, a second runner is refused while the lease is fresh, takes over once it is stale, marks the orphaned attempt rejected, keeps the killed attempt's edits, and completes the task from a resume pack rebuilt from the SQLite ledger.

Backendboundedcode v0.1.0-alpha.33 test files, 9 artifactsEvidence
Tested in CIenvironment:Lab

Task-branch isolation and git integrity checks

Agent work is checkpointed as commits on a task branch without modifying main, and a worktree whose git admin directory has been redirected to a planted repository blocks the task with an integrity decision before host git can run the planted filter.

Securityboundedcode v0.1.0-alpha.32 test files, 5 artifactsEvidence
Tested in CIenvironment:Lab

Hardened container arguments for agent execution

The container runner always builds engine arguments with no network, all capabilities dropped, no-new-privileges, a non-root user, read-only mounts where requested and tmpfs masks over secret paths, and refuses sensitive host mounts such as the home directory, ~/.ssh and the Docker socket.

Securityboundedcode v0.1.0-alpha.31 test file, 5 artifactsEvidence
Tested in CIenvironment:Lab

Containment of a prompt-injected agent

An agent that follows injected instructions cannot weaken verification — a rewritten verification config is ignored because configuration is read from the base commit and the path is protected — and a host secret behind a planted symlink never reaches a context pack.

Securityboundedcode v0.1.0-alpha.34 test files, 8 artifactsEvidence
Tested in CIenvironment:Lab

Fail-closed secret handling

The full verification gate errors when no secret scanner is available instead of passing silently, secret-looking paths (.env files, keys, Terraform state) are classified as secrets while ordinary source files are not, and API keys are redacted from escalation packets.

Securityboundedcode v0.1.0-alpha.33 test files, 7 artifactsEvidence
Tested in CIenvironment:Lab

Token-budgeted context packs

The context-pack builder keeps its estimated token count within the configured budget, truncating sections when inputs exceed it, and renders the same pack for the same inputs.

AI infrastructureboundedcode v0.1.0-alpha.31 test file, 3 artifactsEvidence
Tested in CIenvironment:Lab

Deterministic, budgeted model-escalation policy

Whether a task escalates from the local model to a frontier model is decided by deterministic rules: routine tasks stay local, repeated failures and high-risk changes escalate, a call budget caps escalation, and an explicit user request overrides the budget.

AI infrastructureboundedcode v0.1.0-alpha.32 test files, 6 artifactsEvidence

← All work