Method
How I work
AI-native, not AI-dependent
AI engineering protocol
I use AI coding agents every day. They scaffold implementations, generate test candidates, summarize logs, propose incident hypotheses, draft documents, review configuration and suggest alternative designs. They do not own the outcome. Before anything is merged or published, I verify architecture, security boundaries, concurrency, correctness, failure behaviour, benchmark method and numbers, and whether the tests are adequate.
Every claim here inherits that rule. AI-assisted code still has to pass review, tests, the race detector, security scans and CI. A claim reaches Tested only when a named test asserts it at a pinned commit. A passing test proves only what it asserts, so each claim names the test that supports it.
Use aggressively
Scaffolding, boilerplate, doc drafts, test generation, config review, log summaries, alternative designs.
Accepted after normal review; CI is the gate.
AI assists, I verify
Concurrency, retries, rate limiting, authentication, policies, failure handling, performance claims.
Merged only with evidence: a test (including -race), a reproducible experiment, a trace, or the documentation line that proves it.
No AI first
Troubleshooting drills, root-cause analysis, system-design practice, explaining my own work.
Attempted alone first; AI may critique afterwards.
In practice
Who did what, stated in the repository
BoundedCode's README says which parts I designed and decided, and that implementation, test runs and first drafts were produced with heavy use of AI coding agents under that design and review.
README · How this was built (GitHub, pinned to commit be1aa9f)Results published even when they were bad
The first validation of BoundedCode passed 0 of 8 real tasks on its frozen build. The report is published unedited, next to the later 5-of-6 result.
Validation report (0 of 8) (GitHub, pinned to commit be1aa9f)A gate that refuses to trust green tests
The tool enforces the same rule I hold myself to: a change only counts as verified when a test that exercises it fails before and passes after.
Claim, tests and limits
A worked example is BoundedCode: the control plane of an AI coding agent that refuses to call a change verified unless a new test fails before the change and passes after it.
Two axes, never one score
Evidence levels
Verification
- Designedenvironment:Lab
- A design doc or ADR states the intended behaviour; no working code yet.
- Implementedenvironment:Lab
- Code is merged at a tagged release and runs; no automated test asserts the claim.
- Tested in CIenvironment:Lab
- An automated test in CI asserts the claimed behaviour, including a failure case where relevant, and passes at the verified commit.
- Measuredenvironment:Lab
- Tested, plus a repeatable experiment committed as machine-readable output with method, environment and raw data.
Environment
- Lab
- Local machine, local clusters or CI runners.
- Reference
- Validated design or infrastructure code that has not been deployed.
- Deployed
- Running in a real, reachable environment I operate.
- Production
- Real users or traffic: employment history, or a project with real usage.
Context
- Employment
- Real work for an employer, sanitized. Source code is proprietary.
- Personal lab
- My own public repositories.
- Reference architecture
- Reference architecture: designed and tested, not deployed.
- Open source
- Contributions to projects I do not own.
- Employment + proprietary
- Capped at Implemented: a test nobody can inspect is not public evidence.
Checked in CI on every change
What the evidence audit enforces
| Rule | The build fails unless… |
|---|---|
| Schema | Every manifest, lock, report, résumé and front-matter file validates against a versioned JSON Schema (draft 2020-12). |
| Identity | Evidence ids are unique across all sources; every referenced id, incident, case study, diagram and directive must exist. |
| Pinning | Public evidence comes only from a source pinned to a tag and commit; the vendored lock and every vendored file must match their recorded SHA-256. |
| Paths | Every implementation, test, report and ADR path must exist at the verified commit. |
| Tested | TESTED requires named test functions that exist in the file and a CI run of the named workflow that succeeded at the verified commit. |
| Measured | MEASURED requires a benchmark-v1 report whose raw data is vendored and matches raw_sha256. Hand-edited numbers fail. |
| Combinations | PRODUCTION requires employment; reference architectures are never deployed or measured; proprietary evidence stops at IMPLEMENTED. |
| Limits | Every claim states at least one limitation; proprietary claims state an explicit boundary. |
| Résumé | Project bullets cite public evidence at TESTED or above; employment statements are a separate, labelled bullet type. |
| Review mode | /review shows only public evidence at TESTED or above. |
| Wording | Reference-architecture case studies may not use “deployed” wording. |
Each rule has a negative test that feeds the audit invalid evidence and expects a rejection. The schemas are published under /schemas.
GitHub is the source of truth
How evidence reaches this page
This build renders snapshot fcebb7139bed from 1 pinned source release. The same snapshot is served as JSON at /api/v1/.
Production
How this site runs
The reasoning behind each choice is recorded in the decision records, including why the Go API exists at all and why pages never depend on it.