AI infrastructure
Choosing a local model with the agent's own tasks
Implementedenvironment:LabPersonal labSelf-reported · private sourcemyagent-local-model-eval
Claim
Benchmarked two local models (Qwen3 4B and Qwen3.6 35B-A3B MoE) on an 8 GB laptop GPU against a hand-labelled evaluation set, using the agent's production prompts, schemas and validators, and chose the model per task from the results.
Status
- implemented
- Capped at Implemented: the code, tests and records are private, so this is self-reported and cannot be verified publicly.
- lab
- Local machine, local clusters or CI runners.
- personal lab
- My own projects. Public repositories carry CI-verified evidence; private ones are self-reported.
Source
A personal project whose repository is private and cannot be published. So this claim is self-reported: it stops at Implemented and links to no artifacts.
Limitations
- The 35B results come from the first 30 items of each set and the 4B results from the full sets, so the comparison is not like for like.
- One run per model; the evaluation set is small and fictional.
- Raw reports: local only.