Skip to content
Ali Akbari
Menu

AI infrastructure

Choosing a local model with the agent's own tasks

Implementedenvironment:LabPersonal labSelf-reported · private sourcemyagent-local-model-eval

Claim

Benchmarked two local models (Qwen3 4B and Qwen3.6 35B-A3B MoE) on an 8 GB laptop GPU against a hand-labelled evaluation set, using the agent's production prompts, schemas and validators, and chose the model per task from the results.

Status

implemented
Capped at Implemented: the code, tests and records are private, so this is self-reported and cannot be verified publicly.
lab
Local machine, local clusters or CI runners.
personal lab
My own projects. Public repositories carry CI-verified evidence; private ones are self-reported.

Source

A personal project whose repository is private and cannot be published. So this claim is self-reported: it stops at Implemented and links to no artifacts.

Limitations

  • The 35B results come from the first 30 items of each set and the 4B results from the full sets, so the comparison is not like for like.
  • One run per model; the evaluation set is small and fictional.
  • Raw reports: local only.

Appears in