Sub-agents are a scalpel, not a speed knob. Here’s what we learned making them work, and the ways they quietly fail.
A “worker” (or sub-agent) is a scoped assistant your main agent, the driver , dispatches to do a self-contained errand: scout a module, run a test suite, make a bounded edit, summarize a directory. It reports back a result and disappears. The pattern is powerful and, done wrong, a reliable source of silent bugs and wasted compute.
We built workers into lean-coder and tested them hard. This is the distilled guidance: when to use one, how to scope it, and the two failure modes that will bite you if you don’t design around them.
What a worker is for
A worker earns its keep in exactly one situation: a self-contained sub-task whose detailed work would otherwise bloat the driver’s context.
Good uses:
- Scouting. “Find where request retries are handled and report the file, the
function, and how it classifies errors.” The worker reads ten files; the driver gets three sentences.
- Bounded, verifiable edits. “In
parser.py, maketokenize()handle an empty
string; return the diff.” One file in, one diff out.
- Noisy, throwaway output. “Run the full test suite and tell me which tests fail
and why.” The 4,000-line log stays in the worker; the driver gets the summary.
The unifying trait: the work is large but the result is small, and the result is checkable. You’re spending a whole separate context window to keep the driver’s window clean.
What a worker is not for
- Not for parallel speed on one machine. (More on this below, it’s the most
common misconception.)
- Not for open-ended work. “Refactor the auth system” is not a worker task; it’s
a driver task. A worker has no memory of the broader plan and can’t make the judgment calls that span the codebase.
- Not for anything you won’t verify. A worker’s report is a claim, not a fact.
If you can’t (or won’t) check it, don’t delegate it.
Failure mode #1: the worker that lies without lying
This is the one that will hurt you, so it gets the most space.
We dispatched a worker to fix a bug. It made its edit, then reported: “fixed.” The file was unchanged.
Here’s exactly what happened, because the mechanism is general. The worker’s first edit used an ambiguous search pattern that matched two functions; our safety guard correctly refused it. The worker recovered, re-read the file, produced a correct second edit, but it bundled that edit tool-call in the same message as its “I’m done, here’s the result” block. The harness took the result as final, and the second edit was never actually executed. The worker sincerely believed it had fixed the file. It hadn’t.
Nothing errored. No exception, no red text. Good reasoning, correct edit, honest intent, and a wrong outcome, reported as success. This is the agent equivalent of “yeah, it’s pushed” while the change sits unsaved in the editor. For a language model, deciding to make an edit and having the edit take effect feel like the same act; the gap between them is invisible from the inside.
Three defenses, in order of importance:
- The driver must verify. This is the real rule. A worker’s success claim is a claim. The driver owns the plan and the ground truth, so the driver re-reads the file, re-runs the test, checks the exit code, before it believes “done.” Self-certification is the thing you design out.
- Separate action from completion. The worker protocol must forbid emitting a result block in the same turn as a tool call. Finish the tool round, observe the effect, then, on a later turn, write the result. Deciding and doing are different turns, always.
- Re-read before reporting. Tell the worker to confirm the change landed (re-read, re-run) before it declares success. Cheap insurance from the worker’s own side.
The principle generalizes to any system where one agent consumes another’s self-report: build it assuming the report can be sincerely wrong. Make success observable, and lodge the authority to declare “done” with the party that can see the ground truth.
Failure mode #2: fan-out that isn’t
The seductive idea: “I have three independent edits, spin up three workers and do them at once, three times faster.” On one machine, this is usually slower.
We measured two workers against the same task on one host: 63 seconds concurrent vs 37 seconds sequential. Running them in parallel was 70% slower.
The reason: local inference serializes on a single backend. Two workers hitting the same model don’t run in parallel, they queue and contend for the same compute while they wait. You pay the coordination overhead of concurrency and get none of the benefit, because the underlying resource (one GPU) was never parallel. Two agents “at the same time” on one device is an illusion; they’re taking turns anyway, just less efficiently than if you’d told them to.
So the concurrency limit in a worker system is mostly about isolation, keeping sub-tasks in separate contexts so they don’t pollute each other, not throughput.
Fan-out only pays off when there’s real independent capacity:
- Across different hosts or accelerators (driver on one box, workers on
separate machines).
- Across genuinely separate model instances.
- For I/O-bound work, a worker mostly waiting on the network, disk, or a slow
external command, where the waiting overlaps.
For compute-bound edits on a single box: run them sequentially and don’t kid yourself.
Choosing the worker’s model: the executor mindset
Workers don’t need to be as capable as the driver. In fact, the ideal worker model is often the opposite of the ideal driver:
- The driver needs judgment, planning, and reliable multi-turn behavior, it
holds the plan and makes the calls.
- A worker needs to execute a bounded task competently and get out. A fast,
non-reasoning code model is often a better executor than a big reasoning model: quicker, cheaper, and less prone to over-thinking a simple edit.
Concretely, we found a mid-size non-thinking coder model did a bounded edit in ~26 seconds, while a small reasoning model took 211 seconds on the same task, drowning in its own chain-of-thought. For a scoped executor task, “thinks less” is a feature.
This is also why such a model is an executor, not a self-certifying agent (see failure mode #1). It does the mechanical work well; the authority to declare success stays with the driver.
Topology: brain and tools are separate axes
One subtlety worth internalizing: a worker has two independent “wheres.”
- Where its tools run (which filesystem
read_file/apply_diff/run_command
act on).
- Where its brain runs (which machine does the inference).
These don’t have to be the same box. You can run a worker’s inference on a fast shared server while its edits land on the local machine’s disk, or vice versa. That decoupling is what makes sensible topologies possible: put the heavy inference where the compute is, put the file operations where the code is, and stop assuming they’re the same place.
The practical upshot: on a big-memory machine, a driver plus a couple of resident workers coexist happily. On a small-VRAM GPU box, don’t try to hold two models at once, use it as the driver and dispatch workers to a beefier host, or run them sequentially.
Permissions: a worker is capped by its parent
A safety note that should be a hard rule in any worker system: a worker’s authority is never above its dispatcher’s. If the driver is read-only, a worker it spawns is read-only. If the driver can edit but not run commands, neither can its worker. Capability flows down and only narrows. A sub-agent should never be a privilege-escalation path, the whole point of a leash is that it’s inherited and capped, not a fresh grant.
The checklist
Before you dispatch a worker, ask:
- Is the task self-contained? Could a stranger do it with only the instruction you’re about to write? If it needs the broader plan, it’s a driver task.
- Is the result small and checkable? Big work, small verifiable output = good worker. Otherwise reconsider.
- Will the driver verify the result? If not, don’t delegate it.
- Am I reaching for a worker for speed? On one machine, that’s usually a mistake, do it inline or sequentially.
- Is the model right for an executor? Fast and literal beats big and ruminative for bounded tasks.
- Is the leash correct and capped? The worker inherits the driver’s authority, never more.
Workers are a scalpel for keeping the driver’s context clean and its judgment focused. Used that way, they’re one of the highest-leverage tools in an agent. Used as a generic “make it faster / do more at once” button, they’ll cost you compute and, worse, hand you confident reports of work that never happened.
This is a field note from lean-coder, an open-source, dependency-free terminal coding agent that treats context as the scarce resource it is. Source and docs: github.com/codemonkeying/lean-coder. Licensed MIT.


Leave a Reply