In this guide
Subagent
A subagent is an agent that another agent calls to handle a bounded piece of work, instead of doing that work inline itself. It runs with its own context window — the pool of conversation and information it can see while working — its own instructions, and often its own restricted set of tools. When it finishes, it hands back a result, not its full transcript. Everything it did along the way disappears from the parent's context.
The word doesn't point to different technology — it points to a relationship. The same agent is a top-level agent when a person talks to it directly, and a subagent when another agent invokes it as part of a larger task. Either way, it was called by another agent, for a piece of that agent's own job, by a system that still controls and spawns it.
Handing work to an agent you don't control the internals of at all, like a different company's own system, is a different problem with its own protocol: A2A.
Why subagents exist
- Context isolation. A subtask that takes twenty tool calls and a lot of back-and-forth to finish would otherwise fill the parent's context window with all of that detail. Delegating it means the parent's context grows only by the size of the final result, not the work it took to get there. This is what lets one top-level agent handle a larger overall task before hitting context limits.
- Specialization. A subagent can get a narrower system prompt and a restricted tool list than the parent. A code-review subagent might get read-only file access and nothing else — more reliable at that job, and safer than a do-everything agent with every tool enabled.
How Claude Code does it
Claude Code's own subagent feature works exactly this way: a project defines subagents as Markdown files with YAML frontmatter under .claude/agents/ (or ~/.claude/agents/ for ones available across every project), each with its own system prompt and its own list of allowed tools. When the main agent needs to search a large codebase for how something is implemented, it can delegate that search to an Explore subagent — which might make dozens of grep and file-read calls — and get back a short, synthesized answer instead of dozens of raw search results cluttering the main conversation.
Parallel subagents
Because a subagent's work is isolated, independent subtasks can run as multiple subagents at once instead of one after another. If a task splits into unrelated pieces — research three topics, run three independent code reviews — dispatching them as parallel subagents and combining the results is faster than doing them serially. No coordination is needed until the results come back.
When to use a subagent
- The subtask's own reasoning or tool output would be large enough to meaningfully crowd the parent's context.
- The subtask benefits from a different persona or a more restricted tool set than the parent has.
- The task splits into genuinely independent pieces that can run in parallel.
When not to use one
- Trivial lookups. Spinning up a subagent has real latency and token overhead. A question a single tool call can answer doesn't need a subagent wrapped around it.
- When the parent needs step-by-step visibility. A subagent hides its intermediate steps by design — if the parent (or the person watching) needs to see and react to each piece of progress as it happens, delegating hides exactly the information that's needed.
- Tightly coupled, evolving state. If the task requires constant back-and-forth against shared state that's awkward to hand off in one message and collect back in one result, a subagent's one-shot delegate/return shape doesn't fit well.
The real tradeoff
Context isolation cuts both ways. It's why subagents scale, and why they can hide mistakes: the parent only sees a subagent's final summary, not the reasoning behind it. An error a human could have caught mid-process can reach the parent already baked into a confident-sounding result. Delegating is a bet that the subagent's summary is accurate — it's not free of risk just because it keeps the parent's context clean.