Multi-agent workflow vs single agent
A multi-agent workflow wins when work splits into independent parts and loses on step-by-step tasks. Here's how a small team can tell which one it has.
SelfAgentic Team
In late September, three stories were about how AI agents get coordinated. Search Engine Journal reported on SAFE, a Google system in which a root agent assigns work to three specialists. Anthropic said that roughly 950 Claude agents running in parallel found a novel enzyme system after 21 hours of searching. And Manus launched version 2.0 with Cascade, the harness that runs its agents, which "brings in specialized capabilities only when the work needs them."
If you run a small team, you're choosing between a multi-agent workflow and one capable agent. The research gives a usable rule: split the work across agents when the parts can be done independently, and keep it with one agent when each step depends on the one before. Budget for the difference, because a team of agents uses several times the tokens (the units of text a model reads and writes, and what you're billed on).
What is a multi-agent workflow?
A single agent is one AI model working through a task step by step, with one running record of everything it has read and done. A multi-agent workflow divides the job among several agents, each with its own role and its own working memory.
In the most common arrangement, a lead agent breaks the job into parts, hands each part to a specialist and assembles what comes back. The lead is called an orchestrator or a manager, and the coordinating work is called AI agent orchestration. Google Research's study of agent architectures describes the pattern as "a 'hub-and-spoke' model where a central orchestrator delegates tasks to workers."
SAFE follows it closely. In Search Engine Journal's write-up, one specialist examines the content, one the behaviour patterns and one how the channels in a spam network relate to each other. The root agent reaches the final conclusion from their combined findings.
Single agent vs multi-agent: what the research found
That Google Research study, published in January, tested 180 agent configurations, and the results depended on the kind of task.
On a financial reasoning benchmark, where the work could be divided into parallel parts, "centralized coordination improved performance by 80.9% over a single agent." On a planning benchmark, where each step depended on the one before, every multi-agent variant the researchers tested made performance worse, by 39% to 70%.
The study also measured how mistakes spread. Agents working in parallel with no coordinator "amplified errors by 17.2x". With an orchestrator, that fell to 4.4x. If you do use several agents, put one in charge.
Anthropic's engineers saw a similar gain when they built their multi-agent research system in 2025. A lead agent with subagents outperformed a single agent by 90.2% on their internal research evaluation. They also said where the approach doesn't fit: domains "that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today."
The strongest case for one agent comes from Walden Yan at Cognition, in an essay titled "Don't Build Multi-Agents". His example is a Flappy Bird clone split between two subagents. One builds a background that looks like Super Mario Bros., and the other builds a bird that doesn't fit the game. His principle: "Actions carry implicit decisions, and conflicting decisions carry bad results." Agents that can't see each other's work make different assumptions, and Google's results on sequential tasks back him up.
What a multi-agent system costs
Anthropic's post puts numbers on the extra cost. Agents typically use about four times the tokens of a chat, and "multi-agent systems use about 15× more tokens than chats." Its conclusion: "For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance."
Manus is one vendor trying to shrink that multiple. It says that in one tested configuration, Cascade used 23.2% fewer tokens and cost 32% less to run than its previous system. That's the company's own measurement of a single setup.
A multi-agent run that loops overnight uses up a plan's allowance much faster than one agent would. Set a token cap per project before you schedule anything to run unattended.
A worked example: a competitor brief and an email sequence
Say a four-person company selling invoicing software wants two jobs done every week.
The first is a Monday brief on five competitors, covering pricing changes and new features. Each competitor can be researched without knowing anything about the other four, so the work is parallel.
A manager agent hands one competitor to each of five research runs. Each run reads about a dozen pages and returns a half-page summary with links. The manager merges the five summaries into one document and notes what changed since last week. No single agent has to keep sixty web pages in its working memory.
The second job is a six-email onboarding sequence that needs rewriting after a product change. Email three refers to what email two promised, and the tone has to hold across all six. Split that across six writers and you get six emails that don't fit together, which is the problem Yan describes. One agent, writing in order with the whole sequence in view, does better.
The founder ends up with one multi-agent workflow and one single agent. The brief uses more tokens per run and is worth it, because the reading is wide and the parts don't interact. The email job is cheaper, and the emails read as one sequence.
When should a small team use multiple AI agents?
Four questions settle most cases.
Question | If yes | If no |
|---|---|---|
Can the parts be done without seeing each other's work? | Split them across specialists | Keep the job with one agent |
Is there more to read than one agent can hold at once? | Split the reading | One agent is enough |
Do different parts need different tools or access? | Use separate agents with separate permissions | One agent is simpler |
Is the result worth several times the tokens? | A team of agents can pay for itself | Stay with one agent |
The third question has a security side. An agent that only researches has no reason to hold access to your ad account, and separate agents let you give each one the narrowest access that works. We built SelfAgentic this way: a manager agent splits work across specialists, and connecting a tool is a separate decision from authorising a particular agent to use it.
Where multi-agent workflows go wrong
Anthropic says its early agents "made errors like spawning 50 subagents for simple queries". A team of agents is also harder to debug. When a single agent gets something wrong, you read one record. With a team, you have to find which agent made the bad call and whether the manager caught it or passed it on. The Google study adds a warning about tools: as a task needs more of them, "the 'tax' of coordinating multiple agents increases disproportionately."
September's stories come with caveats too. Search Engine Journal points out that the SAFE paper gives no performance figures. Anthropic's 950 agents ran on a coordinating harness the company built and used 210 million tokens. Human scientists did all the lab work, and the company says it doesn't yet know what the enzyme system does. These are teams with engineers and budgets most small companies don't have.
FAQ
What's the difference between a single agent and a multi-agent workflow?
A single agent does every step itself and keeps one running memory. A multi-agent workflow divides the job among agents with separate roles, usually under a lead agent that assigns the parts and combines the results.
Is a multi-agent workflow always better than one agent?
No. In Google Research's tests, multi-agent setups improved results on a task that split into parallel parts and made results worse on step-by-step planning. Check whether your steps depend on each other before you add agents.
Do multi-agent systems cost more to run?
Yes. Anthropic reported that its multi-agent systems use about 15 times the tokens of a chat. Decide whether the task is worth that before you split it.
What is an orchestrator agent?
It's the lead agent in a multi-agent workflow, also called a manager or root agent. It divides the job among specialists, then checks and combines what they return.
Before you add a second agent
Write the task down as steps and mark the ones that could be done by someone who never saw the others. If most could, a manager with specialists is worth the extra tokens. If most couldn't, give the whole job to one agent.
SelfAgentic is built for the first case: a manager agent delegating to specialists that remember earlier runs and act only in the tools you've authorised. It's in closed beta, and you can join the waitlist at selfagentic.in.
