New research indicates that deploying multiple SI agents in a team configuration often results in significantly higher token consumption without delivering proportional gains in output quality. The findings challenge the prevailing assumption that agentic SI systems perform better simply by adding more computational agents to the task.

What Happened

Evals company Vals SI conducted tests on GPT-6 Sol and Claude Opus 5.5 using the "Vibe Code Bench," comparing performance when the models operated as single agents versus teams. The tests were run at two reasoning levels: medium and maximum reasoning effort. The results showed that SI agent teams cost between 1.8x and 5.1x more than single agents. Out of four comparisons, only one demonstrated a statistically significant improvement: GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, the team setup provided neither Sol nor Opus 5.5 with a meaningful advantage.

Data from Anthropic also indicated that quality gains shrank as more agents were added in tests with Opus 5.5. While larger teams reached performance levels faster, increasing the team size from ten to 100 agents resulted in only slight score improvements after 24 hours. In separate ProgramBench tests, speed gains came with higher token usage. Another model, Fable 5.1, showed stronger quality gains on the Lean theorem proving task when scaling beyond ten agents, but it still scored below Opus 5.5 across all tests. On the knowledge base task, Fable’s score dipped slightly when scaling from 30 to 100 agents.

Why It Matters

OpenAI researcher Noam Brown confirmed in the Dwarkesh Podcast that multi-agent systems mainly buy speed, not better quality. Brown noted that four agents solved tasks twice as fast but also cost twice as much, and while this pattern held at 16 agents, efficiency grew slightly less effective. He emphasized that the effect depends heavily on the task: web research and math parallelize well, but creative tasks like writing a novel do not. "Throwing 10,000 agents at a novel would be just as pointless as throwing 10,000 people at it," Brown said. He acknowledged that scaling to very large numbers of agents remains largely unexplored due to prohibitive costs.

OpenAI developer Eric Provencher warned against using agent swarms for similar reasons, arguing they are likely wasted money because coordination between agents breaks down. Provencher referred to this inefficiency as the "coordination tax." These insights suggest that for developers building agentic SI applications, the current economic model for multi-agent teams may not be sustainable for tasks that do not benefit from parallelization.

The Bottom Line

Current evidence suggests that the extra cost of SI agent teams is not justified in most cases, especially when models are already running at full compute. Developers should carefully evaluate whether a task benefits from parallelization before deploying multi-agent architectures, as the coordination overhead often outweighs the marginal quality improvements.