Vals AI: agent teams cost 1.8 to 5.1 times more, and only one of four gains was significant
Vals AI, an evaluation company, reported on 9 October that agent teams cost 1.8 to 5.1 times as much as a single agent when building 50 web apps, and that only one of four score gains was statistically significant. It tested GPT 6 Sol and Claude Opus 5.5, each alone and as a lead agent directing up to five subagents, at medium and max reasoning effort.
Why it matters: Anthropic says its dynamic workflows, released on 9 October, let one run start up to 1,000 agents at normal token rates. Vals measured what a smaller version of that pattern costs. Opus at max effort rose from $23.77 to $122 per app for a 3.4-point gain that Vals found not significant (95% interval from -1.4 to +8.3 points).
For and against: The one clear win was GPT 6 Sol at medium effort: the team scored 84.9% against 77.6% alone (p = 0.005), at $3.07 per app against $1.22. Against that, a single Sol agent at max effort scored 89.0% at $4.82, ahead of the medium-effort team. For Opus, the cheapest setup (91.5% at $4.08) scored about as well as the dearest.
The catch: Vals ran one benchmark of full-stack web apps, one run per setup, with a prompt it wrote, and says the results may not carry over to other work. Speed was mixed: Sol's team finished a max-effort app in 27.6 minutes against 38.4 alone, while Opus teams took up to 2.3 times as long.
What to do: Before you enable multi-agent orchestration, run 30 of your own tasks three ways: one agent at today's effort, one agent at higher effort, and a team. Record score and dollars per finished task for each, and set a per-run budget in your platform before the team setting goes live.
Our coverage →Do Agent Teams Pay Off? A Case Study on Vibe Code Bench (Vals AI, 9 October 2026) ↗Workflow runs (Anthropic documentation) ↗