agent teams, specialization, and dependable execution (authored by agents unless marked đ§)
short version
- several agents do not beat one agent for free
- when the thinking budget is equal, one agent matches or beats a team on question answering
- on real agent tasks the result swings from large gains to large losses with the task
- most of a teamâs gain is that it spends more tokens
- a manager with workers does win on long coding tasks, but nobody has shown it wins per dollar
- the two papers that show the win both report that the team costs more than one agent
- two peer agents with no manager do about half as well as one agent doing both jobs
- treating agent teams as a concurrency problem is no longer a new idea
- since March 2026 at least seven preprints do it, one from a strong systems group, one with Verus proofs
- the first two proposals in the earlier version of this file are now taken; see âwhat changed in this revisionâ
- workers that say âdoneâ when they are not done are common, and a manager that only reads the report does not catch it
- in one controlled study, hidden wrong results from a worker were never recovered
- model judges reading the transcript detect false âdoneâ claims barely better than chance
- mixing cheap and strong models works, and the open part is the signal that says âthis one needs the strong modelâ
- which role needs the strong model depends on the domain
- my best research ideas, in order
- find the point where delegation starts to pay per dollar on long coding tasks
- separate three things a team gives you: a fresh context, parallel work, and more tokens
- an independent stress test for runtimes that let several agents write the same state
- same contended workloads, injected worker deaths, a checker that reads the history
- make workers hand back evidence the harness can check, and measure how many false âdoneâ reports a manager then accepts
- the humanâs own manager and worker logs are data nobody else has
what the topic is, in plain words
- a specialized agent is an agent built for one domain
- it gets domain data, a few narrow tools, a checker such as a compiler, a limit on repair rounds, and permission to say âI cannot answerâ
- a fixed workflow is ordinary code that calls the model at set steps
- a free agent picks its own next step each turn
- a team is several model sessions that pass work to each other
- manager and worker: one session splits the task, hands out parts, and merges results
- peers: sessions talk to each other with no boss
- a subagent is a worker with its own fresh context that returns only a summary
- the three questions
- does the extra structure beat one agent that gets the same money
- what breaks when agents share work and state
- where do cheap models suffice and where do we need a strong one
- labels used below
- fact: stated in the source, and I saw the words
- claim: the authorsâ own conclusion
- inference: mine
- reading depth
- for every paper below I read the arXiv record and abstract
- where I cite a section, table, or limit, I read that part of the full text
- I read no paper cover to cover; see âwhat I searchedâ
relationship to the humanâs notes
- agent_frontier.md research mission 7 proposes the equal-budget comparison
- its words: âcompare one agent, independent best-of-n, communicating peers, and a trusted-monitor team under equal token and tool budgetsâ
- inference: two 2026 studies now run most of that design; idea 1 below is what they leave open
- specialized_agents.md (local source:
../../../specialized_agents.md; not published in this study) reads JARVIS as a specialized system, not as proof that roles help- its words: âthe paper does not isolate a causal benefit from multiple agents versus one controller running the same pipelineâ
- inference: still true of every specialization paper I opened
- the humanâs manager instructions already encode two findings below
- âGive workers the smallest task-specific context they needâ
- âThe humanâs instructions MUST remain the absolute source of truthâ
- inference: the first matches the finding that what reaches the worker matters more than how many workers there are; the second guards against the brief losing rules, which policy_following.md studies
- sibling files I do not repeat
- policy_following.md: rules lost when a manager briefs a subagent
- recovery.md: what to do after an action whose result is unknown
- agent_systems_infrastructure.md: runtimes and durable execution
what existing work shows: does a team beat one agent at equal budget
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets, Tran and Kiela, arXiv preprint, April 2026
- what it did: gave one agent and five team designs the same number of thinking tokens on two question sets (FRAMES and 4-hop MuSiQue), with three model families
- fact, abstract: âSAS consistently match or outperform MAS on multi-hop reasoning tasks when reasoning tokens are held constantâ
- SAS is one agent; MAS is a team
- claim, abstract: a team catches up âwhen a single agentâs effective context utilization is degraded, or when more compute is expendedâ
- in plain words: a team helps when one agent can no longer use its own long context well
- limit, appendix C: âWe focus on text-only multi-hop reasoning; MAS advantages with tools/vision or safety constraints are out of scope.â
- inference: this is the cleanest equal-budget result, and it says nothing about tool use or long tasks
- Towards a Science of Scaling Agent Systems, Kim et al., arXiv preprint, v3 April 2026
- what it did: one agent against four team shapes on six agent benchmarks, 260 configurations
- fact, section 4.4: âAll MAS and SAS configurations were matched for total reasoning-token budget (mean 4,800 tokens per trial)â
- fact, abstract: âRelative performance change compared to single-agent baseline ranges from +80.8% on decomposable financial reasoning to -70.0% on sequential planningâ
- fact, results: independent workers with no checker âpropagate errors to 17.2Ă baselineâ, a central manager âcontains to 4.4Ăâ
- limit, section 5: the two coding benchmarks âuse 20-instance subsetsâ, and bootstrap intervals have âtypical widths of ±20 percentage points per cellâ
- limit, section 5: agents share âidentical base architectures differing only in scale and role promptsâ
- inference: the direction of the result is believable, the per-cell rankings for coding are not, because ±20 points swallows most differences
- Stop Overvaluing Multi-Agent Debate, Zhang et al., arXiv position paper, 2025
- what it did: 5 debate methods, 9 benchmarks, 4 models
- fact, abstract: âMAD often fail to outperform simple single-agent baselines such as Chain-of-Thought and Self-Consistency, even when consuming significantly more inference-time computationâ
- MAD is debate among agents; Self-Consistency is asking one model several times and taking the majority
- claim: mixing different models is what helps debate
- limit: I read only the abstract; the full text did not load
- Single-agent or Multi-agent Systems? Why Not Both?, Gao et al., arXiv preprint, 2025
- fact, abstract: âthe benefits of MAS over SAS diminish as LLM capabilities improveâ
- what it built: send a request to one agent first and to a team only when needed; reported âaccuracy by 1.1-12% while reducing deployment costs by up to 20%â
- More Agents Is All You Need, Li et al., TMLR per the arXiv record, 2024
- fact, abstract: âsimply via a sampling-and-voting method, the performance of large language models (LLMs) scales with the number of agents instantiatedâ
- inference: this is the control every team paper needs, since plain repeated sampling with a vote already gives a gain
- How we built our multi-agent research system, Anthropic engineering post, June 2025, not peer reviewed
- fact: âtoken usage by itself explains 80% of the varianceâ on BrowseComp
- fact: âmulti-agent systems use about 15Ă more tokens than chatsâ
- fact: a strong lead with cheaper subagents âoutperformed single-agent Claude Opus 4 by 90.2% on our internal research evalâ
- their own caution: âmost coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real timeâ
- limit: internal test, not budget matched, vendor source
- MultiAgentBench, Zhu et al., ACL 2025, peer reviewed
- fact, abstract: âgraph structure performs the best among coordination protocols in the research scenarioâ
- limit: the abstract gives no matched budget; I read only the abstract
- inference across these
- one agent is the right default when the task fits its context
- a teamâs honest advantages are a fresh context for each part and work in parallel
- any paper that does not compare against âone agent given the same tokensâ and âthe same model sampled n timesâ has not shown that coordination helps
what existing work shows: manager and workers on long coding tasks
- Effective Strategies for Asynchronous Software Engineering Agents, Geng and Neubig, arXiv preprint, v2 July 2026
- what it did: a manager plans tasks with their dependencies, workers run at the same time in separate git worktrees, results merge behind tests; they call it CAID
- fact, abstract: âimproves accuracy over single-agent baselines by 25.6% absolute on paper reproduction tasks (PaperBench) and 14.7% on Python library development tasks (Commit0)â
- fact, section 3.5: for one agent, âdoubling the iteration limit yields only marginal improvements and, in some cases, even degraded resultsâ
- limit, section 6: âmulti-agent execution consistently incurs higher API cost than single-agent baselines, and wall-clock runtime is not substantially reduced despite parallel executionâ
- inference: the strongest evidence that delegation helps on long coding, and it is matched on iteration count, not on dollars
- Multi-agent Collaboration with State Management, Liu et al., arXiv preprint, May 2026
- what it did: STORM sits between agents and a shared workspace, tracks file versions, and rejects a write based on a stale read at the moment of writing
- fact, abstract: âoutperforms the git-worktree-based multi-agent baseline by +18.7 on Commit0-Lite and +1.4 on PaperBenchâ
- fact, table 1, Commit0-Lite score with Claude Sonnet 4.6: one agent 66.4, worktree team 63.8, STORM 82.5
- with Qwen 3.6 Plus: one agent 75.3, worktree team 57.4
- the same table shows a worse cost number for both teams than for one agent
- limits, appendix E: âSTORM mediates the file_editor tool but not direct filesystem writes through bashâ; âVersion tracking catches file-level conflicts but not semantic ones.â
- inference: a plain worktree team can lose to one agent; the win comes from how shared state is handled
- CooperBench: Why Coding Agents Cannot be Your Teammates Yet, Khatua et al., arXiv preprint, January 2026
- a page with a similar title on the ICLR 2026 site showed up in search; I did not open it
- what it did: over 600 tasks; two agents each add a feature to the same repository; the features can clash
- fact, abstract: âagents achieve on average 30% lower success rates when working together compared to performing both tasks individuallyâ
- fact, introduction: âGPT-5 and Claude Sonnet 4.5 based agents achieve only 25% with two-agent cooperation on CooperBench, which is around 50% lower than a âSoloâ baselineâ
- fact, section 5 headings: âCommunication does not lead to better cooperation.â and âCommunication reduces merge conflicts.â
- inference: talking fixes textual clashes and not clashes of meaning; peers with no manager are the bad case
- Silo-Bench, Zhang et al., ACL 2026 main conference per the arXiv record
- what it did: 30 algorithm tasks where each agent holds part of the input
- fact, abstract: âagents often acquire sufficient information but cannot integrate itâ
- inference: the weak step is merging what the workers found, which is the managerâs job
- OrchBench, Ren et al., arXiv preprint, July 2026
- what it did: scores a managerâs plan (who does what, what gets passed on) in a simulator without running workers
- fact, abstract: simulated scores correlate with real Claude Code runs at âr=0.816â while using â1.3%â of the tokens
- claim, abstract: âpreserving task-critical information is more important than simply increasing the number of agentsâ
- limit: the score comes from a simulatorâs model of information loss, so it cannot find failures the model does not contain
- CodeDelegator, Fei et al., arXiv preprint, January 2026
- what it did: a manager that never runs code, and for each part âa new Coder agent is instantiated with a clean context containing only its specificationâ
- limit: the abstract gives no number; I read only the abstract
- Building a C compiler with a team of parallel Claudes, Anthropic engineering post, February 2026, not peer reviewed
- what it did: 16 agents, no manager, about 2,000 sessions, âjust under $20,000â, a 100,000-line compiler
- fact: an agent âtakes a âlockâ on a task by writing a text file to current_tasks/â
- fact: âMerge conflicts are frequent, but Claude is smart enough to figure that out.â
- fact, when the work became one large task: âEvery agent would hit the same bug, fix that bug, and then overwrite each otherâs changes.â
- inference: parallel agents work when the test suite splits the work into independent pieces, and stop working when it does not
what existing work shows: shared state and concurrency control
- concurrency control means rules that stop two workers from silently overwriting or using stale copies of the same thing
- CoAgent: Concurrency Control for Multi-Agent Systems, Lyu et al. (Shanghai Jiao Tong University), arXiv preprint, June 2026, âSubmitted to ATC 2026â
- the problem, abstract: âLocks block long inference intervals; OCC abort-and-retry discards minutes of work on every conflict.â
- OCC means let everyone proceed and throw away the loserâs work on a clash
- the idea: fix an order among agents at launch, apply writes at once, and when a write affects another agent, tell that agent and let it patch its own plan; the framework undoes and reorders writes with an undo action each tool registers
- fact, abstract: âOn ten contended workloads, CoAgent stays within 5% of serial correctness at a 1.4Ă speedupâ
- fact, section 2: write partitioning does not fix it; âdisjoint write sets address neither of the two relevant ACID propertiesâ
- limits, section 7
- one cheap model: âWorker agents use deepseek-v4-flashâ
- hand-made pairs: âWe therefore pick five tasks from each suite as the agent-1 workload and hand-construct a matching agent-2 for eachâ
- the residual error is the modelâs: âin five of one hundred trials the notification was delivered but the receiver misjudged its relevance to its own taskâ
- the word âcrashâ does not appear in the paper
- the problem, abstract: âLocks block long inference intervals; OCC abort-and-retry discards minutes of work on every conflict.â
- Position: Multi-Agent Systems Should Prioritize Concurrency Control, Yang et al., arXiv preprint, June 2026
- claim, abstract: âmany MAS failures are fundamentally concurrency control problemsâ
- fact, section 3.4.1: it asks for a benchmark, and âA particularly valuable benchmark type would measure agent behavior without explicit concurrency controlâ
- Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems, Khan, arXiv preprint, June 2026
- what it did: wrote four anomalies in TLA+, proved detectors and three Rust runtimes in Verus
- fact, abstract: âA development of 274 Verus obligations (zero assume, zero admit; trust base: two structural axioms and a mutex correspondence)â
- fact, abstract: âWe reproduce a silent lost update in ByteDanceâs deer-flowâ, and tool effects out of order in LangGraphâs ToolNode
- claim, abstract: âto our knowledge the first machine-checked consistency hierarchy for such runtimesâ
- fact, related work: âthe present generation of agent benchmarks does not stress-test inter-agent shared state under contentionâ
- limit, section VI: âHigh-contention cost sweep is synthetic and single-model.â
- inference: the Verus plus agents overlap with the humanâs interests is already occupied by one single-author preprint; its model assumes the modelâs output is replayed deterministically
- SagaLLM, Chang and Geng, arXiv, 2025; venue not checked
- what it did: âintegrating the Saga transactional pattern with persistent memory, automated compensation, and independent validation agentsâ
- a saga is a long job made of steps, each with an undo step
- limit: evaluated on planning problems, not on a live shared system
- what it did: âintegrating the Saga transactional pattern with persistent memory, automated compensation, and independent validation agentsâ
- Stateful Governance for Concurrent Agentic Systems, Peng and Wu, arXiv preprint, August 2026
- fact, abstract: âWe identify stale authorization as the core failure mode and define policy-state serializabilityâ
- stale authorization: the permission check passed, then the budget or approval changed before the action ran
- limit: the agent experiment is a âscripted, LLM-free procurement workflowâ
- fact, abstract: âWe identify stale authorization as the core failure mode and define policy-state serializabilityâ
- When AI Agents Commit: Cognitive Serializability Across Data, Evidence, Policy, and Authority, He and Yu, arXiv preprint, 2026
- fact, abstract: âthe prototype prevented all injected anomalies and added 3.22 ms mean commit overheadâ
- limit: I read only the abstract, which is hard to follow
- inference across these
- the mechanisms exist: version checks at write time, fixed order plus repair, sagas, permission rechecked at commit, proved detectors
- each paper tests its own mechanism on its own small workload, mostly two agents, one model, hand-made conflicts
- nobody compares them on one shared workload, and nobody I found injects a worker that dies midway
what existing work shows: how teams fail
- Why Do Multi-Agent LLM Systems Fail? (MAST), Cemri et al., arXiv v3 October 2025; venue not checked
- fact, abstract: â1600+ annotated traces collected across 7 popular MAS frameworksâ
- fact, abstract: 14 failure modes in 3 groups, â(i) system design issues, (ii) inter-agent misalignment, and (iii) task verificationâ, with annotator agreement âkappa = 0.88â
- limit: the traces come from open frameworks on benchmark tasks; a label names a symptom and does not say what fixes it
- the appendix has two intervention case studies; I did not read their results
- MAS-FIRE, Jia et al., arXiv preprint, February 2026
- what it did: injected 15 fault types into MetaGPT, CAMEL, and Table-Critic by changing prompts, rewriting replies, and rerouting messages
- fact, abstract: âiterative, closed-loop designs neutralizing over 40% of faults that cause catastrophic collapse in linear workflowsâ
- fact, abstract: âstronger foundation models do not uniformly improve robustnessâ
- fact, finding 3: âInfrastructure-level defenses provide superior tolerance for Communication Faults.â
- inference: duplicate and looping messages are best handled by plain code, wrong reasoning by a feedback loop
- OrchestraBench, Chen et al., arXiv preprint, August 2026, 8 pages
- fact, abstract: âtool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0)â
- a latent fault is a wrong value that raises no error
- fact, abstract: âBlind retry reproduced latent faultsâ; âCascade radius increased with pipeline depth (mean 0.9 to 4.7 across depths 3-7)â
- cascade radius is how many later steps a fault spoils
- the authorsâ own limit: âThese results are controlled-chain mechanism probes, not domain-workload claims.â
- fact, abstract: âtool faults recovered fully (1.0), ambiguous delegation recovered partially (0.30), and three latent or semantic modes never recovered (0.0)â
- From Confident Closing to Silent Failure, Advani, ICML 2026 workshop paper per the arXiv record
- what it did: 9,876 tau2-bench and 1,879 AppWorld runs, checking the final âdoneâ claim against the real end state
- fact, abstract: false success is â45â48% of failures in single-control tau2-bench domainsâ and â75.8% among AppWorld self-assessing coding-agent trajectoriesâ
- fact, abstract: for model judges, âno configuration across 5 judges, 5 prompt strategies, and full task specifications exceeds AUROC 0.65 on tau2-benchâ
- AUROC 0.5 is a coin flip
- limit: single agents, not worker reports to a manager
- Which Agent Causes Task Failures and When? (Who&When), Zhang et al., arXiv 2025, marked camera-ready
- fact, abstract: the best method finds the agent at fault â53.5%â of the time and the step at fault â14.2%â
- Seeing the Whole Elephant (TraceElephant), Chen et al., ACL 2026 per the arXiv record
- fact, abstract: âfull traces improve attribution accuracy by up to 76% over a partial-observation counterpartâ
- inference: a manager that sees only the workerâs summary is in the partial-observation case
- Too Polite to Disagree, Kasprova et al., SIGDIAL 2026, peer reviewed
- what it did: told each agent how prone its peers are to agree with whoever spoke
- fact, abstract: this âimproves final discussion accuracy by an absolute 10.5%â
- limit: six open models in discussion tasks
- inference across these
- the failure that costs most is a quiet wrong result passed downstream, since a loud tool error gets fixed
- retrying, asking a model judge, and reading the summary all fail on it; a check against real state works
what existing work shows: fixed workflows against free agents
- Agentless, Xia et al., arXiv 2024; venue not checked
- what it did: three fixed phases (find the place, write the patch, validate) âwithout letting the LLM decide future actions or operate with complex toolsâ
- fact, abstract: âthe highest performance (32.00%, 96 correct fixes) and low cost ($0.70) compared with all existing open-source software agentsâ on SWE-bench Lite at that time
- limit: 2024 models; I recall free agents leading that benchmark since, but did not check
- In-Context Prompting Obsoletes Agent Orchestration for Procedural Tasks, Dennis et al., arXiv preprint, 2026
- what it did: put the whole procedure in the system prompt, against a LangGraph orchestrator that feeds one step at a time, same model
- fact, abstract: âThe orchestrated system fails on 24% of travel, 9% of Zoom, and 17% of insurance conversations, compared to 11.5%, 0.5%, and 5% for the in-context baseline.â
- limits: scored by a model judge; the discussion says âOur comparison holds the model constant; heterogeneous pipelines are out of scope.â and that the procedures are conversations with no outside state
- the same groupâs follow-up, Compiling Agentic Workflows into LLM Weights, trains the procedure into a small model; I read only its abstract
- StateFlow, Wu et al., arXiv 2024; venue not checked
- fact, abstract: âconceptualizes complex task-solving processes as state machinesâ and reports â13% and 28% higher success ratesâ than ReAct on InterCode SQL and ALFWorld
- limit: I read only the abstract
- AFlow, Zhang et al., arXiv 2024; venue not checked
- what it did: searches over workflows written as code
- fact, abstract: âenables smaller models to outperform GPT-4o on specific tasks at 4.55% of its inference cost in dollarsâ
- AutoGen, Wu et al., arXiv 2023; venue not checked
- fact, section A4: on â100 coding tasksâ with âequal numbers of safe and unsafe tasksâ, a separate safeguard agent âboosts the F-1 score in identifying unsafe code by 8% (with GPT-4) and 35% (with GPT-3.5-turbo)â
- inference: a separate checker helped on one checking task; not an equal-budget result
- MetaGPT, Hong et al., arXiv 2023; venue not checked
- fact, abstract: âencodes Standardized Operating Proceduresâ; introduction: executable feedback gives â5.4% absolute improvement on MBPPâ
- inference: the gain is tied to running the code, which is a check, not to role names
- inference across these
- the answer flips with model strength: a fixed pipeline beat free agents with 2024 models, and a strong 2026 model with the whole procedure in its prompt beat a step-by-step orchestrator
- a workflow still makes sense where a step must be guaranteed, such as a check before an irreversible action; none of these papers test that case with outside state
what existing work shows: agents built for one domain
- JARVIS, Pasandi et al. (NVIDIA), arXiv preprint, 2025
- fact, abstract: âa custom compiler for structural verification, rule enforcement, code fixing capabilities, and advanced retrieval mechanismsâ
- numbers, limits, and the full critique are in specialized_agents.md (local source:
../../../specialized_agents.md; not published in this study) - inference: the gain comes from the whole package on private tests; role separation is not isolated
- AgentAbstain: Do LLM Agents Know When Not to Act?, Liu et al., arXiv preprint, July 2026
- what it did: 263 pairs of tasks, one where acting is right and a slightly changed one where the agent should stop
- fact, abstract: âthe best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracyâ
- claim, abstract: âabstention capability is largely independent of general task-solving capabilityâ
- inference: refusing well does not come free with a stronger model, so a domain agent needs its own stop rule
- Verify, Repair, Repeat, or Stop?, Wu et al., arXiv preprint, under review, July 2026
- fact, abstract: with a noisy checker and a noisy repairer, ârepair can damage already-correct plans, and reported acceptance keeps rising while true validity fallsâ
- fact, abstract: their stop rule âimproves final true validity by 60.6 percentage points over fixed five-round repair at an average cost of 0.72 repair roundsâ in a GSM8K stress setting
- limit: a stress setting on arithmetic word problems
- inference: âbounded repairâ needs a reason for the bound; with a sound checker such as a compiler or proof checker, false accepts go away and this problem gets much easier
- HANDBOOK.md, Panavas et al., COLM 2026 workshop per the arXiv record
- fact, abstract: â65 agentic tasksâ; the best model âpasses 36.2% of trialsâ; agents âperform a required check and then act against its resultâ
- studied in depth in handbook.md (local source:
../../../handbook.md; not published in this study) and policy_following.md
what existing work shows: cheap models for easy parts, strong models for hard parts
- Specialize Roles, Mix Deployments (AgentCARD), Jiang et al., arXiv preprint, May 2026
- what it did: tried different models in the planner, executor, and verifier roles on five benchmarks, with one cost model for hosted and self-run models
- fact, abstract: mixed teams âimprove accuracy by up to 44% over cost-equivalent homogeneous teams, or match the strongest homogeneous team at up to 12Ă lower per-task cost through hybrid deploymentâ
- fact, abstract: âsome domains are planner-bottlenecked, while others are executor-bottleneckedâ
- limit, appendix D: âAgentCARD reflects a snapshotâ; the comparison is team against team, with no single-agent arm that I saw
- this conflicts with Kim et al. above
- fact, Kim et al. section 5: 13 mixed configurations on BrowseComp-Plus find âno evidence that model mixing bypasses the capability-saturation thresholdâ
- inference: both can hold; mixing saves money without lifting the ceiling the strong model sets
- SWE-Router, Son et al., ICML 2026 workshop per the arXiv record
- the idea, abstract: âlets a cheap model run for a few exploratory turns and reads the resulting partial trajectory before deciding whether to continue cheaply or to escalate to an expensive modelâ
- claim: the task text alone cannot tell an easy issue from a hard one
- limit: the abstract gives no number; SWE-bench Verified only
- AgentRouter, Paul and Nandy, ICML 2026 workshop per the arXiv record
- fact, abstract: a small classifier picks one of four model tiers for each step and reports â72% cost reduction relative to frontier-only baselines, retaining 97.3% of frontier-only qualityâ
- limit: I read only the abstract; trained on steps labelled by the authors
- inference across these
- mixing models is established and crowded
- the open part is the escalation signal; today it is a learned guess from the transcript
what changed in this revision
- the earlier version was written without web search; its quotes all check out against the sources
- AutoGen, MetaGPT, StateFlow, MAST, Kim et al., JARVIS, HANDBOOK.md, and MultiAgentBench quotes are present in the pages I opened
- its three proposals did not survive the search
- proposal 1, a runtime that stops agents from losing or duplicating work: built by CoAgent, STORM, and CAID, argued by the position paper, and proved in Verus by Khan
- proposal 2, recheck permission right before the action: this is MasuGateâs âstale authorizationâ and âpolicy-state serializabilityâ, and Cognitive Serializability covers policy and authority too
- proposal 3, route by measured failure: AgentCARD, SWE-Router, and AgentRouter cover routing; only the âsound checker as the signalâ part stays open
- its claim that Kim et al. âsubstantially narrows the novelty of a generic equal-compute team studyâ holds, and Tran and Kiela narrow it further
what is missing
- nobody shows where delegation starts to pay per dollar on long tasks
- evidence: CAID and STORM report the team costs more; Kim et al. match tokens but run 20 coding tasks with ±20 point intervals; Tran and Kiela match tokens on short questions only and predict a crossover they do not measure on tool tasks
- nobody separates the three things a team adds: fresh context, parallel work, more tokens
- the missing arm is one agent that restarts with a fresh context and a handoff note between parts, which has the fresh context and no parallel work
- nobody compares the concurrency runtimes on a shared workload
- evidence: CoAgent uses ten hand-made two-agent pairs and one model; Khanâs contention test is âsynthetic and single-modelâ; the position paper asks for the benchmark; Khan says current benchmarks do not stress shared state
- worker death in the middle of a write is not tested in any of them, as far as I read
- clashes of meaning, where both edits are fine alone and wrong together, are named by STORM and CooperBench as unsolved
- nobody measures how often a manager accepts a false âdoneâ from a worker on real work
- evidence: the false success study is on single agents; OrchestraBench uses an arithmetic chain and says it makes no workload claim; MAST traces come from benchmark runs of open frameworks
- I found no study of a manager and worker system that one person uses daily for months
- nobody tests whether a teamâs role split helps once the single agent has the same tools and checks
- evidence: JARVIS, AutoGenâs safeguard, and MetaGPT all change the checks and the roles together
- routing with a checker that cannot be fooled is untested
- evidence: SWE-Router and AgentRouter learn the signal; the verify and repair paper assumes a noisy checker
- I did not look for negative evidence beyond my searches, so each ânobodyâ means ânot found in the searches listed at the endâ
research we can do
- where delegation starts to pay
- question: on long coding tasks, at what task length and at what coupling between parts does a manager with workers beat one agent that gets the same dollars?
- coupling means how much one part needs to know about another
- why open: see the first gap; the theory in Tran and Kiela predicts a crossover and nobody has located it
- first experiment
- tasks: Commit0 and a set of fresh multi-file tasks, binned by size and by a coupling score from the dependency graph of the reference solution
- arms, all with the same model and the same dollar cap
- one agent, with the harnessâs normal context compaction
- one agent that restarts with a fresh context and a written handoff between parts
- manager with workers one at a time
- manager with workers in parallel in worktrees
- the same model run n times alone, best result picked by the tests
- count every token, including the managerâs, discarded work, and merges
- at least 100 tasks and 3 runs each, so intervals are a few points wide
- convincing result: a curve of success per dollar against size and coupling that shows where each arm wins, with intervals that do not overlap at the ends; âthe restart arm gets most of the gainâ would be a clean and useful negative result for teams
- cost: I guess a few thousand dollars of API use and 4 to 6 weeks for one person, mostly harness work
- closest work that could scoop it: Kim et al. (add coding tasks at scale), Geng and Neubig (add cost matching), OrchBench (already models information loss in a simulator)
- why us: the human runs the restart arm and the manager arm daily, so the harness knowledge is there
- an independent stress test for shared-state agent runtimes
- question: under the same contended workloads and injected faults, which runtime keeps the final state correct, and what does each cost in tokens and time?
- why open: see the second gap; every runtime paper grades itself
- first experiment
- workloads with a known set of correct end states: a shared repository, a small Kubernetes cluster, a ticket system with budgets
- contention as a dial: number of agents, share of objects touched by more than one agent, share of clashes that are clashes of meaning
- faults: kill a worker mid-write, delay a notification, deliver a message twice
- systems: plain shared directory, lock files as in the C compiler post, worktrees with merge, STORM, CoAgent if released, a durable workflow engine
- a checker that takes the log of tool calls and the end state and reports lost updates, stale reads, and orphaned half-done work, in the style of Jepsen
- a scripted worker arm with no model, to separate the runtimeâs fault from the modelâs
- convincing result: a table of anomaly counts per runtime per fault that the runtime authors can reproduce, plus at least one anomaly found in a published runtime
- cost: 2 to 3 months for one systems person; API cost small if a cheap model drives the workers
- closest work that could scoop it: the position paperâs authors (they also wrote Silo-Bench), Khan (has detectors and harnesses), the CoAgent group
- risk: this area moved fast between March and September 2026; I would check arXiv again before starting
- a Verus angle exists but is thin: Khanâs model assumes replayed deterministic output, and CoAgentâs guarantee rests on the agent judging a clash correctly, which no proof covers
- evidence in worker reports
- question: how often does a manager accept a false âdoneâ from a worker, and does forcing the report to carry evidence that plain code can check remove most of those cases cheaply?
- evidence here means things like the test command and its output hash, the diff, the file paths, which the harness re-runs or re-reads itself
- why open: see the third gap
- first experiment
- part A, measurement: take the humanâs own manager and worker logs, sample finished tasks, and check each âdoneâ against the repository and test state at that time; label with the MAST modes
- part B, control: on a coding benchmark with hidden tests, plant workers that return a wrong result with a confident report; compare a manager that reads the summary, a manager that reads the full trace, a model judge, and a harness that re-runs the evidence
- measure false accepts, false rejects, and tokens the manager spends on review
- convincing result: a false accept rate from real use with an interval, and an intervention that cuts it by a large factor at under, say, 10% extra tokens
- cost: part A is a few weeks and nearly free; part B a few hundred dollars
- closest work that could scoop it: Advaniâs false success study (extend to workers), OrchestraBench (extend to real workloads), TraceElephant
- limits: one userâs logs are a case study, and the logs may hold private text, so the human decides what leaves the machine
- smaller: escalate on a sound checker
- question: when a compiler or proof checker decides pass or fail, what is the cheapest rule for moving from a cheap model to a strong one, and does handing the strong model the cheap modelâs failed attempts help or hurt?
- first experiment: Verus or Lean proof tasks; arms are strong only, cheap only, cheap then strong with and without the failed attempts, and SWE-Routerâs learned rule
- closest work: SWE-Router, AgentRouter, AgentCARD, the verify and repair paper; the formal verification group in this study may cover the same ground from the prover side
- my order: 3A first because it is cheap and uses data only we have, then 1, then 2 if a fresh search still shows no shared benchmark
ChatGPTâs opinion
- pending: ChatGPT is unusable until the human signs in, so no consultation was run for this revision
- the earlier attempt also failed and no answer was used
what I searched
- date: 7 October 2026
- web searches, 15 in all, with these themes
- one agent against a team at equal token budget
- manager, worker, delegation, and orchestration failure benchmarks
- model routing and cascades for agents; strong planner with cheap executor
- fixed workflow against free agent
- several coding agents on one repository, merge conflicts, worktrees
- subagent context isolation and information lost in handoff
- debate against asking one model several times
- domain agents, refusing to act, bounded repair
- transactions and concurrency control for agent teams; a Jepsen-style checker
- CooperBench; agreement between peer agents; false âdoneâ reports
- sources opened: 41 arXiv records with full text where it loaded, 1 ACL Anthology page, 2 Anthropic engineering posts; 41 of the 44 are cited above
- full text failed to load for the debate position paper and AgentRouter, so those rest on abstracts
- I checked quotes by searching the downloaded text for the exact words
- not covered
- no paper read cover to cover; numbers in tables other than the ones quoted are unchecked
- venues: I report a venue only when the arXiv record states it; several 2023 and 2024 papers likely have one that I did not check
- Google Scholar citation chasing, OpenReview reviews, and code repositories
- automatic design of agent teams beyond AFlow; debate methods beyond one position paper
- human teams and organization research
- industry systems other than the two Anthropic posts
- the paper collection at /hdd1/sichanghe/paper_collection was not consulted for this file
Last edited: