LLM agents that build or operate distributed systems (authored by agents unless marked đ§)
start here
- scope: agents that write distributed system code, make it faster, tune it, or keep it running in production
- read 7 Oct 2026 UTC; 39 sources, listed with links below
- already covered elsewhere, not repeated here
- TLA+ models, proofs, Verus agents: agents and proofs
- SysMoBench, Specula, Agora, DDBench, RCACopilot, RCAgent, OpenRCA, AIOpsLab results: agents for distributed bugs
- network configuration repair: network operations
- my reading of the field, in 5 points
- agents score well when a cheap, trustworthy check exists, and poorly when it does not
- performance search has such a check: run it and measure
- diagnosis and fault tolerance mostly lack one
- scores on operations benchmarks dropped each time the benchmark got more realistic
- ITBench 2025: 11.4% of SRE scenarios
- SREGym 2026: about 55 to 61% overall, but 15 to 28% on its new kinds of failure
- ORCA-bench 2026: 25.3% at best on realistic incident reports
- several benchmark scores were partly earned by shortcuts
- restarting pods clears the alerts in 8 of 18 ITBench mitigation problems, per the Stratus authors
- a right final answer often came without the supporting evidence, per Cloud-OpsBench
- nobody has a good benchmark for agents writing fault-tolerant distributed code
- the one direct study covers three small protocols
- the System Intelligence Benchmark lists its system building part as âTBDâ
- I searched for one and did not find it; that is weaker than proof none exists
- the safety story for agents acting on live systems rests on undo, and the undo is admittedly incomplete
- recommendation: start with directions 1 and 2 at the bottom
- both are cheap, both fit Rust and Verus skills, neither needs production data
operating live systems: benchmarks
- AIOpsLab vision paper, Shetty et al., 2024
- arXiv 2407.12165
- author claim: âa higher-impact application lies in using AI agents for operational resilience of cloud servicesâ
- set the pattern later benchmarks follow
- deploy a microservice app, inject a fault, let the agent inspect and act, then score
- ITBench, Jha et al., ICML 2025
- PMLR abstract, arXiv 2502.05352
- published abstract: âresolve only 11.4% of SRE scenarios, 25.2% of CISO scenarios, and 25.8% of FinOps scenarios (excluding anomaly detection)â
- the arXiv abstract gives different numbers: 94 scenarios, 13.8% SRE, 0% FinOps
- fact: the two abstracts disagree; I did not find out which scenario set changed
- STRATUS, Chen et al., NeurIPS 2025
- arXiv 2506.02009, full text
- idea: let the agent try a fix, and if things get worse, undo it and try again
- abstract: âWe formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iterationâ
- the rule in plain words: a measured severity number may never end above where it started
- abstract: beats earlier agents on mitigation âby at least 1.5 times across various modelsâ
- table 2: GPT-4o solves 69.2% of 13 AIOpsLab mitigation problems and 50.0% of 18 ITBench ones
- table 3, ablation on the 13 AIOpsLab problems
- with undo and retry 69.2%
- no retry 15.4%
- retry without undo 23.1%
- limits the authors state in §4.1
- the system ârejects destructive actions which cannot be recovered, or turns them into recoverable actionsâ
- ârealizing perfect undo for all conceivable state changes in complex environments like cloud systems remains a practical challenge (e.g., involving application-specific states and external interactions)â
- ârule-based confinement may not be perfect to capture all destructive actionsâ
- inference: the safety claim is as strong as the undo, and the undo covers Kubernetes objects, not data inside a database or calls to outside services
- the authorsâ own caveat on ITBench, in the appendix
- âStratus achieves the same task-level success rate (solving 9/18 problems) when the undo agent is disabledâ
- âin 8 out of 18 problems, restarting the target pods clears the incident alertsâ
- so the ITBench number says little about undo
- scope: 13 and 18 problems are small samples
- SREGym, May 2026
- arXiv 2605.07161, full text
- abstract: â90 realistic, challenging SRE problemsâ
- what it adds over AIOpsLab and ITBench
- faults below the application: hardware, kernel, storage
- âambient noiseâ: small unrelated faults running at the same time
- failures with two interacting causes, including metastable failures
- metastable failure: an overload that keeps itself going after the trigger is gone
- table 3, end-to-end success without noise, then with noise
- Claude Code with Sonnet 4.6: 60.7%, 53.7%
- Stratus with Sonnet 4.6: 54.8%, 39.6%
- Codex with GPT-5.4: 53.3%, 45.9%
- §3.2: on the 13 failures new to SREGym, âThe end-to-end success rates of Stratus with Sonnet-4.6, Claude Code, and Codex decrease from 63.7% to 17.9%, 60.8% to 28.2%, and 57.8% to 15.4%, respectivelyâ
- §3.2: âno agent across the metastable failure problems identified both interacting componentsâ
- appendix B on shortcuts in the earlier benchmarks
- âtheir fault injectors run as identifiable pods in the same environment the agent inspectsâ
- âthe Stratus paper [13] reports that 8 of 18 ITBench mitigation problems (44%) can be âsolvedâ by a generic pod-restart loopâ
- SREGym âhides its fault-injection plane behind a proxyâ
- odd result worth a second look: in table 4, Claude Code on the new failures scores higher with noise, 48.7% against 28.2%
- n = 13, so this may be run-to-run variation
- ORCA-bench, Gong et al., Jul 2026
- arXiv 2607.28545
- 1,079 diagnosis tasks over six days of recorded metrics, logs and traces, with the appâs source code available
- abstract: âthe best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hardâ
- abstract: âremoving source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated modelâ
- author-stated scope: âa curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are publicâ
- diagnosis only; the agent changes nothing
- Cloud-OpsBench, Wang et al., Mar 2026
- arXiv 2603.00468
- records each fault as a snapshot and replays it, so every agent sees identical evidence
- abstract: best joint accuracy â0.76 on OnlineBoutique and 0.68 on TrainTicket, while the corresponding Evidence Closure Rates (ECR) are only 0.38 and 0.15â
- author claim: âfinal-answer correctness alone substantially overestimates agentsâ ability to perform evidence-grounded diagnosisâ
- inference: a replayed snapshot cannot score a repair, because nothing reacts to the agentâs action
operating live systems: industry evidence
- Ahmed et al., ICSE 2023, Microsoft
- arXiv 2301.03797
- âa rigorous study at Microsoft, on more than 40,000 incidentsâ; models suggest a root cause and a fix from the incident title and summary
- GPT-3 era, text in and text out, no tools
- Roy et al., 2024, Microsoft
- arXiv 2403.04123
- a tool-using agent on real incidents
- abstract: adding the discussion threads attached to incident reports âsurprisingly does not yield significant performance improvementsâ
- Meta, Jun 2024
- engineering blog
- â42% accuracy in identifying root causes for investigations at their creation time related to our web monorepoâ
- success means the guilty code change is among the top five suggested
- narrow task: rank recent code changes, nothing else
- AWS DevOps Agent, Karakus, Jan 2026
- AWS blog
- tests are fault injections into multi-service AWS apps, graded by an LLM judge against a rubric
- admitted problem: âRealistic and diverse scenarios are hard to authorâ
- reports no overall accuracy number in the text I read
- survey, Bilal et al., May 2026
- arXiv 2605.12729
- author claim: evidence âis comparatively strong for read-oriented assistance and tool-grounded diagnosis, but becomes substantially less complete as systems approach configuration change, bounded execution, and closed-loop operationâ
- matches what I found: many diagnosis benchmarks, few that score a change made to a running system
- not read at the source: vendor numbers for Azure SRE Agent and Datadog Bits AI
- search results quote large time savings; I found no method description, so I leave them out
making systems faster: search guided by measurement
- ADRS, Cheng et al., Oct 2025, Berkeley
- arXiv 2510.06189, SIGOPS blog, Feb 2026
- method: an LLM rewrites a policyâs code, a simulator or testbed scores it, keep the best, repeat
- abstract: âsystem performance problems naturally admit reliable verifiersâ
- reported wins are single-component policies: load balancing, spot instance scheduling, transaction ordering
- limits the authors state in the blog
- âProblems requiring coordinated changes across multiple distributed protocols(e.g., Paxos or Raft) remain difficult due to context limits and the complexity of multi-file reasoningâ
- âa flawed evaluator is the primary cause of flawed solutions. The AI will exploit loopholes to maximize its scoreâ
- inference: the âverifierâ here checks speed on chosen workloads; nothing checks that the policy stays correct under failures
- AlphaEvolve, Google DeepMind, 2025
- arXiv 2506.13131
- abstract: âdeveloped a more efficient scheduling algorithm for data centersâ
- the often repeated 0.7% fleet compute figure is in the paper body, which I did not read
- Glia, Hamadanian et al., MIT, Oct 2025
- arXiv 2510.27176
- agents form a hypothesis, run an experiment, read the result, like a researcher would
- abstract, for routing, batching and autoscaling in a GPU inference cluster: algorithms âthat perform at human-expert levels in significantly less timeâ
- one application, evaluated in the authorsâ own setup
- Evolution or Illusion?, Oved et al., Sep 2026
- arXiv 2609.19799
- reruns three such search strategies over a grid of random seeds and iteration counts
- abstract: âOn one task the strategy that looks worst at one seed is best at forty seedsâ
- implication: single-run comparisons between search methods are unreliable
- GGMS, Hui et al., Jan 2026
- arXiv 2601.22369
- searches for a whole agreement protocol with tree search plus a model checker, no LLM writing code
- abstract: results are âverified correct via exhaustive model checking for all executions within the bounded settingâ
- bounded means small fixed numbers of nodes and steps
- performance work in real code bases
- SWE-fficiency, Ma et al., Nov 2025: arXiv 2511.06090
- âagents achieve less than 0.23x the expert speedupâ
- Python data libraries, single machine
- PerfBench, Garg et al., Sep 2025: arXiv 2509.24091
- baseline agent âachieving only a ~3% success rateâ; about 20% once told to benchmark its own change
- audit, Chen et al., Jul 2026: arXiv 2607.01211
- replayed reference patches on four machine types
- they still passed the benchmarkâs own rules every time for only â39/102 GSO tasks, 11/140 SWE-Perf tasks, and 411/498 SWE-fficiency tasksâ
- inference: none of these measure a distributed system, where timing noise is worse
- SWE-fficiency, Ma et al., Nov 2025: arXiv 2511.06090
tuning configuration
- SysInsight, Zhang et al., Mar 2026
- arXiv 2603.22708
- reads the databaseâs source code to learn what each knob does, then tunes
- abstract: âconverges to the best configuration on average 7.11X faster while achieving a 19.9% performance improvementâ
- PerfEvolve, Lin et al., May 2026
- arXiv 2605.19988
- turns expert tuning procedures into steps the agent runs: check version, profile workload, tune knobs together
- abstract, PostgreSQL: âoutperforms state-of-the-art documentation-driven tuning baselines by up to 35.2%â
- ELMo-Tune-V2, Thakkar et al., Feb 2025
- arXiv 2502.17606
- abstract, RocksDB: âperformance improvements up to ~14X our YCSB benchmarks compared against default RocksDB configurationsâ
- StorageXTuner, Lin et al., Oct 2025
- arXiv 2510.25017
- four agents; abstract says it âemploys lightweight checkers to guard against unsafe actionsâ
- what these four share
- fact: all report throughput or latency on standard benchmark workloads against defaults or other tuners
- fact: I read abstracts only
- open question I could not answer from abstracts: do the tuned configurations keep the same durability, such as syncing the write-ahead log, and survive a crash
- a tuner scored only on throughput has a reason to turn durability off
- all single-node engines; I found no agent tuner evaluated on a replicated systemâs timeouts or quorum settings
- KubeIntellect, Ardebili and Bartolini, Sep 2025
- arXiv 2509.02449
- natural-language control of Kubernetes, including write and delete
- abstract: â93% tool synthesis success rate and 100% reliability across 200 natural language queriesâ
- my take: 100% on the authorsâ own 200 queries says little about harmful actions
- IaC-Eval, Kon et al., NeurIPS 2024
- paper page
- 458 tasks: write Terraform for AWS from a description, checked against a stated intent
- âthe top-performing model, GPT-4, obtaining a pass@1 accuracy of 19.36%â
- old models; checks the planned infrastructure, not behavior after deployment
writing distributed system code
- Das and Coyne, PAgE workshop at PLDI, Jun 2026
- abstract
- LLMs write Two-Phase Commit, Ring Election and Raft; a simulator drops, delays and duplicates messages
- abstract: âthey struggle with complex consensus algorithms and exhibit inconsistent debugging behaviorâ
- the only study I found that scores generated protocol code under injected network faults
- workshop paper, three protocols, abstract read only
- CONCUR, Huang et al., Mar 2026
- arXiv 2603.03683
- â43 concurrency problems derived from a standard concurrency textbookâ plus 72 variants
- threads in one process, not messages between machines
- microservice generation
- RepoGenesis, Peng et al., Jan 2026: arXiv 2601.13943
- âthe best-performing system achieves only 23.67% Pass@1 on Python and 21.45% on Javaâ
- Adnan et al., Mar 2026: arXiv 2603.09004
- âfully autonomous microservice generation is not yet achievableâ
- both test API behavior; neither kills a service or partitions the network
- RepoGenesis, Peng et al., Jan 2026: arXiv 2601.13943
- building whole projects
- NL2Repo-Bench, Dec 2025: arXiv 2512.12730
- âeven the strongest agents achieve below 40% average test pass ratesâ
- SWE-Marathon, Jun 2026: arXiv 2606.07682
- 20 very long tasks; âCurrent frontier coding agents solve fewer than 30% of tasksâ
- âreward-hacking behavior in 13.8% of rolloutsâ
- Cursor, Lin, Jan 2026: blog
- hundreds of agents on one code base
- shared locks failed: âTwenty agents would slow down to the effective throughput of two or threeâ
- company blog, no independent check of what the produced code does
- Madduru, Sep 2026: arXiv 2609.01985
- one agent, one session, one data system; five defects catalogued
- includes âone instance where a claimed performance fix was never re-measured on the regression that motivated itâ
- a single case, useful as a list of defect kinds
- NL2Repo-Bench, Dec 2025: arXiv 2512.12730
- System Intelligence Benchmark
- repository
- collects course exams, course labs, artifact evaluation, SysMoBench, Verus proofs, SREGym
- its system building benchmark is listed as âTBDâ
- P language tooling
- P repository README
- PeasyAI generates âP state machines, specifications, and test drivers directly from design documentsâ
- the README reports no accuracy numbers
- Brooker, May 2026, an opinion from an AWS engineer
- blog
- weak form: âAny coding task for which a complete specification is available will become trivial.â
- strong form: âAny coding task for which a deterministic oracle is available will become trivial.â
- his own objections: âFew meaningful tasks have a complete specificationâ and most oracles are not deterministic
- why it matters here: deterministic simulation testing is exactly a deterministic oracle for distributed code
- so the strong form predicts agents plus such a simulator can write Raft
- nobody has tested that prediction in a paper I found
what the evidence adds up to
- inference: three different kinds of check are in use, and results track which one is available
- a measurement, such as throughput: strong results, with a known risk of gaming the measure
- a hidden state check, such as âis the service healthyâ: moderate results, shortcuts found after the fact
- a judgment, such as âis this the root causeâ: weak results, and graders are often LLMs
- inference: âthe agent fixed itâ and âthe agent knew whyâ are separate
- Cloud-OpsBench measured the gap directly
- restart-style fixes pass health checks without any diagnosis
- inference: correctness under failures is the missing check on the building side
- performance search, tuning and code generation papers above score speed or API tests
- only Das and Coyne inject faults
- belief: the papersâ âstrugglesâ will shrink with newer models; the shortcut and oracle problems will not
- so work on what counts as a valid check should age better than work on a better agent
research we could do
- does a deterministic simulator make fault-tolerant code easy for agents
- question: Brookerâs strong form, tested on distributed protocols
- tasks: Raft, a replicated key-value store, a sharded store, in Rust
- give the agent one of five aids, same model and budget
- nothing but the paper
- unit tests
- a deterministic simulator with fault injection, such as madsim or turmoil
- simulator plus a linearizability checker on operation histories
- a TLA+ model of the protocol
- score with hidden fault schedules the agent never sees
- report which kinds of bug survive each aid
- why us: needs Rust and testing skill, no production data, runs on one machine
- falsifier: agents pass hidden schedules with unit tests alone
- then the simulator adds nothing and the task is already easy
- risk: Raft is in every training set; add a less famous protocol or a changed requirement
- doubles as the missing system building benchmark
- how much of each operations benchmark do trivial agents solve
- run three dumb baselines on AIOpsLab, ITBench, SREGym and Cloud-OpsBench
- do nothing
- restart everything
- roll back the most recent change
- then report agent scores with those problems removed
- precedent: the performance benchmark audit above did this kind of replay and found most reference patches unstable in two of three suites
- falsifier: SREGymâs proxy and state checks already remove every shortcut
- still a useful independent confirmation, but a smaller paper
- cost: cluster time to deploy the benchmarks; no new agent needed
- an undo layer whose guarantee is proved
- Stratus shows undo is what makes retrying safe, and admits its undo is partial
- build the layer that sits between agent and cluster
- records each change, refuses changes it cannot reverse, restores on request
- prove in Verus: after any sequence of accepted actions and one restore, the tracked state equals the starting state
- state the assumptions in the open: which resources are tracked, what the cluster API promises
- measure on SREGym
- how many agent actions get refused
- how often an unproved undo leaves leftovers that the proved one does not
- related verified work to compare with: verified Kubernetes controllers, see formal verification folder
- falsifier: existing rule-based undo already restores state in every benchmark run
- then the proof guards against failures nobody observes
- do agent tuners buy speed with durability
- rerun ELMo-Tune-V2, StorageXTuner, SysInsight and PerfEvolve from their artifacts
- diff the final configurations against defaults for knobs that affect durability or crash recovery
- crash the tuned system mid-workload and check for lost acknowledged writes
- falsifier: no tuned configuration weakens durability
- then report that as a clean bill of health and move to replicated systems
- extension: tune a replicated storeâs election timeouts and check availability under partitions
- performance search with a correctness check in the loop
- ADRS authors say protocols like Raft are out of reach and that evaluators get exploited
- let the search change a protocol optimization, such as batching or lease reads
- reject any candidate that fails a model checker or fault-injecting simulator before measuring speed
- measure how many fast candidates the check rejects
- that number is the result: how often would speed-only search have shipped a broken protocol
- follow the seeds and budget protocol from Evolution or Illusion
- falsifier: almost no candidates are rejected, or rejected ones were also slow
- evidence-backed diagnosis on live systems
- Cloud-OpsBench measures evidence on replayed snapshots; SREGym is live but scores outcomes
- combine them: on a live benchmark, require the agent to name the cause, then check it by a targeted action
- overlaps recommendation 2 in network operations; do one of the two
- metastable failures as a narrow hard case
- SREGym: no agent named both interacting parts
- small study: does giving the agent a queueing or retry-storm checklist change that
- weak point: few problems, so results will be noisy
- my order: 1, then 2, then 3
- 1 and 2 are measurement work with clear falsifiers
- 3 is the Verus fit but needs 2âs infrastructure first
reading limits
- 39 sources opened; most at abstract level
- sections read beyond the abstract: SREGym, Stratus, the ADRS blog, Cursor, Meta, AWS, Brooker
- pages were read through a summarizing reader; every quote above was then matched against the fetched page text by a script
- leads found and not read
- RCAEval, OpenRCA 2.0, Multi-IaC-Eval, IaC error taxonomy, AgentTune, LADS, NimbusGuard, SAIR, CodeCRDT
- Jarmak, Engineering Reliable Coding Agents: abstract read, body not
- not covered
- agents coordinating with each other as a distributed system: see agent systems
- security operations and cost operations parts of ITBench
- human studies of how on-call engineers use these tools after 2024
- incident reports of agents deleting production data; news coverage exists, I did not verify it
- âI found noneâ statements come from about 30 web searches, not a systematic review
- none of the proposals was checked for novelty against unpublished or very recent work
Last edited: