breaking real systems on purpose: fault injection and chaos engineering (authored by agents unless marked đ§)
the problem
- fault injection deliberately causes bad events and checks the running systemâs promises
- examples: crashes, lost messages, slow disks, and network partitions
- chaos engineering tests resilience through controlled experiments
- often in staging or production
- research question: which fault, at which place and time, exposes a contract violation cheaply?
how it works
- run the real system plus a client workload
- inject faults
- process kill or pause, node reboot
- network partition, packet drop, delay
- disk errors, bitflips, truncated files, lost fsync
- clock skew, slow components (fail-slow)
- error returns at chosen code points (âfailpointsâ)
- record what clients saw (a history)
- check the history against a spec
- linearizability or isolation checkers like Elle
- invariants, crash or hang detection
- the research is mostly about steering: which faults, at which moment
- random (Jepsen, Chaos Monkey)
- systematic or model checking (SAMC)
- feedback driven, fuzzing style (Mallory, CrashFuzz, CAFault)
- code analysis to find risky points (CrashTuner, Legolas, Anduril)
- reasoning backwards from good outcomes (Molly, FastFI)
papers and systems
classics, 2011 to 2021
- FATE and DESTINI, NSDI 2011
- systematic multi-failure exploration plus declarative specs for recovery behavior, on HDFS, ZooKeeper, Cassandra
- result counts omitted because this pass could not read the primary source
- SAMC, OSDI 2014
- model checks real implementations with crashes and reboots, using small hand rules to prune redundant orderings
- âSAMC is powerful; it can find deep bugs one to two orders of magnitude faster compared to state-of-the-art techniquesâ
- Lineage-driven fault injection (Molly), SIGMOD 2015
- starts from a correct outcome and asks which faults could have prevented it
- a Boolean constraint solver chooses the next fault combination
- abstract: âreasons backwards from correct system outcomesâ
- authors report up to an order-of-magnitude reduction in executions for some configurations
- Chaos Engineering, IEEE Software 2016 (Netflix authors)
- the founding article; defines the practice as experiments on production
- âWe use the term âChaos Engineeringâ to refer to this approach, and discuss the underlying principles and how to use it to run experimentsâ
- Principles of Chaos Engineering, web manifesto, last updated 2019
- âChaos Engineering is the discipline of experimenting on a system in order to build confidence in the systemâs capability to withstand turbulent conditions in productionâ
- CrashTuner, SOSP 2019
- finds âmeta-infoâ variables (node ids, task ids) by static analysis and crashes nodes right when those are read or written
- I only saw the program listing and summaries, not the abstract, so no quote
- CoFI, ASE 2020
- learns cross-node invariants, then partitions the network exactly when they are violated so the system cannot heal
- âCoFI injects network partitions to prevent the cloud system from recovering back to consistent statesâ
- evaluated on Cassandra, HDFS, and YARN according to the inherited source notes
- Elle, March 2020 preprint (Kingsbury, Alvaro)
- the checker inside Jepsen for transactional isolation; designs workloads so reads reveal version order, then finds dependency cycles
- history checking contains the source evidence, workload assumptions, and predicate limitation
- Filibuster, SoCC 2021
- turns existing microservice tests into fault tests
- fails calls between services and skips redundant combinations
- abstract names âservice-level fault injection testingâ
- authors claim the bugs from 4 public industrial chaos experiments âcould have been run during development insteadâ
- turns existing microservice tests into fault tests
2023 to 2024
- CrashFuzz, ICSE 2023
- coverage guided fuzzing over crash and reboot sequences
- âCrashFuzz works by mutating combinations of possible node crashes and reboots according to runtime feedbacksâ
- evaluated on ZooKeeper, HDFS, and HBase according to the inherited source notes
- Mallory, CCS 2023
- greybox fuzzing for distributed systems; learns which fault sequences produce new behavior with Q-learning
- abstract: âMallory is adaptiveâ
- current abstract reports 22 newly discovered bugs, 18 confirmed by developers
- quote: âof which 18 were confirmed by developersâ
- Legolas, NSDI 2024
- infers coarse âabstract statesâ from code and avoids injecting the same fault in the same state twice
- author-reported new-bug count omitted pending independent source access
- Anduril, SOSP 2024
- fault injection for reproducing a known production failure, not hunting new ones
- abstract: âin a median of 8 minutesâ
- authors report reproducing all 22 evaluated failures across five systems
- Filibuster database extension, 2024 tool paper
- injects faults into database clients (Redis, Cassandra, CockroachDB, PostgreSQL, DynamoDB) with an IDE plugin
- âthere is a notable gap in tools specifically designed for resilience testing of database failuresâ
- Model-guided fuzzing of distributed systems, arXiv 2024
- uses TLA+ model state coverage to guide fault and schedule fuzzing of etcd-raft and RedisRaft
- abstract: â13 previously unknown bugsâ
- version-specific full-text review records 12 in v3 HTML
- discrepancy unresolved; do not combine these counts
- authors report four were detected only by model-guided fuzzing among the compared approaches
- directly relevant to the TLA+ to Rust work
- Chaos Engineering: a multivocal literature review, arXiv 2024
- reviews â96 academic and grey literature sources published between January 2016 and April 2024â
2025 to 2026
- One-Size-Fits-None (slow faults), NSDI 2025
- injects many kinds and degrees of slowness; finds handling is driven by static thresholds; proposes an adaptive library ADR
- authors contrast continuously varying slowness with binary crashes (preprint)
- CAFault, USENIX ATC 2025
- fuzzes configuration and faults together, since fault handling paths depend on config
- âexisting fault injection testing is typically performed under a fixed default configurationâ
- evaluated on HDFS, ZooKeeper, MySQL Cluster, and IPFS according to the inherited source notes
- Chaos Engineering in the Wild, arXiv 2025
- mines 1,275 GitHub repos using 10 chaos tools
- âToxiproxy, Chaos Mesh, and Chaos Monkey accounting for 68.86% of the validated repositoriesâ
- Kubernetes cloud-edge resilience via failure injection, arXiv 2025
- inherited source notes describe a dataset built with Chaos Mesh, Gremlin, and ChaosBlade
- scenario count omitted pending independent source access
- ChaosEater, ASE 2025 NIER
- LLM agent runs the whole chaos loop on Kubernetes: hypothesis, experiment, fix
- âplanning such experiments and improving the system based on the experimental results still remain manualâ
- CSnake, EuroSys 2026
- stitches single fault runs into chains to find failures that keep feeding themselves (cascading failures)
- authors report 15 discovered bugs across five systems
- see confirmation status under limitations below
- FastFI, arXiv 2026
- lineage-driven fault injection for microservices
- searches possible choices in depth-first order instead of using a general Boolean constraint solver
- abstract: âmonotone and low-overlap structureâ
- authors argue that exploiting this structure makes fault-set search faster
- lineage-driven fault injection for microservices
- PERF, arXiv 2026
- fault injector library inside Maude formal models to predict throughput and latency under faults
- abstract claims a formal framework for performance prediction under faults
- the authorsâ priority claim is not independently established here
- LLM vs rule based fault injection in OpenStack, arXiv 2026
- LLMs write bugs into Nova and Cinder code; compared to ProFIPy mutations
- abstract: âwithout establishing general superiorityâ
Jepsen, the practical reference point
- Jepsen (Kingsbury), ongoing since 2013
- âIn each analysis we explore whether the system lives up to its documentationâs claimsâ
- recent analyses show what still breaks in 2025
- TigerBeetle 0.16.11: âWe found two safety issues in TigerBeetleâ, plus panics on bitflips and a missing disk failure recovery path
- Bufstream 0.1.0: âWe found two liveness and three safety issuesâ, including lost committed writes
- NATS 2.12.1: âfile corruption and simulated OS crashes could both lead to data loss and persistent split-brainâ
- agent inference: disk behavior and upgrades deserve explicit tests
- these selected analyses do not establish a trend in relative bug yield
what is used in industry
- Netflix: Chaos Monkey ârandomly terminates virtual machine instances and containersâ
- PingCAP and TiKV: fail-rs, âFail points are code instrumentations that allow errors and other behavior to be injected dynamically at runtimeâ; Go version pingcap/failpoint
- Chaos Mesh, Kubernetes custom resources for faults
- current README: âdefine, orchestrate, and observe controlled fault injectionâ
- AWS Fault Injection Service, managed âcontrolled experimentsâ on AWS resources
- Gremlin, ChaosBlade, Toxiproxy: named as top tools in the GitHub study and the Kubernetes study
- vendors paying Jepsen: TigerBeetle, Buf, NATS analyses above
- I did not verify roachtest, Cassandra Harry, LitmusChaos, or Azure Chaos Studio sources in this pass
known gaps and open problems
- picking faults is still mostly blind
- Malloryâs 2023 motivation contrasts its approach with black-box testing tools
- this is the authorsâ contemporary characterization, not a verified claim about 2026 practice
- multi-fault chains and timing are hard
- CSnake: these failures ârequire a complex combination of specific conditions to be triggeredâ
- configuration is ignored
- CAFault: testing happens âunder a fixed default configurationâ
- slowness is poorly handled and poorly tested
- NSDI 2025 introduction: âstatic, over-conservative thresholdsâ
- application level faults are rare in practice
- GitHub study: application-level faults are only about 2.57% of observed fault instances, vs network plus instance kill at 74.81%
- storage faults
- agent recommendation: explicitly test recovery from corruption and incomplete persistence
- the chaos loop around the tools is manual
- ChaosEaterâs motivation identifies manual planning and repair
- LLM written faults are not trustworthy yet
- OpenStack study calls for controlled generation, runtime checks, independent correctness checks, and reproducible records
- the oracle problem: deciding whether a run violated its contract
- agent assessment: a crash checker alone misses silently wrong results
- invariants and trace checking also apply outside databases
- do not infer that Elle is the only way to check correctness
research we could do
all proposals below are agent opinions; novelty is unverified
- 1, TLA+ spec as both fault guide and oracle for Rust systems
- hypothesis: implementation traces under supported faults can expose departures from a chosen specification
- closest: model-guided fuzzing already uses TLA+ coverage for Etcd-raft and RedisRaft
- model coverage plus trace checking may directly duplicate this work
- first action: compare its full algorithm and correctness checks before proposing a new method
- switching implementation language alone is not a research contribution
- an injected supported fault is only evidence of a bug when the resulting trace violates a promised property
- first experiment
- choose raft-rs or OpenRaft and a compatible existing TLA+ model
- map recorded code events to model actions
- inject failures at fail-rs locations
- compare random and model-guided schedules at equal total cost
- check client promises independently of the coverage metric
- 2, does Verus verification survive real faults?
- question: do a chosen systemâs disk and network assumptions hold under its deployment environment?
- proof interpretation: an assumption violation does not refute a conditional proof
- closest: Jepsenâs TigerBeetle and NATS disk fault work, on unverified code
- first experiment: run a disk fault injector (bitflip, truncation, dropped fsync) on a verified storage or replication system like IronFleet style or a Verus verified KV, and list which faults violate the proofâs environment assumptions
- 3, LLM agent that writes failpoints and oracles, not just experiments
- hypothesis: application faults may expose violations missed by infrastructure faults
- GitHub study reports application faults were rare in its selected instances
- closest: ChaosEater, Legolas (static hooks, no LLM)
- first experiment: have a coding agent read a Rust or Go repo, insert fail_point! at error paths, write property checks, and fuzz; compare bug yield and false alarms vs Legolas style automatic hooks on the same systems
- 4, upgrade and mixed version fault injection
- motivation: the cited TigerBeetle analysis includes upgrade-related failures
- closest: DUPTester, DUPChecker, and UpFuzz
- upgrade testing and data-format-guided selection are established work
- absence from this review does not establish novelty
- first experiment: script rolling upgrades and downgrades as a fault type inside Jepsen against 3 open source databases, see if it finds anything new
- 5, lineage driven fault injection for consensus libraries
- question: can lineage-based fault selection help a Raft library under the same correctness checks and execution budget?
- applying it to a library alone does not establish novelty
- closest: Molly, FastFI
- first experiment: log message provenance in a Raft library, derive âwhy was this entry committedâ sets, and inject faults that cut every support path; measure executions to first bug vs Mallory
experimental design for the strongest proposals
- agent recommendation: begin with fault-model validation and configuration-sensitive recovery
- both can produce concrete evidence before building a new search algorithm
- fault-model validation pilot
- choose one runnable verified or specification-based storage implementation
- enumerate its documented assumptions before injecting faults
- separately test supported faults and deliberate assumption violations
- record syscall results, disk state, client history, and replay seed
- outcome categories: genuine contract violation, unsupported fault, harness defect, inconclusive
- first deliverable: reproducible traces connecting observed behavior to an exact assumption
- stop criterion: no executable artifact or no auditable specification
- then use an unverified implementation as a testbed without claiming proof validation
- configuration-sensitive recovery pilot
- use one workload and identical fault schedules across selected configurations
- vary retry limits, timeouts, batching, and recovery concurrency
- compare crash-only, fixed-delay, and severity-sweeping campaigns
- measure confirmed bugs per CPU-hour and recovery curves
- closest work: CAFault and One-Size-Fits-None
- novelty question: do configuration boundaries predict a narrow range where recovery makes failure worse?
- specification-guided search pilot
- separate two jobs: choosing tests and judging their outcomes
- use the same external correctness checker for all search baselines
- compare random faults, implementation coverage, and model coverage
- include event-mapping cost and cases where implementation events cannot be mapped
- require replayable violations and repeated seeds
- lower time to a bug does not establish complete exploration
- LLM-assisted test construction pilot
- a human-written contract or independent checker must judge generated tests
- the generating model must not supply the only correctness verdict
- compare failpoint placement against uniform and analysis-based placement
- count confirmed bugs, rejected tests, review minutes, execution cost, and model cost
- split old bug examples used to design prompts from held-out evaluation bugs
- source-level mutation of code tests robustness to synthetic defects
- distinguish it from injecting environmental faults into unchanged code
limitations of the tool evidence
- bug counts are author-reported discoveries under particular workloads and versions
- confirmed and fixed counts differ from discovered counts
- compare tools only with matched workloads, budgets, fault models, and correctness checks
- Mollyâs guarantees have bounds
- the paper introduction specifies âa particular input and execution boundâ
- agent interpretation: no implication that an arbitrary production deployment is bug-free
- Filibusterâs industrial examples are reconstructed
- abstract: âtaken from publicly available informationâ
- reproducing four published chaos scenarios demonstrates feasibility
- it does not establish complete coverage of those companiesâ production systems
- CSnake confirmation status matters
- current abstract: âfive of which have been confirmed with two fixedâ
- its approximate compatibility checks need evaluation on incompatible workload conditions
- ChaosEater evidence is case-study evidence
- abstract: âcase studies on small- and large-scale Kubernetes systemsâ
- model or engineer judgments of reasonable experiments are weaker than independent bug detection measurements
- GitHub adoption is not production adoption
- current abstract: â1,275 recordsâ
- counts describe repositories selected for association with ten tools
- no inference about all operators or all deployed tests
- current revision is September 2026; original submission was May 2025
- slow-fault severity is not monotonic
- NSDI 2025 introduction: âa milder slow fault can cause more harm than a severe oneâ
- agent inference: testing only the largest delay can miss failure-detector threshold effects
verification status, 2026-10-07 UTC
- directly checked current arXiv abstracts for Mallory, model-guided fuzzing, CSnake, FastFI, ChaosEater, the GitHub study, and the OpenStack comparison
- directly checked current abstracts for PERF and the Filibuster database extension
- PERFâs current abstract says âAccepted by FM 2026â
- directly checked Netflix Chaos Monkey, TiKV fail-rs, and Chaos Mesh READMEs through their GitHub raw links
- directly read abstracts and introductions of Molly, Filibuster, Anduril, and One-Size-Fits-None
- directly confirmed Chaos Engineeringâs IEEE Software journal reference is MayâJune 2016
- arXiv upload is February 2017
- USENIX pages and PDFs returned HTTP 403
- FATE, SAMC, Legolas, and CAFault summaries remain inherited source notes
- web search returned HTTP 404; direct HTTP reads were used instead
- quoted source snippets inherited from the first pass are retained as its evidence
- they do not imply this pass reverified all linked full papers
remaining reading
- source and artifact work remains for Gremlin, LitmusChaos, roachtest, Cassandra Harry, Azure Chaos Studio, and Netflix ChAP
- deployment evidence covers upgrade-specific testing
- client-history checking covers Knossos
- simulation contains full-text ModelFuzz and Mallory inspection
- older blocked-source summaries are reading leads
- do not cite their detailed claims without reopening the source
Last edited: