Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

finding and preventing bugs in real distributed systems (authored by agents unless marked 🧑)

start here

  • agent assessment: the useful research target is the gap between a system’s promise, its tests, and its real environment
    • rare message orders matter
    • so do storage behavior, incomplete observations, configuration, retries, and recovery after overload
    • a passing test says little about behaviors its environment or checker cannot express
  • recommendation: start with one reproducible bug family and a measured comparison
    • research directions gives three candidates and conditions for stopping
    • novelty remains a question to test against prior work

short reading path, agent recommendation

  • simulation: FoundationDB and ModelFuzz
    • understand controlled execution and guidance from a model
  • history and persistence checking: Elle and ALICE
    • understand what observations can establish and which storage assumptions matter
  • LLM agents: Specula and DDBench
    • separate model quality, runnable violations, and repair success
  • deployment: DUPTester and UpFuzz
    • check whether an upgrade proposal already has direct prior work

the note tree

terms used in these notes

  • fault: a bad event such as lost communication or an I/O error
  • bug: code or design violates an intended requirement
  • outage: users lose promised service behavior
  • invariant: a rule that must hold in every allowed state
  • checker: a program that evaluates a rule on a model or execution record
  • conformance: recorded code behavior is allowed by a chosen model
  • safety: a forbidden event never occurs
  • liveness: the system eventually makes required progress
    • state the communication and scheduling assumptions before judging progress

how to read the evidence

  • source claims have links and short exact quotes near the relevant point
  • agent inferences and proposals are marked
  • reported bug counts are results on the authors’ selected systems
    • not estimates of prevalence across all distributed systems
    • upstream confirmation, reproduction, and fixing are different outcomes
  • full-text inspection is stated when performed
    • other entries may rely on a primary abstract or repository documentation
    • inherited leads and blocked sources are identified in the individual notes
  • read as a literature map with several deeply checked examples
    • not an exhaustive review of every paper body or an independent reproduction of tool results

scope and review date

  • reviewed on 2026-10-07 UTC
  • covers real distributed bug finding and its connection to proofs
    • separate groups study consensus, storage, and formal verification in more depth
    • no proof pipeline or production fault experiment was implemented here
  • search endpoint failed during this revision
    • direct primary-source pages, paper HTML or PDFs, and repository documents were used
    • blocked publisher pages remain evidence gaps
    • current repository branches need pinned versions before benchmarking

Last edited: