Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

distributed storage and databases (authored by agents unless marked 🧑)

start here

  • recommendation: begin with existing artifacts and a falsification experiment
    • choose a research direction only after checking its closest prior work
    • the priorities below are agent opinions
  • first candidate: test the boundary between a transaction protocol and its storage engine during recovery
    • transactions and regions supplies the MongoDB 2025 comparison
    • verification boundaries supplies the proof-assumption angle
    • proposed contribution: explain which executable integration tests a formal contract can provide
    • stop if MongoDB’s existing model-generated tests already do the same thing
  • deferred candidate: regional data movement under conflicting transactions
    • SPANStore and SkyPIE mechanism checks remain blocked
    • this is retired from active project selection
    • compare PolyBase’s row reassignment with Bonspiel’s geographic concurrency control
    • measure movement cost and slow responses during shifting demand before inventing a policy
    • stop if the combined policy adds no benefit at equal load and guarantees
  • third candidate: garbage collection and recovery in object-backed storage
    • stores and recovery examines versioned data, deletion, repair, and durability
    • select one concrete race between publishing a new version and reclaiming the old version
    • stop if the selected design already states and enforces the needed invariant
  • fourth candidate: recoverable table snapshots during regional failover
    • tables on object stores distinguishes individual copied objects from a complete table version
    • compare existing catalog recovery and replication barriers before proposing a new mechanism
  • agent storage is another branch of this study
    • LLMs and storage is the predecessor’s broad survey
    • its strongest candidate requires an independent correctness test of a released implementation
    • treat its unverified paper and artifact claims as leads rather than established facts
    • see its audit for corrections and limits

deeper second-pass studies, 7 Oct 2026

  • each adds to the first-pass files below and ends with ranked research candidates
    • these second-pass workers obtained no ChatGPT opinion at the time
    • the coordinator completed a later Extra High consultation below
    • combined shortlist and remaining work reconciles these studies with the earlier candidates
  • distributed transactions
    • ≈55 sources, 14 read in full
  • consistency guarantees
    • ≈45 sources, mostly abstracts and introductions
    • 8 Oct follow-up inspected VeriStrong and Isolde definitions, arguments, and implementation descriptions
  • key-value stores and storage engines
    • ≈75 sources, none read end to end
  • verified storage
    • ≈25 sources, 2 read in full
  • data that spans regions
    • partial: ≈40 sources, cut short by the Claude usage limit
    • 8 Oct follow-up inspected SkyStore, Macaron, Skyplane, and Akkio mechanisms
    • a complete 2022–2026 conference sweep remains undone
  • partial: a follow-up on object-table recovery, generated file systems, and checkpoint-dependent deletion

reading map

what these notes establish

  • primary evidence identifies close comparisons and narrows plausible questions
  • experiments are proposals
    • no performance, correctness, or proof result was produced here
  • a paper’s measured result belongs to its workload and assumptions
  • “not found in this review” does not mean “nobody has done it”
  • detailed notes separate quoted evidence, interpretation, and proposed work

coverage and limitations

  • added transaction, regional, storage/recovery, and verification literature beyond the inherited LLM survey
  • inspected primary PDFs and proceedings
    • individual notes record their read depth and search scope
    • search failures prevented an exhaustive search through October 2026
  • the requested Opus low and Fable medium workers were unavailable in this session’s agent interface
    • used the available inherited model for parallel literature workers and review
  • earlier ChatGPT Extra High consultation attempts failed
    • initial attempt selected Extra High but failed before submission
    • retry successfully verified Extra High and submitted the prompt
    • the helper then returned account_ui_retry_required with no assistant answer
    • the later successful consultation is recorded below
    • recommendations were instead checked by independent context-free reviewers
  • unresolved
    • full citation-chain search for each shortlisted direction
    • artifact buildability and reproducibility
    • novelty and practical benefit of each proposed combination

review outcome

  • independent review corrected definitions, baseline assumptions, and metric interpretation
    • RIFL client reliability is now explicit
    • MongoDB’s existing checkpoint and rollback extensions are baseline work
    • DBA-Bench Safe Pass measures safe task completion
    • SpecFS artifact availability is separate from buildability
  • process lesson: keep failure promises beside each proposed test
    • this prevents testing an excluded failure as if it violated a guarantee

recommended next step

  • reproduce one MongoDB storage-contract test
    • the artifact README documents test generation and a separate WiredTiger build
    • buildability remains untested
    • compare its existing action model with the specific recovery sequence we want to study
    • if a meaningful omission remains, write the smallest additional executable test
    • otherwise move to regional-placement measurements

follow-up review, 8 Oct 2026

  • reviewed the five second-pass studies and the combined shortlist against the original research goal
  • material corrections applied
    • separated promised guarantees from tests that deliberately break assumptions
    • allowed an inconclusive result from a sound incomplete history checker
    • qualified Rust memory-safety and unsupported absence-of-prior-work claims
    • corrected links to the systems verification study
    • required a compatible fault injector before scheduling persistent-memory tests
  • review assessment
    • literature breadth supports a broad reading map
    • full-paper reading, artifact checks, and novelty assessment remain incomplete
    • publication does not complete every proposed research study
  • source of assessment
    • independent reviewer record: /tmp/cx_storage_review.md
    • original human goal: “do extensive literature review of each one!”
    • original human presentation requirement: “Results should be in a neat doc tree in my notes”
  • validation
    • local links in the owned folder resolved
    • HTML-only mdbook validation passed
      • the new shortlist rendered with its navigation entry
    • full local build failed because Lua was unavailable
      • exact renderer error: “/usr/bin/env: ‘lua’: No such file or directory”
  • concrete blockers
    • search tool returned HTTP 404
      • direct retrieval of known primary documents remained possible
    • temporary ChatGPT retry returned picker_effort_not_verified
    • the support-provided saved route then returned picker_deadline selecting gpt-6.1-sol
      • those failed attempts obtained no assistant answer
      • a later saved-route retry completed
        • see the successful consultation below
  • remaining scope
    • see combined shortlist for incomplete literature and experiment prerequisites
    • no artifact experiment or new verification result was produced

follow-up validation

  • a second independent reviewer checked nine primary sources
    • the new mechanism descriptions and quoted passages matched their sources
    • clarified persistent-memory emulation and Macaron-TTL
    • review record: /tmp/cx_storage_followup_review.md
  • the four named regional-paper holes are closed
    • this is evidence of progress rather than completion of all citation chains
  • full citation-chain review remains incomplete
    • Extra High consultation is now complete

successful Extra High consultation, 8 Oct 2026

  • ChatGPT, GPT-6.1 Sol, Extra High, saved conversation
    • opinion: “checker → replicated snapshots → proof/simulator mapping → combined store”
    • context: ranks cheap falsification of four supplied candidates
      • it did not see the notes directory or run experiments
      • this ranking excludes the MongoDB reproduction candidate
  • primary-source checks changed the proposals
    • machine-checked isolation characterization is already provided in Rocq
    • table-aware replication and validated destination commits already exist
    • durable serializable transactions and testing unverified storage bindings already exist
  • accepted research discipline
    • establish exact history and failure assumptions before writing proofs or finding bugs
    • compare against the corrected existing mechanism before asserting a contribution
  • remaining scope
    • citation-chain completeness, checker proof bodies, catalog recovery predecessors, and artifact feasibility
    • the consultation and publication do not establish novelty or experimental results

consultation review outcome

  • independent review checked six primary documents supporting the consultation corrections
    • no material factual error was found
    • review record: /tmp/cx_storage_consult_review.md
  • next useful work is to choose one precise contract and inspect its existing implementation
    • a broad claim about new snapshot tracking or new weak-isolation theory is no longer supported

final assessment against the human goal

  • independent reviewer: “the storage literature survey can now close with disclosed access and coverage limits”
    • scope: this storage slice only
    • source: /tmp/cx_storage_goal_final.md
  • completed bounded comparisons
    • Plume theory/code inventory and Viper artifact boundaries
    • Ferrite crash-model/execution connection
    • GoTxn write/flush, incomplete commitment, and recovery contract
  • source-access blockers handled through deferral
    • Viper paper theorem remains unread
    • SPANStore and SkyPIE mechanisms remain unread
    • dependent project selection is deferred rather than presented as a new research gap
  • future research
    • artifact builds, experiments, full proof audits, and exhaustive citation closure
    • none is claimed as a result of this survey
  • publication evidence
    • repository notes and storage navigation were pushed through isolated worktrees
    • HTML-only build and publication-tree link checks passed
    • the consultation-correction deployment succeeded
    • fetched the live shortlist with HTTP 200 and its S3 Tables correction present
    • deployment of this final assessment must be confirmed separately

Last edited: