distributed storage and databases (authored by agents unless marked đ§)
start here
- recommendation: begin with existing artifacts and a falsification experiment
- choose a research direction only after checking its closest prior work
- the priorities below are agent opinions
- first candidate: test the boundary between a transaction protocol and its storage engine during recovery
- transactions and regions supplies the MongoDB 2025 comparison
- verification boundaries supplies the proof-assumption angle
- proposed contribution: explain which executable integration tests a formal contract can provide
- stop if MongoDBâs existing model-generated tests already do the same thing
- deferred candidate: regional data movement under conflicting transactions
- SPANStore and SkyPIE mechanism checks remain blocked
- this is retired from active project selection
- compare PolyBaseâs row reassignment with Bonspielâs geographic concurrency control
- measure movement cost and slow responses during shifting demand before inventing a policy
- stop if the combined policy adds no benefit at equal load and guarantees
- third candidate: garbage collection and recovery in object-backed storage
- stores and recovery examines versioned data, deletion, repair, and durability
- select one concrete race between publishing a new version and reclaiming the old version
- stop if the selected design already states and enforces the needed invariant
- fourth candidate: recoverable table snapshots during regional failover
- tables on object stores distinguishes individual copied objects from a complete table version
- compare existing catalog recovery and replication barriers before proposing a new mechanism
- agent storage is another branch of this study
- LLMs and storage is the predecessorâs broad survey
- its strongest candidate requires an independent correctness test of a released implementation
- treat its unverified paper and artifact claims as leads rather than established facts
- see its audit for corrections and limits
deeper second-pass studies, 7 Oct 2026
- each adds to the first-pass files below and ends with ranked research candidates
- these second-pass workers obtained no ChatGPT opinion at the time
- the coordinator completed a later Extra High consultation below
- combined shortlist and remaining work reconciles these studies with the earlier candidates
- distributed transactions
- â55 sources, 14 read in full
- consistency guarantees
- â45 sources, mostly abstracts and introductions
- 8 Oct follow-up inspected VeriStrong and Isolde definitions, arguments, and implementation descriptions
- key-value stores and storage engines
- â75 sources, none read end to end
- verified storage
- â25 sources, 2 read in full
- data that spans regions
- partial: â40 sources, cut short by the Claude usage limit
- 8 Oct follow-up inspected SkyStore, Macaron, Skyplane, and Akkio mechanisms
- a complete 2022â2026 conference sweep remains undone
- partial: a follow-up on object-table recovery, generated file systems, and checkpoint-dependent deletion
- tables on object stores adds existing retention mechanisms to the recovery baseline
- stores and recovery adds SquirrelFS, SysSpec, and Lakestream comparisons
- a full conference sweep remains undone
reading map
- transactions and data across regions
- isolation, commit protocols, deterministic ordering, clocks, geographic placement
- 18 primary papers spanning foundations and 2023â2025 additions
- key-value, file, and object stores
- architecture, redundancy, repair, crash recovery, object-backed databases
- foundational papers and recent FAST work
- storage correctness across client, server, and disk
- RIFL, IronFleet, Perennial, GoJournal, DaisyNFS, Grove, PoWER, PILOT
- distinguishes the proofâs contract from deployment assumptions
- tables built on object stores
- Delta Lake, Iceberg, S3 consistency and asynchronous regional replication
- snapshot publication, retry records, retention, and recoverable destination versions
- LLMs and storage systems
- agent branches and transactions, generated storage code, model-state storage
- inherited work with a separate follow-up audit
- related slices
- consensus and replication
- finding distributed bugs
- other distributed systems areas
- boundaries were assigned by agents
- relevant overlap is retained
what these notes establish
- primary evidence identifies close comparisons and narrows plausible questions
- experiments are proposals
- no performance, correctness, or proof result was produced here
- a paperâs measured result belongs to its workload and assumptions
- ânot found in this reviewâ does not mean ânobody has done itâ
- detailed notes separate quoted evidence, interpretation, and proposed work
coverage and limitations
- added transaction, regional, storage/recovery, and verification literature beyond the inherited LLM survey
- inspected primary PDFs and proceedings
- individual notes record their read depth and search scope
- search failures prevented an exhaustive search through October 2026
- the requested Opus low and Fable medium workers were unavailable in this sessionâs agent interface
- used the available inherited model for parallel literature workers and review
- earlier ChatGPT Extra High consultation attempts failed
- initial attempt selected Extra High but failed before submission
- retry successfully verified Extra High and submitted the prompt
- the helper then returned account_ui_retry_required with no assistant answer
- the later successful consultation is recorded below
- recommendations were instead checked by independent context-free reviewers
- unresolved
- full citation-chain search for each shortlisted direction
- artifact buildability and reproducibility
- novelty and practical benefit of each proposed combination
review outcome
- independent review corrected definitions, baseline assumptions, and metric interpretation
- RIFL client reliability is now explicit
- MongoDBâs existing checkpoint and rollback extensions are baseline work
- DBA-Bench Safe Pass measures safe task completion
- SpecFS artifact availability is separate from buildability
- process lesson: keep failure promises beside each proposed test
- this prevents testing an excluded failure as if it violated a guarantee
recommended next step
- reproduce one MongoDB storage-contract test
- the artifact README documents test generation and a separate WiredTiger build
- buildability remains untested
- compare its existing action model with the specific recovery sequence we want to study
- if a meaningful omission remains, write the smallest additional executable test
- otherwise move to regional-placement measurements
follow-up review, 8 Oct 2026
- reviewed the five second-pass studies and the combined shortlist against the original research goal
- material corrections applied
- separated promised guarantees from tests that deliberately break assumptions
- allowed an inconclusive result from a sound incomplete history checker
- qualified Rust memory-safety and unsupported absence-of-prior-work claims
- corrected links to the systems verification study
- required a compatible fault injector before scheduling persistent-memory tests
- review assessment
- literature breadth supports a broad reading map
- full-paper reading, artifact checks, and novelty assessment remain incomplete
- publication does not complete every proposed research study
- source of assessment
- independent reviewer record: /tmp/cx_storage_review.md
- original human goal: âdo extensive literature review of each one!â
- original human presentation requirement: âResults should be in a neat doc tree in my notesâ
- validation
- local links in the owned folder resolved
- HTML-only mdbook validation passed
- the new shortlist rendered with its navigation entry
- full local build failed because Lua was unavailable
- exact renderer error: â/usr/bin/env: âluaâ: No such file or directoryâ
- concrete blockers
- search tool returned HTTP 404
- direct retrieval of known primary documents remained possible
- temporary ChatGPT retry returned picker_effort_not_verified
- the support-provided saved route then returned picker_deadline selecting gpt-6.1-sol
- those failed attempts obtained no assistant answer
- a later saved-route retry completed
- see the successful consultation below
- search tool returned HTTP 404
- remaining scope
- see combined shortlist for incomplete literature and experiment prerequisites
- no artifact experiment or new verification result was produced
follow-up validation
- a second independent reviewer checked nine primary sources
- the new mechanism descriptions and quoted passages matched their sources
- clarified persistent-memory emulation and Macaron-TTL
- review record:
/tmp/cx_storage_followup_review.md
- the four named regional-paper holes are closed
- this is evidence of progress rather than completion of all citation chains
- full citation-chain review remains incomplete
- Extra High consultation is now complete
successful Extra High consultation, 8 Oct 2026
- ChatGPT, GPT-6.1 Sol, Extra High, saved conversation
- opinion: âchecker â replicated snapshots â proof/simulator mapping â combined storeâ
- context: ranks cheap falsification of four supplied candidates
- it did not see the notes directory or run experiments
- this ranking excludes the MongoDB reproduction candidate
- primary-source checks changed the proposals
- machine-checked isolation characterization is already provided in Rocq
- table-aware replication and validated destination commits already exist
- durable serializable transactions and testing unverified storage bindings already exist
- see verified storage
- accepted research discipline
- establish exact history and failure assumptions before writing proofs or finding bugs
- compare against the corrected existing mechanism before asserting a contribution
- remaining scope
- citation-chain completeness, checker proof bodies, catalog recovery predecessors, and artifact feasibility
- the consultation and publication do not establish novelty or experimental results
consultation review outcome
- independent review checked six primary documents supporting the consultation corrections
- no material factual error was found
- review record:
/tmp/cx_storage_consult_review.md
- next useful work is to choose one precise contract and inspect its existing implementation
- a broad claim about new snapshot tracking or new weak-isolation theory is no longer supported
final assessment against the human goal
- independent reviewer: âthe storage literature survey can now close with disclosed access and coverage limitsâ
- scope: this storage slice only
- source:
/tmp/cx_storage_goal_final.md
- completed bounded comparisons
- Plume theory/code inventory and Viper artifact boundaries
- Ferrite crash-model/execution connection
- GoTxn write/flush, incomplete commitment, and recovery contract
- source-access blockers handled through deferral
- Viper paper theorem remains unread
- SPANStore and SkyPIE mechanisms remain unread
- dependent project selection is deferred rather than presented as a new research gap
- future research
- artifact builds, experiments, full proof audits, and exhaustive citation closure
- none is claimed as a result of this survey
- publication evidence
- repository notes and storage navigation were pushed through isolated worktrees
- HTML-only build and publication-tree link checks passed
- the consultation-correction deployment succeeded
- fetched the live shortlist with HTTP 200 and its S3 Tables correction present
- deployment of this final assessment must be confirmed separately
Last edited: