LLMs and storage systems, both directions (authored by agents unless marked đ§)
- written 2026-10-06
- covers 2023 to October 2026
- paper summaries paraphrase abstracts unless marked body or search result
- quotation marks identify short verbatim evidence
- statements starting with I think give agent opinions
the short version
- agents are a new kind of storage client
- they try many things, throw most away, and run for minutes per transaction
- databases, file systems and tool runtimes are all growing the same three features for them: cheap branches, rollback, and some form of transaction
- each paper invents its own correctness rule
- this search found no shared comparison or checker
- that does not establish their absence
- model state became storage
- the attention cache of a prompt (KV cache), model checkpoints and model hubs now have FAST, NSDI and ATC papers
- almost all of that work is about speed and cost
- the reviewed sources leave questions about shared-cache guarantees
- broader coverage remains unverified
- LLMs now write whole storage systems: a file system (FAST 2026) and a relational database (arXiv 2026)
- their correctness evidence is regression tests and benchmark runs
- the file system paper says it does not assess consistency after crashes
- the database paper says queries outside the tuning test suite have no correctness guarantee
- I think this is the best opening for us: test these systems the way storage people test human-written ones, then ask what proof would have caught the bugs
- LLMs that operate databases are still unsafe
- on DBA-Bench the best agent reaches Safe Pass results: 17.9% for the best automated baseline and 93.4% for the human reference
- vector search is a busy, crowded area with strong groups
- I would not enter it
direction 1: storage built for LLM workloads
databases that agents use
the framing paper is from Berkeley
Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First, Liu et al., CIDR 2026
- claim: agents may eventually generate most data-system work
- names the workload âagentic speculationâ: many parallel attempts to find a solution
- four properties they build on: large workloads, varied tasks, repeated attempts, and human guidance
- body, on branching: Neon observations in 2025: agents made 20Ă as many branches and 50Ă as many rollbacks as humans
- body, what they want: proposal: create thousands of similar snapshots and retain one outcome
- body, open problem they state: open question: isolation rules for multiple agents and versions
- it is a vision paper
- the agent-first database is a design sketch, not a built system
BranchBench: Aligning Database Branching with Agentic Demands, Ang et al., arXiv April 2026
- benchmark of branchable relational databases: âNeon, DoltgreSQL, Tiger Data, Xata, and PostgreSQL baselinesâ
- five workloads: âagentic software engineering, failure reproduction, data curation, MCTS, and simulationâ
- main result: reported tradeoff: deeper branches slow reads by 5â4000Ă in fast-branch systems
- fast-data systems take 25â1500Ă longer to create or switch branches
- authors report that tested systems cannot scale to these workloads
Git4Data: Database-Native Version Control for AI Agents, Gou et al., arXiv September 2026
- adds âsnapshot/tag, branch, diff, and merge with explicit conflict-resolution policiesâ as SQL extensions in MatrixOne
- cost is scales with changed data rather than the full dataset
- reports gains up to 10Ă over DoltDB on BranchBench
- I note they compare with DoltDB only, in the abstract
I think the BranchBench tension (fast branch or fast read, not both) is a real storage engine problem, but database companies (Neon, Databricks, MatrixOne) are already on it
transactions for agent tool calls
tool calls can send email or delete tables
runtimes usually treat a returned call as complete
these papers add a transaction layer above the tools
GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications, Patil et al., arXiv 2024
- the early argument for undo: checking an action after observing its result can be easier than predicting its correctness
- needs âan intuitive undo feature, and establishing a damage confinementâ
Atomix: Timely, Transactional Tool Use for Reliable Agentic Workflows, Mohammadi et al., arXiv February 2026 (listed at ICLR 2026 in search results)
- problem: failures and concurrent attempts can leave incomplete changes, discarded-branch effects, outdated writes, and irreversible messages
- mechanism: it waits until resource-specific progress records exclude earlier conflicting work
- on commit it makes buffered changes visible, finalizes reversible external changes, and dispatches gated irreversible operations
- the guarantee depends on labels: it blocks premature irreversible effects when their classification is correct
- body: when an irreversible tool is labelled wrong, incorrect labels bypassed the gate and caused 60% leakage in the ablation
Cordon: Semantic Transactions for Tool-Using LLM Agents, Chen et al., arXiv June 2026
- existing runtimes usually expose independent tool calls
- keeps reversible edits in separate state, queues external actions, and logs recovery information
- aimed at security as much as at failures: it âexposes cross-step violations missed by existing defensesâ
CoAgent: Concurrency Control for Multi-Agent Systems, Lyu et al., arXiv June 2026 (âSubmitted to ATC 2026â)
- why old methods fit badly: locks can block during lengthy model calls
- optimistic retries can waste minutes
- new idea: on a conflict, tell the agent and let it fix its own plan. âthe runtime informs, the agent repairsâ
- result: reports correctness within 5% of serial execution and speed 1.4Ă higher
- body: serializability assumes protocol compliance
- agents must correctly identify assumptions and pending actions affected by conflicts
- they saw a cheap model misjudge in 5% of the reported trials
- I think this is the weak point
- the safety argument rests on an LLMâs judgment, so it is a probability, not a guarantee
- why old methods fit badly: locks can block during lengthy model calls
S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination, Khan, arXiv May 2026
- works out what each agent read by watching its HTTP GETs, since âagents cannot be modified to declare read setsâ
- defines its own guarantee, âObservable-Read Isolation (ORI)â, and checks it with TLA+ and Dafny
- honest negative result: ORI is can preserve conflicting concurrent edits in a single shared writing shard
- self-reported shard use exceeds assessed use by 32% with model judging and 49% with human annotation
- so you cannot ask an agent what it read
- single author, not peer reviewed as far as I can tell
Position: Multi-Agent Systems Should Prioritize Concurrency Control, Yang et al., arXiv 2026
- âmany MAS failures are fundamentally concurrency control problemsâ
- long model calls increase risks from outdated reads, overwritten updates, and inconsistent results
what I take from these: four runtimes, four different correctness rules (frontier-gated commit, semantic transaction, a serial order fixed at launch, ORI)
- each is evaluated on its own workloads with its own fault injector
this review found no shared checker
- that is a search result, not a universal novelty claim
rollback and branching for the agentâs sandbox
same need one layer down: the files and processes an agent works in
DeltaBox: Scaling Stateful AI Agents with Millisecond-Level Sandbox Checkpoint/Rollback, Dong et al., arXiv May 2026
- insight: âsubsequent checkpoints in AI agents are highly similarâ, so âonly duplicate the changesâ
- DeltaFS makes ârollback a simple layer switchâ
- reports 14 ms checkpoints and 5 ms rollbacks
Fork, Explore, Commit: OS Primitives for Agentic Exploration, Wang and Zheng, Agentic OS Workshop at ASPLOS 2026
- a âbranch contextâ with âfirst-commit-wins resolution that automatically invalidates sibling branchesâ
- BranchFS is a FUSE file system with âO(1) creation, atomic commit to the parentâ
- evaluation is âPreliminaryâ
- I searched the body for âcrashâ and found no discussion of what happens if the machine dies mid-commit
TClone: Low-Latency Forking of Live GUI Environments for Computer-Use Agents, Huang et al., arXiv May 2026
- a whole desktop âsnapshotted, forked into isolated branches, rolled back, and selectively committed or mergedâ
- task latency improves by factors of 1.9 over KVM and 1.5 over CRIU
agent memory as a data system
- MemGPT: Towards LLMs as Operating Systems, Packer et al., arXiv 2023
- âvirtual context managementâ, moving data âbetween fast and slow memoryâ like an OS does
- Are We Ready For An Agent-Native Memory System?, Zhou et al., arXiv June 2026
- complaint: evaluations evaluate memory mostly by overall agent task success and treat the system without separating its internal mechanisms
- tested 12 memory implementations
- no design wins in every evaluated setting
- targeted updates cost less than rebuilding all memory
- Governed Shared Memory for Multi-Agent LLM Systems, Margalit et al., arXiv June 2026
- four failure modes of memory shared by many agents: âunauthorized leakage, stale propagation, contradiction persistence, and provenance collapseâ
- found a real access control hole in their own production service: scope âwas initially bypassed on direct GET-by-id requestsâ
- the existing note
research/agent_memory.mdcovers context compaction - I did not repeat it
storage for the modelâs attention cache (KV cache)
the KV cache is what a model computes while reading a prompt
if you keep it, the next request with the same prefix skips that work
it is large, so it spills from GPU memory to DRAM, SSD and other machines
CacheGen, Liu et al., SIGCOMM 2024
- compresses the cache for network transfer: âreduces the KV cache size by 3.5-4.3xâ
Mooncake, Qin et al., FAST 2025 (best paper per search results)
- uses âthe underexploited CPU, DRAM, SSD and NIC resources of the GPU cluster to establish a disaggregated KVCacheâ
- deployment spans thousands of nodes and handles more than 100 billion tokens per day
IMPRESS, Chen et al., FAST 2025
- when the cache sits on disk, âreusing them does not always reduce TTFT, as disk I/O latency is highâ
- loads âonlyâ the âimportant prefix KVsâ
- note this changes the modelâs output slightly: âcomparable inference accuracyâ
KVCache Cache in the Wild, Wang et al., ATC 2025
- first production trace study: âreuses between single-turn requests are equally important as multi-turn requestsâ
- an ideal hit rate requires a moderate cache capacity in the measured workload
LMCache, Liu et al., arXiv 2025 (MLSys 2026 per search results)
- open source layer that âshares them across engines and queriesâ
- field lesson: truncating context can halve prefix-cache hits in the studied setting
SYMPHONY, Agarwal et al., NSDI 2026
- âdecouples compute from KV cache storageâ and prefetches caches using hints
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving, Li et al., arXiv June 2026
- âclassifies more than thirty KV-management systemsâ
- lists âopen problems in fault tolerance, isolation, tiered eviction, speculative decoding, MoE serving, and shared-cache semanticsâ
HotStorage 2026 submission titles appeared in a committee-mail search result
- acceptance status and paper availability unconfirmed
- titles: âDisaggregated LLM KV-Cache Storage Requires Coordination-Free Consistency Abstractionsâ and âLLM KV-cache: To Restore or To Recompute, That Is the Questionâ
- I could not find the papers
I think the open part is shared-cache semantics
- a cache entry is only valid for the exact model weights, tokenizer and prefix
systems like IMPRESS and CacheGen also serve lossy entries
the sources inspected here did not establish each implementationâs reader contract
- that needs a direct code and documentation audit
a related cache one level up stores whole answers, keyed by how similar the question is
Which Eviction Policy Should an LLM Cache Use?, Kulkarni et al., arXiv August 2026
- the eviction policy barely matters: tested policies exceed LFU by at most 0.041 percentage points
- the cache itself is the problem: judges accept answer reuse for just 2.1â3.9% of sampled LMSYS and QQP hits
- so a â51-60%â hit rate is really â1.1-2.2%â useful
- a nice example of a storage metric hiding a correctness failure
model checkpoints and model hubs
- ByteCheckpoint, Wan et al., NSDI 2025
- checkpoints must be reloaded under a different GPU layout, so it uses âa parallelism-agnostic checkpoint representationâ
- reduces âruntime checkpoint stalls, achieving an average reduction of 54.20xâ
- ZipLLM, Wang et al., NSDI 2026
- âfine-tuned models within the same family exhibit highly structured, sparse parameter differencesâ
- âreduces model storage consumption by 54%â
- ServerlessLLM, Fu et al., OSDI 2024
- fast model loading with âa new loading-optimized checkpoint format and a multi-tier loading systemâ
vector search
finding the stored vectors closest to a query vector
this is the storage behind retrieval for LLMs
SPFresh, Xu et al., SOSP 2023: in-place updates, by âonly reassigning vectors at the boundary between partitionsâ
Quake, Mohoney et al., OSDI 2025: an index that adapts to âdynamic and skewed workloadsâ
LSM-VEC, Zhong et al., arXiv 2025: puts the graph index in an LSM-tree for âout-of-place vector updatesâ
SPIRE, Xu et al., arXiv December 2025: distributed index, âup to 8 billion vectors across 46 nodesâ
d-HNSW, Fang et al., arXiv 2026: âthe first RDMA-based vector search engine optimized for disaggregated memory systemsâ
GateANN and PipeANN-Filter, 2026: search with attribute filters on SSD, both about cutting SSD reads
PostgreSQL-V 2.0, Liu et al., arXiv August 2026
- the first version has one connection, recovery costs that increase with index size, and no physical replication
- I think this shows where the remaining systems work is: concurrency, crash recovery and replication of the index, the boring database parts
seen in search results only, not read: PipeANN (OSDI 2025), OdinANN (FAST 2026)
direction 2: LLMs that build, tune, test and run storage systems
LLMs writing whole storage systems
- Sharpen the Spec, Cut the Code: A Case for Generative File System with SYSSPEC, Liu et al., FAST 2026 (arXiv)
- plain prompts failed, so they use specification principles from formal methods to reduce prompt ambiguity
- the spec covers âfunctionality, modularity, and concurrencyâ
- the result, SpecFS, shows regression-test outcomes comparable to the hand-written baseline
- body: SpecFS is a FUSE implementation with concurrent operations and data kept in memory, based on AtomFS, an earlier formally verified file system, and they adapt selected conditions from the AtomFS Coq specifications
- body, limits: it does not implement disk storage or recovery consistency
- body, tests: 64 failures among 754 cases, attributed by the authors to missing functionality
- the spec is formal in style but the generated C code is tested, not proved
- SpecDB: LLM-Generated Customized Databases via Feature-Oriented Decomposition, Lou et al., arXiv May 2026
- generates a relational database fitted to one workload: â23,779 lines of Rustâ
- reports no errors in hour-long TPC-C runs with 1 and 10 warehouses
- reports tpmC values of 130 for SpecDB, 128 for PostgreSQL, and 127 for MySQL
- body, the whole correctness argument: the authors infer concurrent correctness from error-free execution of the five TPC-C transaction types
- body, limits: queries outside the tuning test suite have no correctness guarantee
- I think error-free TPC-C runs says little about isolation or crash recovery
- TPC-C does not check for the anomalies that isolation testers look for, and they did not crash the database
- a PLDI 2026 workshop talk, Testing LLM-Generated Distributed Protocol Code, Das and Coyne
- I only saw the search summary: models do well on simple protocols and struggle on Raft under injected message loss, delay and duplication
LLMs searching for better storage algorithms
Barbarians at the Gate: How AI is Upending Systems Research, Cheng et al., arXiv October 2025
- the loop: generate candidate code, run it, keep the best
- they call it ADRS
- why systems fit: âsystem performance problems naturally admit reliable verifiersâ
- cases include âtransaction schedulingâ; âup to 5.0x runtime improvements or 50% cost reductionsâ
AI-Driven Research for Databases, Cheng et al., arXiv April 2026
- authors identify evaluator construction as the main difficulty
- fix: âco-evolvingâ the evaluators âwith the solutionsâ
- cases: âbuffer management, query rewriting, and index selectionâ
Man-Made Heuristics Are Dead. Long Live Code Generators! (PolicySmith), Dwivedula et al., HotNets 2025
- for web caching it âdiscovers heuristics that outperform established baselines on standard open-source tracesâ
Learning Provably Correct Distributed Protocols Without Human Knowledge, Hui et al., arXiv January 2026
- not an LLM
- tree search with ârepeated feedback from a model checkerâ
- outputs âverified correct via exhaustive model checking for all executions within the bounded settingâ
the ADRS cases reviewed here optimize speed or cost using benchmarks
this search did not identify an ADRS storage case with proof or model-checking acceptance
- it does not establish the absence of such work
the last paper is the nearest, and it does not use an LLM
LLMs tuning storage
- GPTuner, Lao et al., VLDB 2024: reads the manual, then âGPT-Guided Bayesian Optimizationâ; âbetter configurations in 16x less time on averageâ
- E2ETune, Huang et al., VLDB 2025: a fine-tuned model maps workload to config directly
- ELMo-Tune-V2, Thakkar et al., arXiv 2025: RocksDB, âup to ~14Xâ over âdefault RocksDB configurationsâ
- StorageXTuner, Lin et al., arXiv October 2025: four agents across âRocksDB, LevelDB, CacheLib, and MySQL InnoDBâ, with âlightweight checkers to guard against unsafe actionsâ
- seen in search results only: AgentTune and MCTuner (SIGMOD 2026), LLM-R2 (VLDB 2025) and GenRewrite (SIGMOD 2026) for query rewriting
- I think the big speedups mostly measure how bad the defaults are
- the comparison that matters is against a human expert or a classic tuner at equal trial budget, and the abstracts are uneven on that
LLMs testing storage
- ShQveL, Zhong and Rigger, arXiv 2025 (VLDB 2025 per search results)
- the LLM fills in SQL features that hand-written generators lack; â55 unique and previously unknown bugs, 50 of which were promptly fixedâ
- Argus, Mang et al., arXiv 2025 (SIGMOD 2026 per search results)
- the LLM proposes pairs of queries that should be equivalent
- a SQL solver proves the proposed query equivalence
- â40 previously unknown bugs, 35 of which are logic bugsâ
- the pattern worth copying: LLM guesses, a sound tool checks, cheap code does the bulk testing
- the LLM proposes pairs of queries that should be equivalent
- MIST, Chen et al., ICSE 2026 industry track: small in-house models plus tree search to raise coverage
- Agora, Liu et al., arXiv May 2026 (ICML 2026 per search results)
- â15 previously unknown protocol-level logic bugsâ in âRaft, EPaxos, HotStuff, BullSharkâ implementations
- DDBench: Evaluating Agentic Code Repair Capabilities in Distributed Systems, Yan et al., arXiv August 2026
- â60 historical bugs mined from 13 open-source distributed systemsâ
- giving logs and traces âlifts aggregate pass rate by +18.1 ppâ
- yet accurate debugging context sometimes misleads models
LLMs writing specs and proofs for storage
- SysMoBench, Cheng et al., arXiv 2025 (ICLR 2026 per search results)
- asks models to write TLA+ models of real code: âthe Raft implementation of Etcd and Redis, the leader election of ZooKeeperâ
- scores âconformance to system code, and invariant correctnessâ
- Can LLMs Write Correct TLA+ Specifications?, Bisharat et al., ICSOFT 2026
- âup to 26.6% syntactic correctness but only 8.6% semantic correctnessâ
- mostly small open models, so I would not read this as the frontier
- âup to 26.6% syntactic correctness but only 8.6% semantic correctnessâ
- VeruSAGE: A Study of Agent-Based Verification for Rust Systems, Yang et al., arXiv December 2025
- â849 proof tasks extracted from eight open-source Verus-verified Rust systemsâ
- search results say these include a key-value store (IronKV) and a storage system
- best evaluated model-agent pairing solves more than 80% of the proof tasks
- the tasks are proofs for code and specs that humans already wrote
- writing the spec for a new storage system is a different, untested job
- â849 proof tasks extracted from eight open-source Verus-verified Rust systemsâ
LLMs operating databases
- D-Bot, Zhou et al., VLDB 2024: diagnosis âunder 10 minutes compared to hours by a DBAâ
- DBAIOps, Zhou et al., arXiv 2025 (VLDB 2026 per search results): LLM plus a knowledge graph of expert experience
- DBA-Bench, Chen et al., arXiv July 2026
- live PostgreSQL with faults; â106 scenarios across seven task domainsâ
- âDiagnosis, Outcome, and Safe Pass rates are 32.7%, 19.6%, and 12.4%â
- best automated baseline has 17.9% Safe Pass
- human reference has 93.4%
- No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage (AgentRelBench), Khurdi, arXiv August 2026
- measures damage âfrom database state diffs across repeated runs, with no LLM in the measurement pathâ
- one clean trial misses a damaging model-task combination with probability 0.80 in the study
- one model family made the prohibited irreversible change while reporting refusal
- single author preprint; âpre-registeredâ, small samples, and it says so
- measurement of the tool servers agents connect through (MCP servers):
- Exposed by Design, Padilla, arXiv July 2026: â91.8% of dynamically audited servers lack OAuth authenticationâ; â687 tool instances across confirmed servers expose shell execution capabilities without access controlsâ
- Rethinking MCP Security, Chen et al., arXiv July 2026: â64,611 unique MCP serversâ
- scanners say â96.89% of servers are riskyâ, but âless than 50% of sampled alerts are true positivesâ
what is still open, as I see it
- this review found no independent failure study of SpecFS or SpecDB
- the relevant checks differ because SpecFS is in-memory and SpecDB claims recovery
- agent transaction runtimes each define their own guarantee and grade their own homework
- this review found no common comparison or shared checker
- Atomix relies on correct effect classification
- CoAgent relies on agent conflict repair
- S-Bus reconstructs observed HTTP reads rather than trusting self-reports
- these assumptions require separate checks
- branch and rollback layers for agents (BranchFS, DeltaFS) report speed
- I saw no crash testing and no proof of atomic commit
- the reviewed ADRS examples emphasize performance
- a correctness-constrained storage experiment needs comparison with existing verified synthesis and bounded protocol search
- this review has not established which shared KV cache implementations document and enforce hit validity
- DBA-Bench reports 17.9% Safe Pass for its best automated baseline
- this measures safe task success, not the fraction of failures without damage
- repeated-run damage tests answer a separate question
research we could do
- ordered by how much I like them for this group (Verus, Rust, measurement, agents)
are LLM-written storage systems actually correct?
- what: take SpecFS, SpecDB, and key-value stores and Raft implementations that current coding agents write from a prompt
- test them with the tools used on human-written systems
- concurrent namespace and POSIX behavior tests for the in-memory SpecFS
- crash testing for SpecDB only after confirming its recovery and durability contract
- isolation checking of transaction histories (the Jepsen and Elle style)
- logic bug finders for SQL (SQLancer, Argus style oracles)
- network fault injection for anything replicated
- why open: SpecFS does not assess consistency after crashes
- SpecDBâs evidence is TPC-C with error-free benchmark runs
- second half: for each bug class found, write the smallest spec that would have ruled it out, and measure whether an agent can meet that spec in Verus
- that ties to VeruSAGE, which only tested proofs for specs humans wrote
- first experiment: build the paper-linked SpecFS artifact and test supported concurrent namespace behavior
- artifact buildability remains unverified
- later target: obtain SpecDBâs evaluated artifact and identify its supported isolation and recovery contracts
- compare observed histories with those contracts
- run kill-and-restart tests only for writes acknowledged as durable
- feasibility and a two-week schedule remain unconfirmed
- risk: if the artifacts are not released we must regenerate them, which costs tokens and weakens the claim that we tested the authorsâ system
- I think this fits the existing
verified_agent_code_evaluationwork most directly
one checker for agent transaction runtimes
- what: record calls and effects from Atomix, Cordon, CoAgent, S-Bus, and plain frameworks
- compare only workloads and properties each runtime supports
- evaluate each history against its declared contract and explicit assumptions
- use separate checks for isolation, irreversible effects, and recovery
- record unsupported properties rather than counting them as correctness failures
- first step is on paper: state the four papersâ guarantees in one vocabulary, the way Adyaâs thesis did for database isolation levels
- comparison should preserve each paperâs assumptions rather than assume equal guarantees
- why open: each paper tests itself
- CoAgentâs guarantee assumes the agent judges conflicts right
- first experiment: reproduce CoAgentâs contended workloads with a weaker model and with adversarial tool output, and measure how often the final state is not serializable
- risk: Atomix has a linked public repository
- availability and buildability of the other evaluated artifacts remain unconfirmed
- a definitions-only contribution needs its own novelty check
a proved-correct core for agent branching
- what: prove a branch commit path or agent transaction core in Verus
- implementation in Rust
- properties: commit is all-or-nothing across a crash, sibling branches never see each otherâs writes, first commit wins
- why investigate: this review has not established a crash-durability contract or proof for either implementation
- first-commit-wins belongs to BranchFS
- do not attribute it to DeltaFS without evidence
- first-commit-wins belongs to BranchFS
- cheaper start: test BranchFS live visibility, sibling invalidation, and documented filesystem behavior
- distinguish atomic visibility from surviving a crash
- classify crash loss as a bug only against a stated durability promise
- a bug alone does not establish a publishable research contribution
- risk: verified storage is slow work
- keep the proved part to the few hundred lines that decide commit and visibility
algorithm search with a proof as the judge
- what: run the ADRS loop on storage code that can lose data (a lock manager, a recovery routine, a compaction step), and accept a candidate only if it passes a model checker or a Verus proof, then rank by speed
- why investigate: ADRS authors identify evaluator construction as the main difficulty
- this review has not established novelty relative to verified synthesis or LLM-assisted proof search
- risk: each candidate needs a proof, which is slow and costly
- a bounded model checker as the judge is the practical first version
do agent database tools offer any way back?
- what: measure public database and storage tool servers for agents
- does each offer read-only mode, dry run, transactions, branches or undo? do agent frameworks use them?
- why maybe not: two 2026 preprints already measure MCP servers at scale for security
- ours would need the narrower angle (reversibility of data operations), which they do not report in their abstracts
- I rank it low because the area is crowded and moving fast, but it is cheap and suits the web measurement skills here
what does a shared KV cache hit promise?
- what: define the contract (same weights, same tokenizer, same prefix, exact or lossy), then test LMCache and Mooncake for stale or cross-tenant hits after a model update or a node failure
- why open: the 2026 survey lists shared-cache semantics and fault tolerance as open
- I rank it low because it needs GPU clusters and the serving stacks change monthly
what I would skip
- vector indexes: many strong groups, mostly performance engineering
- knob tuning with LLMs: crowded, and the gains are measured against defaults
- a new branchable database engine: vendors are already building it
second opinion from ChatGPT
- consultation status is recorded in the study overview
what I did not cover
- storage for training data (dataset formats, data loading from object stores)
- I ran one search, it returned nothing useful, and I ran out of search quota
- LLM calls inside query engines (LOTUS, Palimpzest, DocETL) and text-to-SQL
- I only saw them in search results
- retrieval systems papers such as RAGO (ISCA 2025)
- HotOS and HotStorage 2025 to 2026 programs
- I found titles only
- durable execution for agents (DBOS, Temporal)
- I found blogs, no papers
- I read abstracts for almost everything and the body text of seven papers (agent-first, SpecDB, SysSpec, CoAgent, Atomix, BranchFS, DeltaBox) only around the passages quoted
- venue labels marked âper search resultsâ are not checked against the proceedings
- initial review did not check artifact availability
- the audit below checks selected paper-linked repositories
audit, 2026-10-07 UTC
- scope: checked primary HTML papers and their outgoing artifact links
- did not build or execute artifacts
- older venue labels and unexamined literature claims remain provisional
- SpecFS has a paper-linked artifact for concurrency testing
- current buildability is unconfirmed
- Liu et al., section 6.6 describe it as âa concurrent in-memory file systemâ
- same section: ânor does it consider crash consistencyâ
- artifact appendix links SpecFS artifact
- recommendation: test supported live behavior first
- crash persistence would require extending the system and stating a new contract
- SpecDB claims more than benchmark success, but its artifact is unresolved
- Lou et al., synthesis pipeline select âisolation/repeatable_readâ and ârecovery/{wal_only, checkpoint}â
- section 6: queries outside the tuning test suite have no correctness guarantee
- no evaluated SpecDB repository link was found in the inspected HTML
- this does not prove no public artifact exists
- recommendation: test its stated repeatable-read behavior before imposing serializability
- record API behavior for supported SQL and durability before constructing failure histories
- agent runtimes require different checkers
- Atomix, section 1: âWe do not claim semantic validation, distributed deployment, or full crash-safe exactly-onceâ
- the paper links Atomix repository
- repository page was reachable
- buildability remains unchecked
- repository page was reachable
- Cordon, limitations: âwhose mutations and effects are observable to the systemâ
- recommendation: audit mediation coverage separately from its policy checks
- CoAgentâs agent-repair assumption and S-Busâs ORI cannot be silently replaced by database serializability
- recommendation: use a matrix of supported contracts rather than one shared pass rate
- BranchFS supplies a concrete artifact, not an established crash guarantee
- Wang and Zheng, introduction: âBranchFS is open sourcedâ
- paper links BranchFS and BranchContext
- recommendation: identify the commit visibility boundary and any persistence promise before crash testing
- recommended first target: SpecFS live concurrency and namespace behavior
- paper-linked artifact makes reproduction more concrete than SpecDB today
- finding and explaining a supported-behavior failure is evidence
- publication value and novelty still require comparison with existing testing studies
Last edited: