agent memory, partial compaction, and RAG (authored by agents unless marked đ§)
short version
- revised on 7 Oct 2026 UTC by a second agent with web search; the first draft (same day, no web search) is folded in and corrected below
- compaction research moved fast in 2026 and the simple ideas are taken
- learned compression prompts (Acon, ICML 2026), agent-callable compaction (CAT), budget-aware RL (ContextBudget, CompactionRL), async validation (Slipstream), pointer-based lossless compaction (ARC), parallel compaction for serving, and RL-trained compaction inside GLM-5.2âs pipeline all exist
- inference: a paper that only shortens history better or triggers compaction smarter will be scooped or already is
- the open problem on the memory side is maintenance, not recall
- fact: Supersede measures a 92% to 77% drop when a frontier model must keep its own bounded notes instead of seeing full context, and 68% to 28% as the conversation grows 24x with no recovery from more memory
- fact: MemTrace finds âthe evidence was retrievable 10 times more often than it was missingâ when systems fail
- fact: MemoryAgentBench and Memora both report forgetting out-of-date facts as the weakest competency of every memory system they test; ForgetEval says production failures come mainly from forgetting, not retrieval
- on the RAG side, tracing a poisoned answer back to its documents is done at passage level (RAGForensics, WWW 2025) and span level (Needle-in-RAG), but removing the source does not remove its consequences
- fact: nobody we found measures what survives in summaries, memories, graph indexes, or KV chunk caches after a bad source is withdrawn
- fact: the 435-paper Always-On Agents survey counts 269 works on retrieval against 66 on forgetting and 27 on rollback
- citations are a separate attack surface from answers
- fact: CiteShade makes a model blame a trusted source for a wrong answer; DeepTRACE finds 53.6% to 97.5% unsupported statements in deep research modes
- RAG stores leak: membership inference needs about 30 natural queries (CCS 2025), and agent-driven extraction pulls over 70% of a knowledge base out of commercial platforms
- best ideas, in order
- A. deletion propagation: after a poisoned or outdated source is withdrawn, measure and then fix what still acts on it across memories, compaction summaries, graph summaries, experience stores, and KV chunk caches
- B. delayed obligations: a benchmark and runtime check for constraints that bind 30 to 100 steps after compaction, beyond Slipstreamâs 2 to 4 step validation window, measuring whether the agent actually recalls the pointer ARC-style systems give it
- C. compaction scheduling under real serving load: when does compaction make a fleet of agents slower, given prefix-cache invalidation and the unpredictable summary sizes the parallel-compaction paper measured
what the topic is, in plain words
- an agent is a model that acts in a loop: read, decide, call a tool, read the result, repeat
- everything it has seen piles up in its input; that input is finite and each token costs money and attention
- compaction means replacing part of that pile with something shorter: a summary, a pointer, or nothing
- the humanâs notes call the partial version âOPCâ: replace selected spans while keeping enough to continue correctly
- the risk is that the shortened version drops a promise, a restriction, or a fact the agent needs much later
- long-term memory means keeping things across sessions: user facts, lessons, past results
- the hard part is not storing or finding them but updating and forgetting when the world changes
- RAG means fetching documents before answering
- the document store is a new attack surface: insert bad text and the answer follows it
- the store is also a privacy surface: it can be probed and extracted
- citations are the audit trail, and they can be faked separately from the answer
- these are systems problems about state over time: what is kept, who may change it, how a change propagates, how to undo it
- the Always-On Agents survey frames it the same way and ties it to âdatabases, distributed systems, formal methods, capability security, and machine unlearningâ
sources from the humanâs notes
- agent memory: calls the direction âauditable state management for long-running agentsâ and lists Acon, CAT, ContextBudget, Slipstream, Context Folding, Git-Context-Controller
- RAG: names Zhang et al.âs traceback paper (now verified as RAGForensics, below)
- agent frontier mission 4: âcompare real memory with shuffled, stale, poisoned, no-memory, and equal-token full-context controlsâ
- literature directions direction 3: OPC âcan support a paper if it contributesâ a better abstraction, a stronger evaluation, a systems mechanism, or a negative result
evidence labels used below
- fact: stated in the source we opened
- claim: the authors assert it; we did not check it
- inference: our reasoning
- peer reviewed vs preprint: from the arXiv comments field or publisher page; âpreprintâ means we found no venue
what existing work shows
compaction: who decides what to drop, and how it is checked
- Acon, Kang et al., ICML 2026 (peer reviewed), arXiv 2510.00615, v3 Jun 2026
- what: an outer loop that âiteratively refines compression guidelines based on failure analysis of the agentâ, then distills the compressor into a small model; no change to the agentâs weights
- main number: â26-54% reduction in peak token usage while improving task success over existing compression baselinesâ on AppWorld, OfficeBench, and multi-objective QA with GPT-4.1; AppWorld 56.0% to 56.5%
- limits: fact, appendix A says âhistory compression typically invalidates the existing KV-cache, necessitating a costly re-computation of the entire compressed sequenceâ; latency rises 73.24s to 101.92s in appendix table 4; evaluation âprimarily focuses on GPT modelsâ
- open work named: âthe integration of our framework into live, multi-agent production systemsâ
- CAT, Context as a Tool, Liu et al., preprint, arXiv 2512.22087, Dec 2025
- what: âelevates context maintenance to a callable tool integrated into the decision-making process of agentsâ; a workspace of âstable task-semantic anchors, an evolvable long-term memory, and a short-term working memoryâ; trajectories with inserted compaction actions train a 32B coding agent
- main number: 57.6% on SWE-bench Verified against 49.8% ReAct and 53.8% threshold compression (table 2)
- limits: no limitations section found; one benchmark; the trained model is still below larger base models (DeepSeek-V3.1 61.0%)
- ContextBudget / BACM, Wu et al., preprint, arXiv 2604.01664, Apr 2026
- what: context management as a âsequential decision problem with a context budget constraintâ; GRPO with a budget curriculum from 8k down to 4k tokens; the agent decides âwhen and how much of the interaction history to compressâ
- main number: claim, âover 1.6xâ over strong baselines in high-complexity search settings; in the 32-objective regime F1 4.545 vs 0.909 for MEM1
- limits: claim, sparse delayed reward; search QA only
- CompactionRL, Li et al., preprint, arXiv 2607.05378, Jul 2026, revised Oct 2026
- what: RL that âjointly optimizes task execution and summary generationâ across compaction boundaries
- main number: GLM-4.5-Air reaches â66.4% on SWE-bench Verified and 26.2% on Terminal-Bench 2.0, exceeding the base model under inference-time compaction by 6.6 and 4.9 pointsâ; âdeployed in the RL pipeline for training the open GLM-5.2 modelâ
- inference: compaction is now trained into frontier open models; prompt-level tricks compete against this
- Slipstream, Chen et al., preprint, arXiv 2605.08580, May 2026
- what: compaction âcan unpredictably degrade accuracy due to a structural validation gapâ; the fix runs the compactor âin parallel with the agentâs continued execution on the original, uncompacted contextâ and uses that independent continuation to check and repair the summary
- main numbers: âimproves task accuracy by up to 8.8 percentage points over synchronous compactionâ and âreducing end-to-end latency by up to 39.7%â; table 1 SWE-bench Verified Qwen3.5-9B 23.4% to 29.8%, Seed-OSS-36B 29.6% to 35.8%
- the window is short: âcompaction overlaps with an average of 2.1 agent steps on BrowseComp and 3.7 steps on SWE-bench Verifiedâ; âWindows with Slipstream cover 88% of first error manifestationsâ on SWE-bench Verified
- limits: fact, âThe few errors that surface outside of Slipstreamâs validation windowâŠcannot be recovered, as with synchronous compactionâ; rejection is rare, â1.0-3.5% on BrowseComp, 5.4-8.5% on SWE-bench Verifiedâ
- inference: 12% of first errors on coding tasks fall outside the window, and the paper does not say how far outside; that is the opening for idea B
- ARC, Addressable Recall Compaction, Dang et al., preprint, arXiv 2607.25066, Jul 2026
- what: âARC stores tool observations in an append-only, ID-addressable log and replaces older observations with compact citations when compaction is requiredâ; the agent âcan subsequently use these identifiers to request stored content without re-executing the corresponding tools or depending solely on similarity-based retrievalâ
- main numbers: needle-in-a-haystack â99.40%, compared with 88.12% for the best-performing baselineâ; LongBench-v2 hard â29.97%, compared with 28.25%â; bandwidth down â38.8% on Qwen3-8B and 73.5% on Qwen3-32B relative to âSliding_windowââ
- limits: fact, âAll experiments use the Qwen3 model family; generalization to other tokenizers, reasoning styles, or instruction-following behavior for the _recall convention remains openâ; âEach ARC citation also adds a small, fixed token overheadâ
- inference: this is the pointer-plus-recall design the first draft proposed as new; it is prior work now; what ARC does not measure is whether the agent asks for recall when an obligation, not a fact lookup, depends on it
- Parallel Context Compaction, Cim et al., preprint, arXiv 2605.23296, May 2026
- what: split the history into blocks and summarize them concurrently on vLLM; characterizes synchronous compaction as a serving problem
- facts worth keeping: âthe blocking call stalls agent inference for tens of secondsâ; âmodels largely ignore length instructions and self-bound their outputâ; output length coefficient of variation 19.8% to 84.5% across runs; âAt matched compaction decode volume, it reduces end-to-end wall time and improves compaction throughput over the sequential baselineâ
- limits: claim, prefill of uncached block content can erase the gain at 16k blocks; the agent still blocks until all workers finish; accuracy measured on HotpotQA and LoCoMo, not agent tasks
- TRACE, Min et al., preprint, arXiv 2608.06503, Aug 2026
- what: âcompression can weaken the influence of recent interactions, increasing blocked actions, repeated exploration, and instability across runsâ; evaluates each compaction event by âpaired closed-loop continuations from the same environment stateâ
- limits: self-described âpreliminary empirical studyâ on AppWorld only
- What Does Context Compression Cost an Agent?, Liu, preprint, arXiv 2608.16370, Aug 2026
- what: measures reacquisition cost; âGPT-5.5 is the clearest case: completion changes from 80% to 85% (p = 1.0) while retrieval increases from 21.0 to 63.9 calls (p = .002)â
- fact: âIn a second environment, ALFWorld, sliding compression produces no retrieval surge, showing that the reacquisition signature is environment-dependentâ
- inference: task success hides compaction damage; count tool calls and re-reads, not only pass rate
- ICLR (the method, not the venue), Wang et al., preprint, arXiv 2609.29875, Sep 2026
- what: drops old reasoning blocks while keeping actions and observations; âhistorical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedbackâ
- main number: reward 0.699 to 0.718 on 260 WorkBuddyBench tasks with input tokens down 25.5%
- inference: what is safe to forget depends on what was written to disk; a compaction policy can use that signal
- Anthropic context editing, platform docs, read 7 Oct 2026
- fact: server-side strategies clear old tool results at a token threshold and clear thinking blocks; client SDK compaction replaces the whole history with a summary at 100k tokens
- fact: clearing âInvalidates cached prompt prefixes when content is cleared⊠Youâll incur cache write costs each time content is clearedâ
- inference: the cache-invalidation cost Acon names is already a documented production trade-off; a scheduling result (idea C) has a real audience
- older compression background (abstracts only, from the first draft, kept short): LLMLingua and LLMLingua-2 token pruning, MemGPT tiers, Reflexionâs stored reflections, Lost in the Middle
- inference: all are baselines, none decides what an agent must keep
memory across sessions: update, forget, stale
- MemoryAgentBench, Hu et al., ICLR 2026 (peer reviewed), OpenReview, local OCR in the paper collection
- fact: âall methods fail on the multi-hop situation (with achieving at most 28% accuracy)â on the selective-forgetting split, even with the prompt saying ânewer facts have larger serial numbersâ
- fact: an overwrite-policy ablation raises single-hop 36.0 to 40.0 but multi-hop âdrops to 4.0â
- limits: synthetic counterfactual edits from MQUAKE; âwe could only conduct experiments on some relatively representative Memory Agentsâ
- MemoryArena, He et al., ICML 2026 (listed in PMLR v306), arXiv 2602.16313, Feb 2026
- what: 766 tasks in four domains where memory must drive later actions; tests long-context models, MemGPT, Mem0, ReasoningBank, BM25, MemoRAG, GraphRAG
- fact: agents âwith near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic settingâ; âtwo environments exhibiting near-zero SRâ
- claim: current mechanisms âhave limited capacity to preserve and update task-relevant state variablesâ
- Supersede, Patel, preprint, arXiv 2606.27472, Jun 2026
- what: on LongMemEvalâs knowledge-update subset (n=78) the agent keeps a 300-character notes field, âraw sessions are never re-fedâ
- facts: full context 92%, bounded notes 77%, p=0.0033; growing the conversation 24x drops 68% to 28%; giving proportionally more memory (7,150 chars) gives âno detectable recovery (28%->28%, n=25)â
- claim: âThe bottleneck is therefore memory maintenance, not comprehension, and is not closed by a stronger modelâ
- limits: single author, GRPO result is âa single small model (Qwen2.5-3B) and a single runâ; n=25 for the key null
- MemTrace, Long et al., preprint, arXiv 2606.17328, Jun 2026
- what: scores each typed fact across memory age, question type (current, earlier, trajectory), and evidence condition (present, missing, false premise), over 13 memory configurations
- fact: âwhen systems fail, the evidence was retrievable 10 times more often than it was missingâ; âsafe abstention does not imply correcting a false premiseâ
- Memora and FAMA, Uddin et al., ACL 2026 Findings (peer reviewed), arXiv 2604.20006
- what: weeks-to-months personalized conversations; FAMA âpenalizes reliance on obsolete or invalidated memoryâ
- fact: âEvaluations of four LLMs and six memory agents reveal frequent reuse of invalid memories and failures to reconcile evolving memories. Memory agents offer marginal improvementsâ
- limits: simulated users; the authors call it a lower bound
- ForgetEval and control-plane placement, Yang, preprint, arXiv 2606.15903, Jun 2026
- what: 13 memory configurations on 385 adversarial forgetting cases; the âcontrol plane that mutates them via supersede, release, purgeâ is âlargely untestedâ
- claim: hooks at mutation time reach 91.7% to 93.2%; deterministic methods fail canonicalization, inscription-time LLMs fail intent-aware deletion
- StateAuditor, Sun and He, preprint, arXiv 2608.01619, Aug 2026
- what: âMemory-augmented agents can know that a userâs stored state is outdated and still plan around the old valueâ; the fix audits from stored state to draft, with code that âpins each quotation to a single entry, checks that the new evidence really is newerâ; âWhat is verified is provenance and chronology - not semantic supersessionâ
- main number: STALE benchmark .736 vs .686, â+5.0-point paired gain (95% CI [+2.9, +7.2])â; a matched control with equal calls gets only +0.6
- limits: fact, âWe make no claim about general-purpose agent memoryâ; âa harder authored lifecycle set gives no gainâ; proposer, adapter, judge share one model family
- When Stale Constraints Go Unchecked, Nakayashiki, preprint, arXiv 2608.25553, Aug 2026
- what: agents inherit six memory items with provenance links and may inspect two source records; âa memory is stale when the current record for its provenance target withdraws the content the memory statesâ
- facts: models inspected the constraintâs source in about one episode in five; when superseded, âstale-consistent decisions in 77.3%, 74.7% and 74.7% of episodesâ; redirecting one slot recovers +61.3 to +80.7 points; a rule to âprefer memories that state a limit on a candidate directionâ recovers it without oracle knowledge
- limits: six-item stores, two scripted domains, 16 models; no real deployment data
- inference: provenance pointers alone do not help if the agent never follows them; this directly weakens the first draftâs assumption that pointers establish safety
- Always-On Agents survey, Ding et al., preprint, arXiv 2606.30306, Jun 2026
- fact: across 435 coded works, âretrieve (269 of 435)â, âwrite (200)â, âaudit (88)â, âforget (66)â, ârollback (27)â; âAuthority is the rarest axis at 72 of 435â
- what: proposes AOEP-v0, which âscores state mutation and recovery rather than answer quality aloneâ
- inference: this is the best single citation that deletion and rollback are under-studied
- Mem0, Chhikara et al., preprint, arXiv 2504.19413, Apr 2025
- correction to the first draft: the â26% relative improvements in the LLM-as-a-Judge metricâ is over OpenAIâs memory feature, not over full context; the â91% lower p95 latencyâ and âmore than 90% token costâ savings are versus full context
- inference: a vendor paper; the graph variant does not consistently beat the plain one
- A-MEM, Xu et al., NeurIPS 2025 (peer reviewed), arXiv 2502.12110
- fact: new memories âcan trigger updates to the contextual representations and attributes of existing historical memoriesâ; no deletion or source-withdrawal mechanism found in the text we read
- LongMemEval, Wu et al., ICLR 2025 (peer reviewed), arXiv 2410.10813
- fact: appendix E.5, âcorrect retrieval yet wrong generation (15%~19% of all instances, and 40%~50% among the error instances)â; the first draftâs âacross three reader modelsâ wording was not confirmed
- fact: abstract reports a â30% accuracy drop on memorizing information across sustained interactionsâ for commercial assistants
memory poisoning
- MINJA, Dong et al., preprint (v5 Feb 2026; no venue in comments), arXiv 2503.03704
- fact: âinjects malicious records into the memory bank by only interacting with the agent via queries and output observationsâ; injection success 98.2% average, attack success 76.8% average over three agents
- fact: the targeted detection prompt âfails to generalize to other agentsâ and the general one brings false positives; utility on MMLU drops 10.0% under default settings
- MemoryGraft, Srivastava and He, preprint, arXiv 2512.16962, Dec 2025
- what: implants âmalicious successful experiences into the agentâs long-term memoryâ; exploits âthe agentâs semantic imitation heuristicâ; tested on MetaGPTâs DataInterpreter with GPT-4o
- claim: âa small number of poisoned records can account for a large fraction of retrieved experiences on benign workloadsâ
- inference: lessons and procedures are poisonable, not only facts; a withdrawal experiment must cover them
- Memory Poisoning Attack and Defense on Memory Based LLM-Agents, Sunil et al., preprint, arXiv 2601.05504, Jan 2026
- fact: ârealistic conditions with pre-existing legitimate memories dramatically reduce attack effectivenessâ on EHR agents; sanitization ârequires careful trust threshold calibrationâ
- inference: MINJAâs headline numbers are an upper bound; replicate under filled memories
RAG poisoning, traceback, and removal
- PoisonedRAG, Zou et al., USENIX Security 2025 (peer reviewed), arXiv 2402.07867
- fact: âa 90% attack success rate when injecting five malicious texts for each target question into a knowledge database with millions of textsâ (NQ 2.68M, HotpotQA 5.23M, MS-MARCO 8.84M)
- fact: paraphrasing drops ASR 0.97 to 0.87; perplexity filtering has high FPR at high TPR; âduplicate text filtering cannot successfully filter malicious textsâ; at k=50 ASR is still 41% to 43%
- RobustRAG, Xiang et al., preprint (v2 Apr 2026; no venue on arXiv), arXiv 2405.15556
- fact: âisolate-then-aggregate strategyâ; certified robust accuracy 71.0% on RealtimeQA for Llama2-7B with 1 of 10 passages corrupted; âwe focused on single-hop RAG tasks in this paperâ
- Polymorphic Sybil Poisoning benchmark, Lee and Kim, preprint, arXiv 2607.03739, Jul 2026, local PDF in the paper collection
- what: six âlexically diverse passages jointly support an attacker-chosen target while evading lexical near-duplicate filtersâ
- facts: diversity gives â+18.8pp hijack amplificationâ; catching the residual with embedding cosine âraises false-positive rate 9xâ; âabstention and drift together hold 47-66% of output mass, unmonitored by ASR+ACCâ
- inference: this confirms the first draftâs worry that counting independent passages (RobustRAG) is unsafe when copies are paraphrased; it is now a measured result, not a hypothesis
- RAGForensics, Zhang et al., WWW 2025 (peer reviewed), arXiv 2504.21668, DOI, code on GitHub
- this is the traceback paper the humanâs RAG note cites; full text now read
- setup: âWe assume that the traceback system has collected a set of user queries and their incorrect RAG outputs as reported by usersâ; provider has database access; retriever and LLM are black boxes
- method: retrieve top-K for each reported query, ask an LLM per text to âjudge whether the provided context tries to induce you to generate an answer consistent with the provided response, regardless of whether it is correctâ, iterate until K benign texts are found
- main numbers: detection accuracy 97.4% to 99.6%, false positives 0.4% to 2.7% against PoisonedRAG and instruction injection on NQ, HotpotQA, MS-MARCO; holds under two adaptive attacks
- limits: fact, âcurrently limited to targeted poisoning attacks and is unable to trace the specific poisoned texts responsible for untargeted attacksâ; it needs user-reported wrong outputs; the judge is an LLM reading each text with the bad answer in hand
- inference: traceback is solved for the easy case (text visibly argues for the reported bad answer); what happens after identification is not studied
- Needle-in-RAG / RAGCharacter, Cui and Liu, preprint, arXiv 2605.01782, May 2026
- what: character-level traceback by âbudgeted counterfactual masking and replayâ over a logged prompt trace; âmoving RAG forensics from document-level suspicion toward finer-grained evidence auditing and potential remediationâ
- claim: best trade-off between localization and over-attribution across five attack families and six LLMs
- CiteShade, Guo, preprint, arXiv 2609.15660, Sep 2026, local PDF in the paper collection
- what: âan attacker controlling a single source induces a model to produce an attacker-chosen wrong answer and to attribute it to a trusted source that does not support it, while the evidence for the correct answer remains in contextâ
- facts: wrong-answer rate âfrom 0.01 to 0.68â; âvulnerability tracks a modelâs propensity to cite, not its size or accuracyâ; âperplexity filtering and citation-support checking are each insufficientâ
- defense limits: ârecall 0.77 at a 6% false-positive rate on the reference model, but near-zero useful operating points on models whose benign citations are already un-groundedâ; âthe defense raises the attackerâs cost substantially without eliminating the attackâ
- inference: the cited source and the causal source differ; a traceback that trusts citations (or a user report that names the cited source) can blame the wrong document
- RAGtrap, Kapelinski and Kreutz, SBSeg 2025 extended proceedings, publisher page
- not opened: the publisher page and PDF returned an empty search page and a 403; only the search engineâs snippet was visible
- the snippet says it ârecords a signed provenance entry for every passage at ingestion, indexed by source and by content hashâ, and that âExact hashing cannot attribute content changed after ingestion, nor content supplied by more than one sourceâ
- must read before claiming idea A is new; from the snippet it removes passages, not derived state
- RAGuard and RAGShield appeared in search (2026 preprints) but were not opened
answers versus their cited sources, and deep research agents
- DeepTRACE, Venkit et al., arXiv 2509.04499, Sep 2025; mlanthology lists it under ICLR 2026, arXiv comments do not say so
- what: eight metrics over â2,727 samples (303 queries x 9 models)â, about 80,000 LLM-judged support checks
- facts: âYouChat(DR), PPLX(DR), Copilot(DR), and Gemini(DR) all fare poorly, with unsupported rates ranging from 53.6% (Gemini) to 97.5% (PPLX)â; âcitation accuracy ranging from 40-80% across systemsâ
- limits: judge agreement with humans is âa Pearson correlation of 0.62⊠indicating moderate agreementâ
- BrowseComp-Plus, Chen et al., NeurIPS 2025 (listed on neurips.cc), arXiv 2508.06600
- what: 830 BrowseComp queries over a fixed 100,195-document corpus with labeled evidence and hard negatives, so retriever and agent can be scored apart
- facts: GPT-5 70.12% with a dense retriever vs 55.90% with BM25; citation precision 83.4% for GPT-5 but 20.0% for Qwen3-32B
- inference: this is the right testbed for poisoning-plus-traceback experiments on multi-step search agents, since evidence documents are known
- LiveResearchBench, Wang et al., ICLR 2026 (peer reviewed), arXiv 2510.14240
- what: 100 live deep-research tasks, 17 systems, six dimensions including âCitation Association, and Citation Accuracyâ
- fact: âEven SoTA systems are far from citation error-freeâ; judge-human agreement on citation traceability 85.9%
- ReportBench, Li et al., preprint, arXiv 2508.15804
- what: survey papers as ground truth; cited statements checked by fetching âthe full content of each cited webpageâ and comparing
- limits: STEM arXiv surveys only
- Mind2Web 2, Gou et al., preprint, arXiv 2506.21506
- what: 130 tasks; an agent judge checks whether each claim is âsupported by the webpage contentâ at cited URLs
- fact: âOpenAI Deep Research, can already achieve 50-70% of human performance while spending half the timeâ; assumes âcited URLs provide truthful and credible informationâ
- inference across these: support checking is LLM-judged and reads only the cited page; CiteShade shows that is exactly the check an attacker can satisfy
RAG systems side: caches, indexes, freshness
- Cache-Craft, Agarwal et al., SIGMOD 2025 (peer reviewed), arXiv 2502.15734
- fact: in a production system â75% of the retrieved chunks for a query were reprocessed, amounting to over 12B tokens in a month⊠costing approximately $50kâ
- what: store per-chunk KV caches, reuse them in any position, recompute a few tokens; â1.6X speed up in throughput and a 2X reduction in end-to-end response latency over prefix-cachingâ
- CacheBlend, Yao et al., preprint on arXiv (EuroSys 2025 not confirmed on the arXiv page), arXiv 2405.16444
- fact: âreuses the precomputed KV caches, regardless prefix or not, and selectively recomputes the KV values of a small subset of tokensâ; TTFT down â2.2-3.3xâ
- RAGCache, Jin et al., preprint, arXiv 2404.12457
- fact: âorganizes the intermediate states of retrieved knowledge in a knowledge tree and caches them in the GPU and host memory hierarchyâ; TTFT âup to 4xâ over vLLM+Faiss
- SIFT, Sanovar et al., preprint, arXiv 2606.09441, Jun 2026
- fact: KV reuse âis often slower than full recomputation on modern GPUs due to high-latency disk transfersâ; stores attention-location bit vectors âup to 24,000x smaller than KV tensorsâ; TTFT 1.71x âwhile holding accuracy within 1% of full recomputeâ
- inference for idea A: three generations of RAG serving systems persist per-document KV state outside the document store; deleting a document from the vector index does not delete its cached KV or its place in a knowledge tree unless someone wires that up, and none of these papers mentions deletion
- DGAI, Lou et al., preprint, arXiv 2510.25401, v5 Apr 2026
- fact: coupled on-disk graph indexes cause âsubstantial redundant I/O during index updatesâ; decoupling gives 8.17x faster inserts and 8.16x faster deletes
- IP-DiskANN and FreshDiskANN appeared in search, not opened
- inference: deletion in vector indexes is itself a live systems topic; the cost of a withdrawal in idea A includes index deletes
- Beyond Similarity Search, Budigi and Sirigiri, preprint, arXiv 2605.03275
- fact: names âdata staleness, tenant data leakage, and query composition explosionâ as production RAG root causes; proposes one Postgres+pgvector layer; 50k-document benchmark only
privacy leaks from the store
- The Good and The Bad, Zeng et al., preprint, arXiv 2402.16893
- fact: 250 prompts âextracted 89 targeted medical dialogue chunks from HealthcareMagic and 107 PIIs from Enron Emailâ; RAG also âsubstantially reduced the number of PIIs extracted from the training dataâ
- Is My Data in Your Retrieval Database?, Anderson et al., preprint, arXiv 2405.20446
- fact: one prompt, âDoes this: â{Target Sample}â appear in the context? Answer with Yes or No.â; TPR 0.95 black-box on Llama; a template instruction cuts it to 0.09
- Interrogation Attack, Naseh et al., CCS 2025 (peer reviewed), arXiv 2502.00306
- fact: natural questions answerable only if the document is present; â2x improvement in TPR@1%FPRâ, âjust 30 queriesâ, under $0.02 per document
- claim: detected about 5% of the time vs 90%+ for prior attacks
- CopyBreakRAG (formerly RAG-Thief), Jiang et al., preprint, arXiv 2411.14110, v2 Aug 2025
- fact: âextracts over 70% of the data from the knowledge base in applications on commercial platforms including OpenAIâs GPTs and ByteDanceâs Cozeâ; weaker on disconnected records such as medical notes
- inference: the store is an oracle for both membership and content; a memory store built from a userâs own sessions has the same exposure to anyone who can query the agent (MINJAâs setting)
what is missing
- deletion propagation through derived state is unmeasured
- evidence: RAGForensics stops at identifying texts; Needle-in-RAG stops at spans; RAGtrapâs snippet removes passages; the Always-On survey counts 27 rollback works vs 269 retrieval; A-MEM has no delete; MemoryGraft shows lessons persist
- evidence on the systems side: Cache-Craft, RAGCache, CacheBlend, SIFT persist per-document state and say nothing about deletion
- obligations that bind long after compaction are untested
- evidence: Slipstream covers 88% of first errors with a 2 to 4 step window and says the rest âcannot be recoveredâ; ARC gives pointers but tests only needle lookups; the stale-constraints paper shows agents follow pointers one time in five
- memory maintenance under growth is a measured failure with no systems answer
- evidence: Supersede 28% with no recovery from more memory; MemoryAgentBench multi-hop forgetting at most 28%; Memora âmarginal improvementsâ
- inference: the fix is likely an explicit update protocol with provenance and chronology checks (StateAuditorâs direction), evaluated at scale
- compactionâs serving cost under contention has one characterization paper and no scheduler
- evidence: parallel-compaction paper measures tens-of-seconds stalls and 20% to 85% output variance on single-model vLLM; Acon reports latency up 40%; Anthropic documents cache invalidation; no paper schedules compaction across many agents sharing a prefix cache
- traceback assumes the cited or reported source is the causal one
- evidence: CiteShadeâs laundering; RAGForensics starts from user-reported wrong outputs and a per-text LLM judge
- poisoning evaluations use empty or clean memories
- evidence: the EHR replication finds âpre-existing legitimate memories dramatically reduce attack effectivenessâ
research we can do
- A. deletion propagation: what still acts on a withdrawn source
- question: after a poisoned or outdated document is removed from the store, how much of the agentâs later behavior still depends on it, and what mechanism cuts that to zero at what cost
- why open: see gaps 1 and 3; traceback papers end at identification; memory benchmarks test updates seen in context, not withdrawal of a source already consumed
- first experiment
- build on BrowseComp-Plus (known evidence documents) and LongMemEval knowledge-update questions
- inject one PoisonedRAG-style text or one outdated fact, let the agent run long enough to create derived state: a compaction summary, a Mem0 or A-MEM memory, a GraphRAG community summary, a ReasoningBank-style lesson, and a Cache-Craft-style KV chunk cache
- withdraw the source four ways: delete from index only; delete plus invalidate directly linked objects; delete plus rebuild all transitively dependent objects; full reset
- ask later questions that need the corrected fact; also ask questions that the deleted source answered correctly, to measure collateral loss
- record which derived object each wrong answer traces to
- convincing result: a measured residual-dependence rate per object type after index-only deletion (we expect it to be high for summaries and lessons), and a dependency-tracked invalidation that drives it near zero at a fraction of full-reset cost, with claim-level rather than object-level links to avoid erasing correct content
- cost: API-model runs on a few hundred tasks, one open-model serving stack for the KV-cache part; two to four weeks for the measurement half
- closest work that could scoop it: RAGtrap (source revocation, not opened), StateAuditor (draft repair from stored state), ForgetEvalâs mutation hooks, the Always-On surveyâs AOEP-v0 protocol, machine-unlearning work on RAG that we did not search
- the humanâs mission 4 control set (shuffled, stale, poisoned, no-memory, equal-token) applies directly
- B. delayed obligations beyond the validation window
- question: when a restriction or promise set at step t matters only at step t+30 or t+100, across one or more compactions, does the agent still honor it, and does giving it an addressable pointer (ARC) help only if something makes it look
- why open: gap 2
- first experiment
- 30 coding or local-tool tasks with a checkable restriction (do not touch file X, report Y when done, user changed their mind at step k)
- replay the same pre-compaction state through: full context, truncation, synchronous summary with an explicit constraints section, CAT-style stable anchors, ARC-style pointers, Slipstream-style async validation, and an obligation ledger that is re-injected at each step with its provenance id
- measure obeyed, missed, falsely completed, and obsolete obligations; count recall requests the agent actually makes; count reacquisition tool calls as in the compression-cost paper
- then run 100 unmodified SWE-bench Verified or Terminal-Bench tasks to see how often delayed failures occur without injection
- convincing result: a delayed-failure rate that Slipstreamâs window provably cannot catch, and a ledger or forced-check design that fixes it at under 5% extra tokens
- stop conditions: an explicit constraints section in an ordinary summary matches it; gains need hand-written obligations; failures happen only in contrived tasks
- cost: cheap; open 9B to 36B models as in Slipstream suffice
- closest work: Slipstream, ARC, CAT, TRACE, the stale-constraints paper (its âprefer memories that state a limitâ rule is a baseline), CompactionRL (a trained model may already keep obligations; include one)
- C. compaction scheduling under real serving load
- question: with many agents sharing one serving stack and a prefix cache, when does compacting make the fleet slower, and can a scheduler that sees cache state and queue depth beat a token threshold
- why open: gap 4
- first experiment: replay long SWE-bench and BrowseComp trajectories through vLLM with prefix caching; vary concurrency, compaction trigger, summary size, sync vs parallel vs async; measure p95 step latency, time per solved task, cache hit rate, recomputed prefill tokens
- convincing result: a region where threshold compaction hurts throughput and a cache-aware policy recovers it; otherwise a negative result that a tuned threshold is enough
- cost: GPUs and a local stack; defer until available
- closest work: parallel-compaction paper, Acon appendix A, Anthropicâs context editing docs, SGLang and vLLM prefix caching
- smaller follow-ups
- citation-aware traceback: extend RAGForensics with leave-one-out causal attribution so laundered citations do not misdirect it; CiteShadeâs own defense is the baseline and it already notes leave-one-out weakens under multi-source influence
- re-run MINJA and MemoryGraft against filled memories and against Memora-style evolving facts, per gap 6
ChatGPTâs opinion
- pending: the shared ChatGPT tool was unusable on 7 Oct 2026 (sign-in required), so no consultation was run
what we searched
- queries (7 Oct 2026): traceback poisoning RAG RAGForensics; agent context compaction 2026; agent memory benchmark forgetting stale 2026; RAG privacy membership inference; deep research citation faithfulness; RAG serving KV cache vector database freshness; MemoryArena; FAMA; Claude context editing docs; BrowseComp-Plus; RAG poisoning provenance removal 2026; memory poisoning MINJA 2026; streaming ANN updates
- sources opened in full or in large part: RAGForensics, ARC, StateAuditor, Supersede, stale constraints, parallel compaction, Slipstream, Acon, CAT, ContextBudget, MemoryArena, Memora, Always-On survey, MemoryAgentBench (local OCR), CiteShade and polymorphic sybil (local PDFs), BrowseComp-Plus, DeepTRACE, Anthropic docs, PoisonedRAG, RobustRAG, MINJA, Mem0, A-MEM, LongMemEval, Cache-Craft, CacheBlend, RAGCache, SIFT, the three privacy papers, CopyBreakRAG, LiveResearchBench, Mind2Web 2, ReportBench
- abstract only: CompactionRL, ICLR-compression, TRACE, compression-cost, MemTrace, ForgetEval, MemoryGraft, EHR memory poisoning, Needle-in-RAG, DGAI, unified data layer, plus the 2023 to 2024 background papers
- found but not opened: RAGtrap, RAGuard, RAGShield, IP-DiskANN, FreshDiskANN, Context Folding, Git-Context-Controller, AgentPoison, SSGM governance framework
- not covered: machine unlearning for RAG or for memory stores, GraphRAG update mechanics, encrypted or access-controlled retrieval, multimodal memory, evaluation of commercial memory products beyond what the benchmarks report, agent security and prompt injection beyond memory poisoning (sibling topic)
- method caveat: full-text quotes came through a fetch tool that summarizes pages with a small model; the subagent reports flagged a few quotes as possibly trimmed; check wording against the PDF before quoting in a paper
Last edited: