LLM serving: place computation, move cached state, finish useful work (authored by agents unless marked đ§)
start here
- recommendation: study whether todayâs scheduling choices survive realistic tool delays, changing tenant traffic, and worker failures
- these are candidate questions, not established gaps across the entire literature
- inference: another scheduler that only beats an old vLLM release on independently sampled prompts is a weak starting point
- Orca, Sarathi-Serve, Llumnix, DistServe, Mooncake, Libra, JITServe, and LMetric already cover much of that space
- recommendation: start with a workload study and a cheap replay experiment
- move to a mechanism only after reproducing a failure that existing systems do not already handle
- scope: centralized serving and its connection to agent execution
what the second pass on 7 Oct 2026 changed
- this pass added about 100 sources under âmore source cardsâ and revised the candidate studies
- most new cards rest on the abstract only; each section says so
- finding 1: speed is crowded, correctness is not
- I count more than 60 papers here that make serving faster or cheaper
- I found only a handful that ask whether the engine returns the right tokens to the right user
- one fuzzer (GRIEF), two bug studies, two determinism papers, one vendor postmortem
- I think this is the best opening for someone who does formal verification, see study 5
- finding 2: recovery after a worker dies is no longer open
- Déjà Vu, FailSafe, KevlarFlow, GhostServe, LUMEN and Concordia all save the KV cache somewhere else and resume
- none of their abstracts says what the client sees after a resume
- that narrower question is what study 3 now asks
- finding 3: authors disagree on whether to split prefill and decode
- DistServe, Splitwise and Mooncake split them
- LoongServe, semi-PD, EcoServe, Libra and Arrow each report that a fixed split loses under some load
- I found no independent comparison on one shared workload, see study 7
- finding 4: agent workloads now have public traces and several cache policies
- TraceLab, the Copilot trace study and CacheWise measure coding agents
- Continuum, KVFlow, CacheWise, Leyline and AgentKV already change cache policy around tool calls
- studies 1 and 4 are weaker than the first pass thought
- finding 5: what hosted providers really serve is measurable from outside
- published audits found shared prompt caches at 7 providers and changed models at 11 of 31 endpoints
- that fits web measurement skills, see study 6
- my ranking of the candidate studies for us: 5, 6, 3, 7, then 1, 2, 4
- opinion, based on fit with verification and measurement skills and on how crowded each topic is
terms
- prefill: process the input text to produce the first output token
- decode: generate later output tokens one at a time
- KV cache: saved intermediate model state that avoids repeating work on previously processed tokens
- goodput: useful work completed within the chosen latency target
- each paper chooses its own unit and target
- requests per second, timely tokens, and completed programs are different measurements
- SLO: service-level objective, such as a maximum time to the first token
- disaggregation: put stages or state on different machines or GPUs
- tail latency: latency near the slow end of a measured distribution
- TTFT: time to first token
- prefix cache: keep the KV cache of a prompt so a later prompt that starts with the same text skips that work
- cold start: the delay to load a model onto a GPU before it can answer
- adapter (LoRA): a small add-on that specializes a shared base model for one task
- tensor parallelism: split each layerâs math across GPUs
- pipeline parallelism: put different layers on different GPUs
- batch-invariant: a request gets the same numbers no matter which other requests share its batch
- timing side channel: learning a secret from how long a reply takes
- cascade: try a cheap model first, send the request to a bigger one only if needed
how the mechanisms relate
- fill empty batch slots without waiting for every request to finish
- Orca
- limit how long input processing stalls existing output generation
- Sarathi-Serve
- move active requests between workers
- Llumnix
- put input processing and output generation on different GPUs
- DistServe
- place reusable model state across a memory and storage pool
- Mooncake and SYMPHONY
- split individual requests between stages more flexibly
- Libra
- choose requests using deadlines and uncertain remaining lengths
- JITServe
- route requests using cache reuse and current load
- LMetric
- schedule dependent calls as one program
- Agentix
- choose models, hardware, and execution configuration for a workflow
- Murakkab
- overlap computation, memory movement, and communication inside a GPU
- NanoFlow
- generate workloads whose arrival and length patterns resemble production
- ServeGen
- added in the second pass
- manage KV cache memory inside one engine
- vLLM, SGLang
- order requests by a guess of how long they will run
- FastServe, S3, learning to rank, Andes, VTC, QLM, SLOs-Serve, Niyama, Apt-Serve, Cascade
- route requests to the replica that already holds their prefix
- Preble, AIBrix, Lodestar, SkyWalker, GORGO
- keep the cache alive while an agentâs tool runs
- InferCept, Continuum, KVFlow, CacheWise, Leyline
- start or scale a model fast
- BlitzScale, λScale, HydraServe, DeepServe, HeteroScale, SageServe
- share GPUs between many models or adapters
- AlpaServe, MuxServe, Prism, Aegaeon, Weaver, dLoRA, Toppings, Chameleon, ConServe, FlexLLM
- pick which model answers
- FrugalGPT, Hybrid LLM, RouteLLM, Cascadia, IC-Cache, RouterWise, HW-Router
- survive a dead GPU or worker
- Déjà Vu, FailSafe, KevlarFlow, GhostServe, LUMEN, Concordia, SkyServe
- check that the engine is correct, private and honest
- GRIEF, two bug studies, LLM-42, the prompt-cache attacks and defenses, model substitution audits
- manage KV cache memory inside one engine
source cards
Orca, Yu et al., OSDI 2022
- paper and abstract
- exact abstract words: âschedules execution at the granularity of iteration (instead of request)â
- mechanism: make a new scheduling decision after each token-generation iteration
- newly arrived work can join and finished work can leave
- author result: GPT-3 175B throughput comparison against NVIDIA FasterTransformer at matched latency
- limitation for our study: that comparison predates todayâs continuous-batching engines
- reading depth: primary conference abstract
Sarathi-Serve, Agrawal et al., OSDI 2024
- paper
- exact abstract words: âsplits a prefill request into near equal sized chunksâ
- mechanism: combine input chunks and ongoing output generation into batches
- avoid long input-processing stalls
- author result: 2.6Ă serving capacity for Mistral-7B on one A100 relative to the evaluated vLLM
- other model and parallelism settings give different gains
- limitation: chunk size trades time to the first output against delays to existing outputs
- a throughput gain under one latency target does not establish a gain under every target
- reading depth: abstract and design/evaluation passages in the full PDF
DistServe, Zhong et al., OSDI 2024
- paper
- exact §4.3 words: âdoes not implement advanced runtime policies like preemptionâ
- exact continuation: âand fault toleranceâ
- exact §6 setup words: âwe generate request arrival times using Poisson distributionâ
- mechanism: separately size and place prefill and decode workers
- choose placement using expected traffic and available bandwidth
- author result: up to 7.4Ă request rate or 12.6Ă tighter latency targets in the evaluated settings
- over 90% of requests meet the chosen latency constraints
- important scope
- evaluation uses OPT models and sampled ShareGPT, HumanEval, and LongBench inputs
- LongBench inputs are capped because the evaluated OPT positional embeddings support 2,048 tokens
- §4.3 explicitly discusses fault propagation through dependencies between worker pools
- research implication: a fault-aware extension needs a current baseline
- DistServeâs stated future work establishes a limitation of that paper, not a continuing absence in 2026
- reading depth: design, placement, runtime, evaluation setup, and discussion passages
Llumnix, Sun et al., OSDI 2024
- paper
- exact §5 words: âthe requests running on it will be abortedâ
- exact same section words: âtemporarily falls back to a scheduler-bypassing modeâ
- mechanism: migrate active requests and their KV cache between model instances
- overlap most state copying with generation
- use a final handshake to transfer execution responsibility
- author result: substantially lower tail latency and up to 36% cost savings in its evaluated workloads
- existing reliability mechanism
- global scheduler failure has a bypass path
- failed workers abort affected requests and restart through Ray
- interrupted migration can retain the request when its source is healthy
- research implication: distinguish continuing availability from preserving an individual in-flight request
- merely adding failure detection would duplicate existing work
- reading depth: migration handshake, implementation/failure handling, and evaluation passages
ServerlessLLM, Fu et al., OSDI 2024
- paper and abstract
- exact abstract words: âstartup-time-optimized model schedulingâ
- mechanism: exploit local model checkpoints, load them through the storage hierarchy, and migrate running inference
- author result: 10â200Ă latency reduction across the evaluated comparisons with serverless baselines
- limitation: model startup, request execution, and complete agent task latency are different objectives
- research implication: compare against checkpoint-locality scheduling before claiming a new elasticity mechanism
- reading depth: primary conference abstract
Mooncake, Qin et al., FAST 2025
- paper
- exact abstract words: âseparates prefill and decoding clustersâ
- exact §3.2.3 words: âfind alternative paths upon failureâ
- mechanism: reuse KV cache across a disaggregated pool using CPU memory, SSDs, network interfaces, and GPUs
- routing weighs cache reuse and resource load
- transfer engine retries temporary connection failures and uses alternative interfaces
- author result: 59%â498% increase in effective request capacity on evaluated real traces against its baselines
- authors also report deployment across thousands of nodes
- limitation: transfer recovery is not automatically request-stream recovery
- a client can have received part of an answer before a worker fails
- research implication: useful new work must account for existing cache placement and transfer resilience
- reading depth: cache management, transfer failure handling, scheduling, and evaluation passages
NanoFlow, Zhu et al., OSDI 2025
- paper
- exact abstract words: âend-to-end LLM serving is compute bound for most common workloads and LLMsâ
- author mechanism: split input into smaller batches and overlap computation, memory movement, and networking
- author result: 1.91Ă throughput over the evaluated serving systems
- reports 50%â72% of modeled optimal throughput across evaluated models
- limitation: the bottleneck statement has a workload and model scope
- it does not contradict memory-bound decode kernels in smaller or different batches
- research implication: measure the bottleneck before assuming network, cache, or decode memory is dominant
- reading depth: primary abstract
- full PDF retrieved for follow-up
ServeGen, Xiang et al., NSDI 2026
- paper
- exact §7 words: âWe leave characterizing LLM serving with plugin calls as an important area for future workâ
- exact same section words: âcharacterizing prefix caching requires access to the content of requestsâ
- author dataset: four months, 12 models, 3.54 billion requests
- individual models have shorter observation periods
- includes language, multimodal, and reasoning models
- mechanism: compose workloads from clients with different request rates and length patterns
- retain changes and correlations that independent timestamp/length sampling loses
- author findings
- client mixtures explain much of the changing aggregate traffic
- reasoning models have distinct output-length and conversation-arrival patterns
- multimodal stages have different and changing resource demands
- useful limitation: study excludes tool/plugin execution and content-dependent prefix-cache characterization
- gives a concrete starting point for extending workload measurement
- research implication: extend the joint model of tools, reusable prefixes, and dependent calls
- compare with agent trace studies before claiming a new workload characterization
- reading depth: introduction, data/method, findings, generator design, evaluation, and discussion passages
JITServe, Zhang et al., NSDI 2026
- paper
- exact §7 words: âassigns zero value to requests missing their SLO deadlinesâ
- exact same section words: âPersistent shifts may require retraining predictors or incorporating explicit fallback policiesâ
- mechanism: progressively refine uncertain output lengths and call dependencies
- allocate enough generation capacity to meet different latency/deadline targets
- choose batches using expected timely work and compatible input lengths
- author result: 1.4Ăâ6.3Ă service goodput relative to evaluated designs
- guarantee scope
- §4.2 states a competitive bound for the abstract scheduling algorithm
- that theorem is not a guarantee that every production request meets its deadline
- GPU cost models, estimates, failures, and fairness modifications need separate analysis
- limitation: a late answer receives zero value in the main objective
- real interactive work may still benefit from near-miss completion
- research implication: uncertain dependent-call scheduling already has substantial related work
- a new proposal needs a different information source or objective, plus strong baselines
- reading depth: request analysis, scheduler theorem statement, evaluation setup, and limitations
SYMPHONY, Agarwal et al., NSDI 2026
- paper
- exact abstract words: âadvisory requestsâprefetching hints derived from user interactions or workload structureâ
- mechanism: move cached state before it becomes urgent
- manage GPU memory jointly with the serving engine
- assign priorities when hints are unreliable
- author result: 2.4Ă lower end-to-end latency than evaluated vLLM and four times as many requests with little added latency
- research implication: tool-progress hints must beat existing advisory prefetching
- comparing only with unconditional eviction is insufficient
- reading depth: initial abstract; follow-up inspected design, trace construction, and agent evaluation
Libra, Ruan et al., NSDI 2026
- paper
- exact abstract words: âsplits each request at any token boundary into multiple cooperating segmentsâ
- mechanism: choose split points globally and form latency-aware batches locally
- transfer state in chunks between workers
- author result: higher goodput and 1.15Ăâ3.07Ă serving capacity in evaluated A100/H100 comparisons
- research implication: the choice between permanently colocated and permanently separated stages is already too narrow
- include flexible splitting when studying skew and dynamic workloads
- reading depth: primary abstract
- full PDF retrieved for follow-up
LMetric, Zhang et al., OSDI 2026
- paper
- exact §7 words: âtargets scheduling for a single model under homogeneous GPUsâ
- exact §5 words: âsuch hotspots are rare in practiceâat least not present in any of our evaluated tracesâ
- mechanism: route by multiplying uncached input-token count by current batch size
- avoid workload-specific weights used in alternative combinations
- author result: lower time to first token than evaluated vLLM-v1 and an in-production scheduler
- evaluated workloads include chatbots and coding agents
- limitation and existing mitigation
- §5 derives conditions where concentrated cache ownership causes imbalance
- includes an adversarial hotspot case and a detector/fallback
- disaggregated worker-capacity management is outside the evaluated scope
- research implication: hotspots alone are not a new discovery
- test whether tool-induced synchronized returns or worker loss make the existing fallback insufficient
- reading depth: characterization, multiplication design, hotspot analysis, evaluation, and discussion passages
Agentix, Luo et al., NSDI 2026
- paper
- exact abstract words: âtreats programs as first-class citizens to minimize their end-to-end latenciesâ
- mechanism: attach program context to individual model calls
- use previous completed calls to prioritize single-threaded and distributed programs
- author result: 4â15Ă program throughput at the same latency against evaluated vLLM
- identity check: the earlier draft calls the 2025 preprint Autellix
- this accepted paper has the Agentix title and matching authors/mechanism/results
- count it as the same research line, not independent confirming evidence
- research implication: dependent-call scheduling is an existing baseline
- reading depth: initial abstract; follow-up inspected scheduling, implementation, and workload sections
Murakkab, Chaudhry et al., OSDI 2026
- paper
- exact abstract words: âdecouples workflow specification from execution configurationâ
- mechanism: expose workflow structure to a profile-guided optimizer and adaptive runtime
- choose models and hardware subject to user-defined targets
- author result: up to 2.8Ă less GPU use, 3.7Ă less energy, and 4.3Ă less cost in its comparisons
- research implication: model/tool/GPU optimization across workflow stages is already being studied
- a candidate extension needs to identify what dynamic behavior the exposed workflow cannot capture
- reading depth: initial abstract; follow-up inspected optimizer, runtime, evaluation setup, and profile generality
more source cards, added 7 Oct 2026
- how to read these
- quotes are exact words from the abstract on the linked page unless a card names a section
- reading depth is the abstract unless a card says âfull textâ
- numbers are the authorsâ own, on their own hardware and workloads
- âfor usâ lines are my inference
- a venue is from my memory when the linked page does not show it
engines and memory inside one machine
- vLLM, Kwon et al., SOSP 2023, arXiv 2309.06180
- âPagedAttention, an attention algorithm inspired by the classical virtual memory and paging techniques in operating systemsâ
- idea: give each requestâs KV cache small fixed blocks on demand, and let requests share blocks
- authors: âimproves the throughput of popular LLMs by 2-4 with the same level of latencyâ
- for us: the block table and its sharing are the state a correctness spec must talk about, see study 5
- SGLang, Zheng et al., arXiv 2312.07104
- ânovel optimizations like RadixAttention for KV cache reuseâ
- idea: keep cached prefixes in a tree so any request that starts the same way reuses them
- for us: this tree is shared by all users of one engine, which is where the leaks and mix-ups below come from
- NanoFlow is in the first list; three OSDI 2026 papers push the same inside-the-GPU line
- DirectKV, Luo and Shen
- âthe first zero-copy KV cache offloading system for modern heterogeneous CPUâGPU platformsâ
- ECHO, Liu et al.
- KV cache offload for models with sparse attention; âup to 2.1Ă higher generation throughput than state-of-the-art systems such as SGLang and vLLM under long-context workloadsâ
- Revisiting Pipeline Parallelism for LLM Serving, Hwang and Ahn
- âpipeline parallelism with our mechanisms outperforms tensor parallelismâ for two Qwen models on four A100s
- DirectKV, Luo and Shen
- Strata, Xie et al., OSDI 2026, arXiv 2508.18572
- âexisting schedulers fail to account for cache-loading delays, leaving systems loading-bound rather than compute-boundâ
- authors: âup to 5x lower Time-To-First-Token (TTFT) compared to vLLM + LMCacheâ
- for us: this disagrees in emphasis with NanoFlowâs âcompute boundâ; long cached contexts move the bottleneck to loading
- LoongServe, Wu et al., SOSP 2024, arXiv 2404.09526
- âelastic sequence parallelism (ESP), to elastically adapt to the variance between different requests and phasesâ
- authors: throughput âup to 3.85 compared to the chunked prefill and 5.81 compared to the prefill-decoding disaggregationâ
- for us: an early paper that reports a fixed split losing
- Helix, Mei et al., ASPLOS 2025, arXiv 2406.01566
- âformulate inference computation of LLMs over heterogeneous GPUs and network connections as a max-flow problemâ
- Mélange, Griggs et al., arXiv 2404.14527
- âthe most cost-efficient allocation for a given service is typically a mix of heterogeneous GPU typesâ
- Vidur, Agrawal et al., MLSys 2024, arXiv 2405.05465
- a simulator of serving performance; âestimates inference latency with less than 9% error across the rangeâ
- for us: the cheap way to run a first replay experiment
- caution from âCalibrate, Then Routeâ below: constants taken from a simulator cost that paper â4.5 goodput pointsâ
ordering requests when their length is unknown
- FastServe, Wu et al., NSDI 2026, paper, arXiv 2305.05920
- âenable preemption at the granularity of each output tokenâ
- idea: start every job at high priority and demote it the longer it runs
- the 2023 preprint claims âup to 31.4xâ throughput over vLLM; the NSDI 2026 abstract says âup to 6.1Ăâ
- I read this as the baseline getting better over three years, and a reason not to trust old speedups
- S3, Jin et al., NeurIPS 2023, arXiv 2306.06000
- âpredicts the output sequence length, schedules generation queries based on the predictionâ
- Efficient LLM Scheduling by Learning to Rank, Fu et al., NeurIPS 2024, arXiv 2408.15792
- âalthough predicting the exact generation length of each request is infeasible, it is possible to predict the relative ranks of output lengths in a batchâ
- Andes, Liu et al., arXiv 2404.16283
- âusers receive the first token promptly and subsequent tokens at a smooth, digestible paceâ
- idea: a user cannot read faster than a fixed speed, so tokens delivered faster than that are wasted effort
- VTC, Sheng et al., OSDI 2024, arXiv 2401.00588
- âthe definition of LLM serving fairness based on a cost function that accounts for the number of input and output tokens processedâ
- âWe prove a 2x tight upper bound on the service difference between two backlogged clientsâ
- for us: one of few serving papers with a proved property; the proof is about the algorithm on paper, not the code
- QLM, Patke et al., SoCC 2024, arXiv 2407.00047
- one queue for batch and interactive requests, ordered by estimated waiting time
- SLOs-Serve, Chen et al., arXiv 2504.08784
- âcustomize the allocation of tokens to meet these SLO requirementsâ
- Niyama, Goel et al., arXiv 2503.22562
- âselective request relegation that enables graceful service degradation during overload conditionsâ
- Apt-Serve, Gao et al., SIGMOD 2025, arXiv 2504.07494
- keeps a smaller âhidden cacheâ in place of some KV cache so more requests fit in a batch
- Cascade (the scheduler, not a model cascade), Adnan et al., arXiv 2608.06557, Aug 2026
- âthe difference between a requestâs service level objective and its predicted remaining service timeâas its per-request latency budgetâ
- idea: one number per request decides both its queue position and whether its cache is reloaded or recomputed
- for us: this group is full
- JITServe, already carded, is the strongest baseline
- every paper depends on a length guess; none of the abstracts reports what happens when the guess model goes stale
splitting prefill and decode, and the papers that push back
- Splitwise, Patel et al., ISCA 2024, arXiv 2311.18677
- âwe propose splitting the two phases of a LLM inference request on to separate machinesâ
- authors: â1.4x higher throughput at 20% lower cost than current designsâ
- TetriInfer, Hu et al., arXiv 2401.11181
- âdisaggregates prefill and decode instances so each can run independentlyâ
- semi-PD, Hong et al., arXiv 2504.19867
- âthe advantage of the disaggregated system lies in the disaggregated computationâ
- names four storage costs of a full split, among them âKV cache transfer overhead between the two phasesâ
- idea: split the GPUâs compute between the phases but keep one copy of weights and cache
- Arrow, Wu et al., arXiv 2505.11916
- âsignificant fluctuations in request input/output lengths lead to imbalanced computational loads between prefill and decode nodes under traditional static node allocationâ
- EcoServe, Du et al., OSDI 2026, paper
- a full split âdepends heavily on high-performance interconnects that such clusters lackâ
- idea: each instance alternates between the phases over time, and instances take turns so one is always free for prefill
- authors: goodput gains of â1.96Ă, 1.99Ă, 2.51Ă, and 2.40Ăâ over âvLLM, Sarathi, DistServe, and MoonCakeâ on 32 L20 GPUs over Ethernet
- SmartGen, Luo et al., arXiv 2607.28150, Jul 2026
- âtransferring enormous key-value (KV) caches between disaggregated nodes can easily saturate the limited inter-node network bandwidthâ
- idea: send only the cache entries decode will need first
- HeteroScale, Li et al., ByteDance, arXiv 2508.19559
- âcritical imbalances between prefill and decode stagesâ
- âBy leveraging a single, robust metric to jointly scale prefill and decode poolsâ
- authors: âDeployed in a massive production environment on tens of thousands of GPUsâ
- Calibrate, Then Route, Tumkur et al., arXiv 2609.16206, Sep 2026
- a small honest measurement on eight A40 GPUs
- âthe calibrated router achieves the highest mean goodput at 0.864, compared with 0.835 to 0.847 for round robin, least loaded, and a length heuristicâ
- âBenefits grow with decode pool size and traffic heterogeneity but disappear in pools with three instances, where queue counts are often enoughâ
- for us: on a split cluster, clever routing bought about 2 goodput points over counting queues; that weakens study 2
- for us: split or not is a real open disagreement, and the answer seems to depend on network speed, load mix and scale
- DistServe, Splitwise, Mooncake and HeteroScale report wins from splitting
- LoongServe, semi-PD, Arrow, EcoServe and Libra report losses from a fixed split
- each paper uses its own hardware and traces, so nobody can tell which condition flips the answer
prefix caches, and routing to the machine that holds them
- Prompt Cache, Gim et al., MLSys 2024, arXiv 2311.04934
- âprecomputing and storing the attention states of these frequently occurring text segments on the inference serverâ
- CacheGen, CacheBlend and LMCache, Liu, Yao et al., University of Chicago
- CacheGen, SIGCOMM 2024: compress the KV cache to send it over a network
- CacheBlend, EuroSys 2025: âreuses the precomputed KV caches, regardless prefix or not, and selectively recomputes the KV values of a small subset of tokensâ
- this is approximate reuse: output can differ from a full recompute
- LMCache: the open-source cache layer for vLLM and SGLang
- âcontext truncation, which is a widely applied technique in industry, can greatly reduce prefix cache hit ratio by halfâ
- the storage note has more on where the cache bytes live
- DroidSpeak, Liu et al., NSDI 2026, arXiv 2411.02820
- âKV cache reuse across distributed nodes running inference of different LLMs, so long as the LLMs have the same architectureâ
- âselectively recomputes a few layers of the KV cache produced by another LLM and reuses the remaining layers, with negligible quality lossâ
- CacheSlide, Liu et al., FAST 2026, paper
- reuses cached segments of agent prompts whose position shifted, and corrects them âusing learned weightsâ
- Bidaw, Hu et al., FAST 2026, paper
- âleveraging LLM-generated responses to predict user access patterns during KV evictionâ
- KVCache Cache in the Wild, Wang et al., ATC 2025, arXiv 2506.02634
- âthe first systematic characterization of the KV$ workload patterns from one of the leading LLM service providersâ
- âreuses between single-turn requests are equally important as multi-turn requestsâ
- âfor a specific request category, the pattern tends to be predictableâ
- âthe overall cache size required for an ideal cache hit ratio is moderateâ
- for us: a production measurement of cache reuse already exists; ServeGenâs remark that this needs request content is partly answered
- Preble, Srivatsa et al., ICLR 2025, arXiv 2407.00023
- âthe first distributed LLM serving platform that targets and optimizes for prompt sharingâ
- âco-optimizes KV state reuse and computation load-balancingâ
- AIBrix, arXiv 2504.03648
- the open-source cluster layer around vLLM: âLLM-specific autoscalers, and prefix-aware, load-aware routingâ plus âa distributed KV cacheâ
- Lodestar, Lim et al., arXiv 2606.00946, May 2026
- âtrains an online reward predictor that it uses to route inference requestsâ
- authors: âlearns these efficient routing strategies within about 5 minutesâ
- SkyWalker, Xia et al., arXiv 2505.24095
- âaggregates regional diurnal patterns through cross-region traffic handlingâ
- idea: send a busy regionâs overflow to a region that is asleep, but keep each conversation where its cache is
- authors: âreducing total serving cost by 25%â
- GORGO, Toniolo et al., arXiv 2602.11688, Feb 2026
- âholistically factors network latency, prefill cost, and queueing delay using tunable parametersâ
- also useful as data: âopen-source chat datasets such as LMSYS-Chat1M and WildChat-4.8M lack long-context, high prefix-reuse data, we release a synthetic dataset, ART-Chat-2.5Mâ
- for us: LMetric, Preble, AIBrix, Lodestar, SkyWalker and GORGO all score a replica by cache match and load
- they differ in the formula; LMetricâs claim is that plain multiplication is enough
serving agents: caches across tool calls
- InferCept, Abhyankar et al., ICML 2024, arXiv 2402.01869
- todayâs engines âtreat each external interaction as the end of LLM generation and form a new request when the interaction finishes, causing unnecessary recomputation of already computed contexts, which accounts for 37-40% of total model forwarding timeâ
- the earliest paper I found on what to do with the cache while a tool runs
- Parrot, Lin et al., OSDI 2024, arXiv 2405.19888
- apps âhave to use the over-simplified request-level API provided by todayâs public LLM services, losing essential application-level informationâ
- idea: let the app tell the service how its requests feed each other
- Teola, Tan et al., arXiv 2407.00326, and Alto, Raghavan et al., arXiv 2403.04311
- both run a multi-step app as a dataflow graph so steps overlap; Teola reports âup to 2.09x speedupâ
- KVFlow, Pan et al., arXiv 2507.07400
- LRU âoften discards KV caches shortly before their reuseâ
- idea: evict by how many workflow steps remain until an agent runs again
- Continuum, Li et al., arXiv 2511.02230, carded in agent systems
- âselectively pins the KV cache in GPU memory with a time-to-live value determined by the reload cost and potential queueing delay induced by evictionâ
- CacheWise, Tiwari et al., arXiv 2606.16824, Jun 2026
- âcollecting a dataset of real-world coding assistant tracesâ
- âreuse-aware eviction guided by lightweight predictions from tool call metadataâ
- authors: âimproves total agent session completion time by up to 3.5xâ
- for us: this already uses tool information to decide eviction, which is most of study 4
- Leyline, Ma et al., arXiv 2606.01065, May 2026
- âAgentic LLMs break this assumption. Their conversations evolve through policy-driven editing: failed tool calls are retried, stale outputs dropped, trajectories pivotedâ
- idea: let the agent tell the engine to cut or replace a span in the middle of a cached context without recomputing what follows
- for us: editing a cache in place is a new way to get wrong outputs; nobody has checked it against a full recompute at scale
- AgentKV, Liu et al., arXiv 2609.14872, and MemDecay, Matam et al., arXiv 2607.10582
- both drop parts of the cache inside one long agent context and accept some quality loss
- model-side work more than systems work; listed so we know it exists
- TraceLab and the GitHub Copilot trace study are carded in agent systems
- with CacheWise that makes three 2026 measurements of coding-agent traffic
starting and scaling models
- ServerlessLLM is in the first list
- BlitzScale, Zhang et al., OSDI 2025, paper, arXiv 2412.17246
- âloading parameters through the compute network between GPUsâ
- âoffload the layer computation from the overloaded serving instances to the scaled ones without waiting for the parameters to be fully loadedâ
- authors: âup to 94 % lower tail latency reductions compared to state-of-the-art autoscaling system (ServerlessLLM)â
- λScale, Yu et al., arXiv 2502.09922
- âenabling distributed inference execution during model transmission â referred to as "execute-while-load"â
- HydraServe, Lou et al., NSDI 2026, arXiv 2502.15524
- âproactively distributes models across servers to quickly fetch them, and overlaps cold-start stages within workersâ
- authors: âreduces the cold start latency by 1.7â 4.7â
- DeepServe, Hu et al., Huawei, ATC 2025, arXiv 2501.14417
- âpre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in secondsâ
- âhas been in production for over a yearâ
- Accelerating Model Loading in LLM Inference by Programmable Page Cache, Liu et al., Huawei, FAST 2026
- loads models faster by changing the kernelâs file cache policy; âreduces the model loading latency by up to 79%â
- SageServe, Jaiswal et al., Microsoft, arXiv 2502.14617
- âwe characterize the LLM serving workloads at Microsoft Office 365â
- âwith over 10 million requests per dayâ
- âcombines short-term request routing to data centers with long-term scaling of GPU VMsâ
- authors: âreduce GPU-hour wastage due to inefficient auto-scaling by 80%â
- the arXiv comment says traces and simulator are released
- SkyServe, Mao et al., EuroSys 2025, arXiv 2411.01438
- âleverages spot replicas across different failure domains (e.g., regions and clouds)â
- authors: âreduces cost by 43% on averageâ
- DynamoLLM, Stojkovic et al., HPCA 2025, arXiv 2408.00741; TAPAS, arXiv 2501.02600; POLCA, arXiv 2308.12908
- the Microsoft line on energy, heat and power caps for inference clusters
- DynamoLLM: âconserves 53% energy and 38% operational carbon emissionsâ
- listed for completeness; I did not go deeper
many models, many adapters, mixed work on shared GPUs
- AlpaServe, Li et al., OSDI 2023, arXiv 2302.11665
- âmodel parallelism can be additionally used for the statistical multiplexing of multiple devices when serving multiple modelsâ
- MuxServe, Duan et al., ICML 2024, arXiv 2404.02015
- âcolocate LLMs considering their popularity to multiplex memory resourcesâ
- Prism, Yu et al., OSDI 2026, arXiv 2505.04021
- âa dynamic bursty-group pattern in which sets of models become active together and shift over timeâ
- idea: let a modelâs memory grow and shrink so idle models give memory to busy ones
- âdeployed in production environments across 10K+ GPUsâ
- Aegaeon, Alibaba Cloud and Peking University, SOSP 2025
- I read only Alibaba Cloudâs own summary, not the paper
- âAegaeon pioneers token-level scheduling, enabling dynamic model switching decisions after each generated tokenâ
- Weaver, Gao et al., ATC 2025, paper
- âa small number of hot models receive the majority of the requests, while most other models remain coldâ
- adapters: dLoRA, OSDI 2024, Toppings, ATC 2025, Chameleon, MICRO 2025
- dLoRA: âdynamically merge and unmerge adapters with the base modelâ
- Toppings: âuses CPUs to compute the lightweight adaption for prefilling as the requested LoRA adapter is being loaded onto GPUsâ
- Chameleon: âcaches popular adapters in GPU memoryâ
- ConServe, Qiao et al., arXiv 2410.01228
- âco-serve latency-critical online requests alongside latency-tolerant offline tasksâ
- FlexLLM, Oliaro et al., NSDI 2026, arXiv 2402.18789
- âco-serve LLM inference and PEFT-based finetuning on shared GPUs by fusing computation at the token levelâ
- BatchGen, Xu et al., OSDI 2026, paper
- âexisting inference engines still rely on execution models designed for interactive servingâ
- idea: treat each sequence of a huge offline job as a small task the runtime can pause and move
- for us: every system here puts more tenantsâ state on one GPU, which makes isolation mistakes cost more
which model answers: routers and cascades
- FrugalGPT, Chen et al., arXiv 2305.05176
- âcan match the performance of the best individual LLM (e.g. GPT-4) with up to 98% cost reductionâ
- Hybrid LLM, Ding et al., ICLR 2024, arXiv 2404.14618
- âa router that assigns queries to the small or large model based on the predicted query difficultyâ
- RouteLLM, Ong et al., arXiv 2406.18665
- trains routers on human preference data; âreduces costs-by over 2 times in certain cases-without compromising the quality of responsesâ
- these three pick by answer quality and price only; the next ones add machine load
- Cascadia, Jiang et al., arXiv 2506.04203
- âthe co-optimization of system deployment and routing strategyâ
- IC-Cache, Yu et al., SOSP 2025, arXiv 2501.12689
- âover 70% of user requests to LLMs have semantically similar counterpartsâ
- idea: show a small model past answers from a big model as examples, then send it the easy requests
- RouterWise, Kasnavieh et al., arXiv 2604.10907, Apr 2026
- âprior routing methods typically assume that each model has a fixed latencyâ
- âachievable output-quality score can vary by up to 87% across retained setupsâ
- HW-Router, Kabir et al., arXiv 2608.14575
- âintegrates real-time hardware signals into model selectionâ
- Cluster, Route, Escalate, Moslem et al., arXiv 2606.27457
- âwhen an output from Stage 1 is judged low-quality, the query is escalated to a stronger modelâ
- for us: a router quietly changes which model a user gets
- that is the same act the substitution audits below try to detect
when a GPU or worker dies
- how often: Story of Two GPUs, Cui et al., arXiv 2503.11901
- â2.5 years of operational data (11.7 million GPU hours) on GPU errorsâ from a 1,056-GPU academic cluster
- âGPU errors on both A100 and H100 GPUs frequently result in job failures due to the lack of robust recovery mechanisms at the application levelâ
- âsignificant overprovisioning of 5% is necessary to handle GPU failuresâ
- Déjà Vu, Strati et al., ICML 2024, arXiv 2403.01876, full text skimmed
- âUpon a failure, the LLM serving system crashes and stalls all in-flight requestsâ (introduction)
- âreplicates KV cache state to avoid losing state and employs fast recovery mechanism to minimize lost work on failuresâ (introduction)
- the first pass missed this 2024 paper; it already did what study 3 hypothesized
- FailSafe, Xu et al., arXiv 2511.14116
- âa single GPU failure can halt execution, trigger costly KVCache recomputation, and introduce long-term compute and memory imbalanceâ
- keeps serving on the GPUs left in a tensor-parallel group; âtwo orders of magnitude lower recovery latency compared to standard fault handling approachesâ
- KevlarFlow, Qian et al., arXiv 2601.22438, Jan 2026, full text skimmed
- âCurrent recovery mechanisms are prohibitively slow, often requiring up to 10 minutes to reinitialize resources and reload massive model weightsâ
- âbackground KV cache replication to maintain high throughput during partial failuresâ
- authors: âreduces mean-time-to-recovery (MTTR) by 20xâ
- GhostServe, Jayakody et al., MLSys 2026, arXiv 2605.00831
- âapplying erasure coding to generate and store the parity shards in host memoryâ
- âallowing the inference process to resume seamlessly without costly full recomputation or state replicationâ
- LUMEN, Cao et al., arXiv 2606.17787, Jun 2026, full text skimmed
- âtreats recovery as a load-aware coordination problem across three decision points: checkpoint placement before failures, interrupted-request distribution at failure time, and serving capacity restoration during model reloadâ
- introduction: âLLM serving jobs encounter failures every few hours on averageâ
- stated limit: âcurrently maintains a single KV checkpoint per request, falling back to full recomputation if the checkpoint holder itself fails; we defer multi-checkpoint replication to future workâ
- Concordia, Gan et al., arXiv 2606.23521, Jun 2026
- âLosing this state after a GPU or communicator failure can discard minutes to hours of workâ
- idea: a small always-running program on the GPU logs changed cache blocks to host memory
- SAVE, Zheng et al., ATC 2025, paper
- memory bit flips that âsilently corrupt resultsâ; evaluated on vision and robotics models, not LLM serving
- for us
- six systems now save the KV cache and resume, so the mechanism is taken
- what I could not find in the abstracts or in my skim of three full texts
- whether a resumed reply is the same text the user would have got without the failure
- what happens to tokens already streamed to the user
- KevlarFlowâs own evaluation says âSome overhead values are negative due to non-determinism in the executionâ, so runs are not repeatable even for the authors
is the engine correct
- GRIEF: Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing, Zhao et al., Maryland and NYU, arXiv 2605.11202, May 2026, full text skimmed
- âtreats timed multi-request traces as first-class inputsâ
- âdiscovers 15 vulnerabilities, 10 confirmed by engine developers, including 2 CVEs, spanning KV-cache isolation failures, cross-request performance interference, and crash or liveness bugsâ
- âconcurrency, caching, and state reuse can induce silent cross-request contamination, noisy-neighbor denial of service, and delayed crashes without malformed inputs or explicit server errorsâ
- how it decides something is a bug: replay the trace alone and compare token probabilities
- §3.4: âit avoids reporting cases where nearly tied logits could legitimately decode differentlyâ
- stated limits (appendix A.2)
- âfocuses on vLLM and SGLangâ; âdistributed deploymentsâ would need more work
- âLike other fuzzers, GRIEF does not prove the absence of bugsâ
- for us: direct evidence that one userâs request can change another userâs answer in the two most used engines
- the paper tests one engine process, not a cluster with cache transfer, migration or recovery
- A First Look at Bugs in LLM Inference Engines, Liu et al., TOSEM, arXiv 2506.09713
- âa comprehensive dataset of 929 real-world bugsâ from 5 engines
- âsix bug symptom types and a taxonomy of 28 root causesâ
- The Foundation Cracks, Jiang et al., arXiv 2506.12320
- â313 bug-fixing commitsâ from HuggingFace Transformers and vLLM
- âthe majority of bugs escape detection due to inadequate test cases (41.73%), lack of test drivers (32.37%), and weak test oracles (25.90%)â
- for us: a quarter of escaped bugs lacked a way to tell right from wrong output, which is the thing a spec gives
- why the same prompt gives different answers
- Defeating Nondeterminism in LLM Inference, Thinking Machines blog, Sep 2025
- âthe primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies!â
- âif weâd like to avoid nondeterminism in our inference servers, we must achieve batch invariance in our kernelsâ
- Yuan et al., arXiv 2506.09501
- âchanging system configuration, such as evaluation batch size, GPU count, and GPU version, can introduce significant differences in the generated responsesâ
- a reasoning model âcan exhibit up to 9% variation in accuracy and 9,000 tokens difference in response lengthâ
- LLM-42, Gond et al., Microsoft Research and UW, arXiv 2601.17768, Jan 2026, full text skimmed
- âdecodes tokens using a non-deterministic fast path and enforces determinism via a lightweight verify-rollback loopâ
- âincurs overhead only in proportion to the traffic that requires determinismâ
- introduction: âSGLang incurs high overhead of up to 56% in deterministic modeâ
- âverifiedâ here means re-run and compared, not proved
- for us: without batch invariance, âis this output rightâ has no exact answer, so tests fall back to fuzzy comparisons like GRIEFâs
- a deterministic mode now exists in SGLang and in LLM-42, which makes exact comparison possible for the first time
- Defeating Nondeterminism in LLM Inference, Thinking Machines blog, Sep 2025
- A postmortem of three recent issues, Anthropic, Sep 2025
- a vendor account of serving bugs that lowered answer quality for weeks
- âsome Sonnet 4 requests were misrouted to servers configured for the upcoming 1M token context windowâ
- âAt the worst impacted hour on August 31, 16% of Sonnet 4 requests were affectedâ
- âsome users were affected more severely, as our routing is "sticky"â
- a compiler bug in âthe approximate top-k operationâa performance optimization that quickly finds the highest probability tokensâ
- for us: a production routing bug and a numeric bug, neither a crash, both found late; the fuzzers and bug studies above do not cover the cluster layer where the first one lived
- StriaTrace, Wu et al., Alibaba, OSDI 2026, paper
- âdetailed tracing only during abnormalitiesâ
- âhas successfully diagnosed hundreds of abnormalities spanning 19 distinct root causesâ
- about slow requests, not wrong ones
- failure and outage studies in the bug-finding folder has an incident study with an inference-engine category
- formal verification: I searched for a proved or model-checked serving engine, scheduler or prefix cache and found none
- one search on 7 Oct 2026, so absence is weak evidence
- VTCâs fairness bound and JITServeâs competitive bound are paper proofs about algorithms
is the shared cache private
- attacks
- The Early Bird Catches the Leak, Song et al., TIFS, arXiv 2409.20002
- ânew timing side channels in LLM systems, arising from shared caches and GPU memory allocationsâ
- âa token-by-token search algorithm to efficiently recover shared prompt prefixesâ
- I Know What You Asked, Wu et al., NDSS 2025
- title and venue confirmed; I read search summaries only, so no quote
- InputSnatch, Zheng et al., arXiv 2411.18191
- âcaching can result in observable variations in response timesâ
- SpliceLeak, Sun et al., arXiv 2606.21842, Jun 2026
- âthe first end-to-end side-channel attack targeting non-prefix KV cache fusionâ, tested on âvLLM integrated with LMCacheâ
- The Early Bird Catches the Leak, Song et al., TIFS, arXiv 2409.20002
- measured on real providers: Auditing Prompt Caching in Language Model APIs, Gu et al., Stanford, ICML 2025, arXiv 2502.07776
- âWe detect global cache sharing across users in seven API providers, including OpenAIâ
- method: âstatistical audits to detect prompt caching in real-world LLM API providersâ
- defenses
- SafeKV, Chu et al., arXiv 2508.08438: share only cache blocks judged not sensitive
- PrefixWall, Pennas et al., arXiv 2603.10726: âmonitors cache reuse across users, flags suspicious sharing, and selectively isolates prefixesâ
- KVGov, Addagada, arXiv 2608.09225: a per-user secret mixed into the cache key, âmaking cache keys cryptographically disjoint across principalsâ
- single author; âthe defense itself is evaluated in simulationâ
- for us: the simple fix is one cache per user, and every defense paper is about getting some sharing back
- none of these abstracts claims a proof that the defended engine leaks nothing through timing
is the provider serving what it says
- Model Equality Testing, Gao et al., Stanford, ICLR 2025, arXiv 2410.20247
- â11 out of 31 endpoints serve different distributions than reference weights released by Metaâ, commercial APIs in summer 2024
- Are You Getting What You Pay For?, Cai et al., Berkeley, arXiv 2504.04715
- âstatistical tests on text outputs are query-intensive and fail against subtle substitutions, while methods using log probabilities are defeated by inherent inference nondeterminism in production environmentsâ
- TOPLOC, Ong et al., arXiv 2501.16007
- the provider sends a short fingerprint of internal activations that a checker can re-compute; â258 bytes of storage per 32 new tokensâ
- for us: outside audits exist but are two years old and hit a wall at nondeterminism
- the determinism work above could move that wall
workload data that anyone can download
- BurstGPT, Wang et al., KDD 2025, arXiv 2401.17644
- â10.31 million traces from regional Azure OpenAI GPT services over 213 daysâ
- includes âSystem response failuresâ
- released with papers above: Splitwiseâs Azure trace, SageServeâs Office 365 traces, ServeGenâs generator, GORGOâs ART-Chat-2.5M, TraceLabâs coding-agent sessions
- Prism, Chameleon and Weaver say they use production or real-world traces; I did not check which are public
candidate studies
- joint workload model for dependent calls, tool pauses, and cache reuse
- hypothesis: preserving their correlation changes which existing serving policy wins
- closest work
- ServeGenâs client mixtures
- Agentixâs program scheduling
- SYMPHONYâs advisory prefetching
- TraceLab and production Copilot trace studies in agent systems
- first experiment
- record model-call boundaries, tool start/end, prompt-prefix identifiers, and final task completion
- replay the same sessions with original timings, shuffled pauses, and independent arrival/length sampling
- freeze model, GPU count, batching configuration, and latency targets
- compare a current serving engine, an advisory cache policy, and program-aware scheduling
- measurements
- completed tasks per GPU-hour and tail task latency
- cache transfers, reprocessed tokens, and tenant fairness
- keep correctness scores separate from serving deadlines
- result that weakens the idea
- correlation preservation makes no material difference across workloads
- an existing generator plus a simple extension already predicts the observed rankings
- novelty status after the second pass: weaker
- three 2026 studies already measure coding-agent traffic: TraceLab, the Copilot trace study, CacheWise
- KVCache Cache in the Wild already measures cache reuse at a provider
- what I still did not find: one generator that keeps tool pauses, prefix reuse and call order together, and a test of whether that changes which policy wins
- the cheapest version is a replay study on TraceLab plus CacheWise traces, not new data collection
- robust routing when cached prefixes and load estimates become stale together
- hypothesis: worker loss or synchronized tool returns can exceed the hotspot assumptions in existing routing
- closest work
- LMetricâs hotspot detector and fallback
- Llumnixâs live migration and scheduler bypass
- Mooncakeâs cache-aware placement
- Libraâs flexible partitioning
- first experiment
- inject delayed load reports, a lost cache-owning worker, and simultaneous returns from tools
- vary shared-prefix concentration and state-transfer bandwidth
- compare existing fallbacks before designing another policy
- possible mechanism after a reproduced failure
- discount old cache/load reports using their age
- cap traffic committed to one cache owner
- reserve capacity for migration or cold recomputation
- measurements
- time to recover steady useful throughput
- peak queue length, worst-tenant latency, and extra transfers
- result that weakens the idea
- LMetric fallback plus standard failover handles the whole stress range cheaply
- novelty status: unresolved
- stale-information load balancing has a long history beyond LLM serving
- second pass: more routers to beat, and a warning
- Preble, AIBrix, Lodestar, SkyWalker and GORGO all route by cache match and load
- Lodestar retrains online, which is itself an answer to stale estimates
- Calibrate, Then Route measured about 2 goodput points between a tuned router and queue counting on 8 GPUs
- I would only pursue this if the first injection experiment shows a collapse, not a few percent
- request recovery with an explicit client-visible stream contract
- hypothesis: protecting a small amount of request metadata yields a useful recovery tradeoff without replicating every KV tensor
- closest work
- Llumnix already aborts failed in-flight requests and restarts workers
- Mooncake already handles transfer failures
- DistServe explicitly leaves advanced fault tolerance outside its implementation
- first experiment
- kill workers before first token, after a client-visible token, and during state transfer
- compare full retry, prefix recomputation, and any existing serving recovery mechanism
- include a fixed random seed and record sampling configuration where supported
- contract to specify before implementation
- whether already delivered tokens can be repeated or replaced
- whether recovery continues the same token sequence or starts a new attempt
- how clients detect a new attempt and cancel old work
- measurements
- recovery time, wasted GPU work, duplicate output, and incorrect request attribution
- evaluate latency under failures as well as normal throughput
- result that weakens the idea
- application retry is cheap enough and acceptable to users
- existing request-resumption machinery already supplies the desired contract
- limitation: preserving numerical and sampling state across hardware changes may be expensive
- exactly the same output is stronger than a usable resumed answer
- novelty status after the second pass: the mechanism is taken, the contract looks open
- Déjà Vu, FailSafe, KevlarFlow, GhostServe, LUMEN and Concordia all keep a copy of the KV cache and resume from it
- so the hypothesis above about âa small amount of request metadataâ must now beat those six, not full retry
- none of them states, in what I read, whether the resumed text equals the text without a failure
- my reasoning for why it usually would not: the resumed request lands in a different batch, and batch changes numbers (Thinking Machines, LLM-42)
- revised question: after recovery or migration, does the user get the same tokens, and if not, how often does the answer change?
- revised first experiment
- run SGLang in deterministic mode, or LLM-42, so that a no-failure run is an exact reference
- kill a worker mid-reply under Déjà Vu-style and LUMEN-style recovery, and under Llumnix migration
- count replies whose remaining tokens differ from the reference, and replies whose final answer differs
- result that weakens it: differences are as rare as ordinary run-to-run noise, or users cannot tell
- this merges naturally into study 5
- tool progress as a fallible scheduling input
- hypothesis: a running tool can provide better estimates than a serving engineâs model of past tool durations
- closest work
- Ask the Tool, Donât Guess in agent systems
- SYMPHONYâs hints
- JITServeâs progressive estimation
- Murakkabâs adaptive workflow runtime
- first experiment
- report tool progress, its age, and confidence through one controlled interface
- compare truthful, delayed, missing, and wrong progress signals
- compare against learned-duration prediction and fixed cache-retention timeouts
- measurements
- post-tool response latency and completed task throughput
- state retained needlessly and useful state evicted early
- result that weakens the idea
- simple duration prediction matches progress hints
- tool implementations cannot provide useful signals at reasonable cost
- novelty status: mechanism already proposed in related work
- robustness to wrong hints and end-to-end co-scheduling need a more precise difference
- second pass: I would drop this as a standalone study
- InferCept, Continuum, KVFlow and CacheWise already decide cache retention around tool calls
- CacheWise already uses âlightweight predictions from tool call metadataâ
- the one piece left is hints that are wrong or late, which fits better as one experiment inside study 1
- a written rule for âeach user gets their own answerâ, then test or prove engines against it
- the rule in plain words: the tokens a request gets back depend only on that requestâs input, the model and the sampling seed
- not on which other requests shared its batch, its cache, its GPU, or a recovery
- why I think this is open
- GRIEF found âsilent cross-request contaminationâ in vLLM and SGLang by fuzzing one engine process
- The Foundation Cracks blames âweak test oraclesâ for a quarter of escaped bugs
- I found no proved or model-checked engine, scheduler or prefix cache
- until 2025 the rule could not be checked exactly, because batching changed the numbers; deterministic modes remove that excuse
- closest work
- GRIEF, the two bug studies, LLM-42, the Thinking Machines post
- the prompt-cache attack papers, which break a weaker rule (timing, not content)
- the storage noteâs question âwhat does a shared KV cache hit promise?â in LLMs and storage
- first experiment, about two weeks
- reference: one request at a time, deterministic mode, no prefix cache
- system under test: same engine with batching, prefix cache, preemption, cache offload, and then a cluster layer (LMCache, a prefill/decode split, migration)
- feed GRIEF-style timed traces and require exact token equality with the reference
- count mismatches per feature turned on
- second step if the first finds bugs or finds none
- write the prefix-cache tree and block table as a small state machine and state the rule as an invariant
- model-check it, or write that component in Rust and prove it with Verus
- the component is small: a tree keyed by token blocks, reference counts, eviction
- approximate reuse (CacheBlend, DroidSpeak, CacheSlide, Leyline) breaks exact equality on purpose; the rule for those needs a bound, which I do not know how to state yet
- measurements
- mismatching replies per million, by feature
- bugs confirmed by maintainers
- cost of the deterministic reference
- result that weakens the idea
- exact equality holds everywhere GRIEF did not already look
- deterministic mode is too slow or too incomplete to serve as a reference
- maintainers treat mismatches as acceptable noise
- novelty status: I think the cluster-level and proof parts are new; one search, so check again before committing
- GRIEFâs group will likely extend to distributed deployments, their appendix names it
- audit hosted LLM APIs from outside, again and over time
- question: in late 2026, which providers share prompt caches across customers, swap or quantize models, or route the same model name to different backends?
- why it fits us: this is web measurement with an LLM endpoint as the site
- closest work
- Gu et al. found cache sharing at 7 providers in early 2025
- Gao et al. found 11 of 31 Llama endpoints differed from the released weights in summer 2024
- Cai et al. showed where output-only tests fail
- Anthropicâs postmortem shows a provider misrouting 16% of one modelâs requests at the worst hour
- it says âdetection and resolution took longer than we would have wantedâ
- first experiment
- rerun Gu et al.âs timing audit and Gao et al.âs equality test on todayâs providers and aggregators
- add repeat probes over weeks to catch changes, and probes from several regions to catch routing differences
- record whether providers fixed what the 2025 audits reported
- measurements
- providers with cross-customer cache hits
- endpoints whose output distribution differs from reference weights
- change over time and by region
- result that weakens the idea
- providers fixed sharing after 2025 and nothing new shows up
- nondeterminism hides everything smaller than a model swap, as Cai et al. warn
- care needed: probing only our own accounts and our own prompts; no attempt to read other customersâ prompts
- novelty status: the methods exist, the longitudinal and multi-region view is what I did not find; unverified beyond one search
- one fair test of âshould prefill and decode be splitâ
- question: on the same hardware and the same traces, when does splitting win?
- why: ten papers give opposite answers on their own setups, listed under âsplitting prefill and decodeâ
- first experiment
- one cluster, two network speeds, three traces (chat from ServeGen, coding agents from TraceLab, long documents)
- run vLLM or SGLang unsplit with chunked prefill, a fixed split, and the open-source adaptive ones (EcoServe released code)
- sweep load and the input-to-output length ratio
- measurements
- goodput at one stated latency target, GPU-hours per million tokens, and the crossover points
- result that weakens the idea
- Libraâs or EcoServeâs evaluation already covers the same grid fairly
- the answer is just âsplit when the network is fastâ, which the papers nearly say
- novelty status: a benchmark paper, lower risk and lower ceiling
- needs more GPUs than the other studies
reading limits and search record
- first pass, 6 Oct 2026: reviewed 14 named research lines
- eight recent full PDFs were retrieved alongside four earlier full PDFs
- source cards state the passages actually read
- retrieval alone is not a full-paper review
- searched primary USENIX OSDI 2025/2026, NSDI 2025/2026, and FAST 2026 programs
- conference pages and linked PDFs were accessible through direct HTTPS
- both configured web search services failed during this session
- second pass, 7 Oct 2026: about 100 more sources
- found by web search, by scanning the OSDI 2025/2026, NSDI 2025/2026, ATC 2025 and FAST 2026 programs for serving titles, and from memory checked against the arXiv page
- read the abstract of each on arXiv or the USENIX page and copied quotes from there
- skimmed full text of six: LUMEN, KevlarFlow, GRIEF, LLM-42, Déjà Vu, Calibrate, Then Route
- read two web posts directly: Thinking Machines on nondeterminism, Anthropicâs postmortem
- Aegaeon and âI Know What You Askedâ rest on secondary pages only
- the arXiv search API was rate-limited, so I could not sweep arXiv by keyword; recent preprints are covered unevenly
- not covered in either pass
- SOSP 2025, EuroSys 2026, ASPLOS 2026, ATC 2026, SIGCOMM 2026 and MLSys 2026 programs were not scanned
- serving of image, video and speech models; mixture-of-experts serving; speculative decoding
- on-device and edge inference
- confidential inference in trusted hardware
- industry stacks beyond their papers: NVIDIA Dynamo, llm-d, KServe, Ray Serve
- pricing and economics of inference
- energy and power got three links and no analysis
- breadth is substantial but not exhaustive
- training resilience and MoE communication deserve separate reviews
- accepted-program coverage does not include every recent preprint or industry implementation
- author performance numbers use different hardware, workloads, engines, and targets
- they are not a ranking across papers
- recommendations and hypotheses above are agent proposals
- no claim of worldwide novelty or demonstrated benefit
closest-work follow-up: tool-aware scheduling already exists
- inspected cached conference PDFs on 8 Oct 2026
- Agentix sections 3â5 and evaluation workloads
- SYMPHONY sections 3.2â3.6, trace construction, and agent evaluation
- Murakkab sections 3.3â3.4 and evaluation setup and profiling generality
- new direct PDF requests returned HTTP 403
- conference landing pages were accessible
- earlier downloaded PDFs supplied the inspected text
- no artifact was executed or performance result reproduced
- Agentix does not require a known execution graph
- authorsâ assumptions: âits execution DAG is initially unknownâ
- Luo et al., NSDI 2026, section 4.1
- its process table tracks completed model service, waiting time, call arrivals, and completions
- its evaluation includes BFCL multi-step tool use and LATS parallel search
- implication: dynamic dependencies and tool-using programs are established scheduling inputs
- a proposal based only on exposing program identity would repeat this work
- SYMPHONY explicitly handles uncertainty in hints
- authorsâ limitation: âadvisory requests arrive early enoughâ
- Agarwal et al., NSDI 2026, section 3.6
- profiles representative agent workflows and hints at possible downstream agents
- supplies invalidation, memory-pressure eviction, and best-effort behavior for missing hints
- evaluates false or missing hints and MetaGPT workloads
- implication: imperfect hints and prefetching during dependent work are existing mechanisms
- the narrower question is whether real joint timing changes their measured value
- Murakkab includes dynamic composition and changing resource demand
- evaluation method: âapproximate workload arrivals using LLM serving tracesâ
- Chaudhry et al., OSDI 2026, evaluation setup
- maps chat arrivals to video Q/A and coding arrivals to code generation
- optimizer uses workflow and model profiles
- a shorter-timescale auto-scaler handles demand changes
- held-out math inputs test some profile generality
- implication: a workflow scheduler cannot claim dynamism alone as its contribution
- remapped arrival traces leave a concrete measurement question about actual workflow correlations
- refined workload proposal
- measure task arrivals, tool durations, dependent model calls, and reused prompt prefixes from the same executions
- preserve original correlations in one replay
- separately shuffle pauses, task arrivals, or prefix associations as controlled comparisons
- compare program-aware scheduling and advisory prefetching before adding a new scheduler
- report task completion and cache-transfer cost separately
- rejection condition: preserving those correlations does not materially change policy ranking or predicted capacity
- remaining uncertainty: available traces may already contain the needed joint information
- inspect TraceLab and CacheWise schemas and replay artifacts before collecting new data
trace artifacts narrow the workload proposal further
- TraceLab already releases a session-aware replay client
- maintainer description: âsession-aware closed-loop workload runnerâ
- TraceLab replay README, pinned revision
- CSV preserves session identity, arrival, round order, prefix length, append length, output length, and post-round waits
- implementation waits for a model response, sleeps the recorded wait, then sends the next round
- prompts use synthetic content and exact output token carry-forward where supported
- inference: the shortlistâs proposed correlated replay mechanism is already substantially implemented
- reuse and audit this artifact before proposing another workload generator
- released trace metadata preserves much of the desired joint evidence
- README field:
timing_events - TraceLab data-format and sanitization documentation
- tool fields include emission, result time, wall latency, internal latency, and continuation identity
- sanitized data removes raw tool arguments and prompt contents
- token counts and command structure cannot establish arbitrary cross-session semantic prefix sharing
- recorded provider cache counts do not directly reveal a different engineâs cache behavior
- README field:
- CacheWise also studies tool-dependent reuse
- primary dataset description: âtimestamps for each messageâ
- Tiwari et al., section 3, arXiv 2606.16824
- describes conversations, tool calls and results, token counts, and human interventions
- repository offers event extraction, workload analysis, and tool-duration prediction
- inspected README and paper characterization passages
- predictor and replay implementation were not audited
- revised decision
- do not claim that joining tool timing, dependent calls, and prefix reuse is missing from existing datasets
- candidate 2 becomes a policy-sensitivity and replay-fidelity assessment using existing traces
- first compare original timing with controlled shuffles under Agentix, SYMPHONY, and CacheWise where artifacts support integration
- a useful contribution requires evidence that a specific replay simplification changes a meaningful conclusion
- no such evidence has been obtained yet
Last edited: