running large language models: inference, serving, hardware (authored by agents unless marked đ§)
takeaways
- agent recommendation: investigate a checked numerical-reproducibility guarantee across execution changes
- reproducibility means the same specified numerical outputs recur
- it does not establish that a model gives correct answers
- Vosti proves a guarantee for its supported single-GPU engine and kernel contracts
- TBIK already demonstrates matching outputs and probabilities across tested tensor-parallel sizes
- a new project needs a narrower unproved state transition or protocol
- first gate: identify the exact guarantee, accessible hardware, and nearest implemented baseline
- one GPU supports a single-device baseline and selected speculative-decoding comparisons
- TP 1/2/4 requires at least four usable GPUs and a stated interconnect
- hardware access and completion time are unconfirmed
- keep generic agent routing as a lower-priority measurement
- Continuum, Autellix, and existing recovery work already cover major mechanisms
- WaferLLM is the humanâs starting point
- đ§ original note: âWaferLLM: Large Language Model Inference at Wafer Scaleâ
- its reported energy advantage depends on the phase and GPU baseline
- without wafer access, a comparison is a modeling project
scope and evidence
- reviewed 2026-10-07 through primary sources
- web search worked in this pass; the earlier pass had no search
- 2026 venue programs were swept: OSDI, NSDI, MLSys, EuroSys, ASPLOS, SOSP 2025
- read in full, text extracted from PDF or arXiv HTML
- foundations: Orca, vLLM, SGLang, DistServe, Sarathi-Serve, Splitwise, Llumnix, Mooncake (FAST 2025 and arXiv versions)
- cache and speculation: CacheGen, CacheBlend, Prompt Cache, Preble, InfiniGen, FlexGen, SpecInfer, EAGLE, EAGLE-3, Medusa, FlashInfer
- agents and operations: InferCept, Parrot, Autellix, Continuum v7, AgentReplay, ServerlessLLM, BlitzScale, DynamoLLM, Andes, VTC
- 2026: WaferLLM, Strata, LMetric, ServeGen
- hardware: DeepSeek-V3 insights, TPU v4, MegaScale-Infer; Groq 2022 in main sections only
- determinism: Vosti, LLM-42, Yuan et al., Thinking Machines post, SGLang post, PromptPeek
- abstract or official page only
- NanoFlow, speculative decoding (Leviathan), Sereno, ContextPilot, GhostServe, RaidServe, DriftBench (plus slides), Beyond the Buzz
- KVFlow, FlashAgents, Murakkab, MarginGate, TBIK (partial), Volta, HijackKV (partial), DeepSeek-V4 report
- vLLM batch-invariance docs: fetch failed, known only through Vostiâs citation
- reading was split across helper agents; their quotes were not re-checked against the text
- the quotes marked âauthorsâ are verbatim from the source
- âinferenceâ marks this reviewâs own conclusions
- speedups below are author claims; nothing was reproduced
the basic problem
- an LLM reads the whole prompt first, then writes one token at a time
- prefill: reading the prompt, lots of math per byte moved
- decode: writing tokens, little math per byte moved, so memory bandwidth limits it
- KV cache: the saved intermediate state for every earlier token
- saves recomputation, but it is big and must be stored, moved, evicted, or rebuilt
- whether a cached value equals a recomputed one is the thread running through this review
- goodput: requests finished inside their latency limits per second
- tokens per second can go up while goodput goes down
- measure first-token delay, gap between tokens, whole-task time, and rejections
- floating-point addition is not associative
- GPU kernels pick tile sizes and reduction orders from tensor shapes
- shapes depend on who else is in the batch, so the same request can produce different bits
literature: batching, memory, and splitting the two phases
- Orca, Yu et al., OSDI 2022, paper
- authors: âwe propose to schedule the execution of the engine at the granularity of iteration instead of requestâ
- mechanism: finished requests leave and new ones join after every token step
- attention runs per request because shapes differ; everything else is batched
- authors report 36.9Ă throughput over FasterTransformer at matched median latency, 175B model, A100s
- evaluation forced every request to its max length and used synthetic Poisson arrivals, single turn
- inference: every later serving paper assumes this; it is the floor, not a baseline to beat
- vLLM, Kwon et al., SOSP 2023, paper
- authors: âwithout affecting the model accuracy at allâ
- mechanism: KV cache in fixed blocks with a per-request block table, like OS paging
- authors report 2â4Ă over FasterTransformer and their own Orca reimplementation
- cost: â20â26% higher attention kernel latencyâ than FasterTransformerâs kernel
- multi-turn: âWe do not store the KV cache between different conversation roundsâ
- inference: the exactness claim is about memory layout; the 2026 determinism tests show vLLMâs default output bits still vary with batch
- SGLang, Zheng et al., NeurIPS 2024, paper
- authors: âKV cache computation depends only on prefix tokensâ
- mechanism: keep every prompt and output in a radix tree; schedule longest cached prefix first
- authors: âwe do not turn on optimizations that will change the computation results so that all systems compute the same resultsâ
- authors report up to 6.4Ă over vLLM v0.2.5; one-month Chatbot Arena deployment hit rates 52% and 74%
- admitted gap: âgreedy cache-aware scheduling ⊠can lead to starvationâ
- replayed ReAct and generative-agent traces, not live tool calls
- Sarathi-Serve, Agrawal et al., OSDI 2024, paper
- authors: âit throttles the number of prefill tokens in each iteration while admitting new requests in a running batchâ
- mechanism: decodes first, prefill chunks fill the rest of a fixed token budget
- authors report 2.6Ă capacity on Mistral-7B and up to 3.7Ă on Yi-34B over vLLM
- cost: chunk size 257 can cost 32% more than 256 because of GPU tile sizes
- authors: âWe leave a quantitative comparison between Sarathi-Serve and disaggregation-based solutions for future workâ
- DistServe, Zhong et al., OSDI 2024, paper
- authors: âDistServe can serve 7.4Ă more requests or 12.6Ă tighter SLOâ
- mechanism: prefill and decode on different GPUs, each with its own parallelism; a simulator picks the split
- KV transfer was under 0.1% of latency on NVLink; cross-node links were 25 Gbps
- authors: âDistServe does not implement advanced runtime policies like preemption [26] and fault tolerance [58]â
- authors on failure: âa fault in a single decoding instance mapped to multiple prefill instances could potentially cripple the entire serviceâ
- Splitwise, Patel et al., ISCA 2024, paper
- authors: âSplitwise clusters achieve up to 1.4Ă higher throughput at 20% lower costâ
- mechanism: same split as DistServe, but on separate machine pools that may use different GPUs or power caps
- authors: âIf the prompt or the token machine fail, Splitwise simply restarts requests from scratchâ
- authors: âwe do not reuse the KV-cache between requests to emulate a cloud service with security guaranteesâ
- uses Azure production traces for lengths only; cluster numbers come from a validated simulator
- Llumnix, Sun et al., OSDI 2024, paper
- authors: âthe KV cache is append-onlyâ
- mechanism: live-migrate a running request with its cache between vLLM instances; copy old blocks while decoding continues
- migration downtime about 20â30 ms regardless of length; recomputing 8k tokens on LLaMA-30B took 3.5 s
- authors: âWhen an instance (or the co-located llumlet) fails, the requests running on it will be abortedâ
- inference: migration is the one general tool for moving state without recompute; no paper checks that migrated and recomputed outputs match
- Mooncake, Qin et al., FAST 2025, paper, earlier tech report
- the arXiv report and the FAST paper differ; the FAST version is the peer-reviewed one
- authors: âKVCache corresponding to the same input prefix can be reused without affecting output accuracyâ
- mechanism: split phases, pool all CPU RAM and SSD as one cache over RDMA, route by cache hit length
- authors report 59â498% more effective requests than vLLM v0.5.1; âover 100 billion tokens dailyâ in production
- authors: âwe recommend a minimum network bandwidth of 100 Gbpsâ
- tech report caveat: âTheoretically, up to only 50% of the KVCache can be reused in our current workloadsâ
- the only paper here with real production tool-and-agent traces, average input 8,596 tokens, output 182
- replay used a dummy LLaMA3-70B architecture, so no output content was checked
- Preble, Srivatsa et al., paper
- authors: âthey are all confined to a single-GPU optimization, while production LLM serving systems are distributed by natureâ
- mechanism: a global radix tree says which GPU holds which prefix; route to it when matched tokens outnumber unmatched ones
- authors report 1.5â14.5Ă average latency over SGLang v0.1.12 round-robin
- admitted limit: âwe do not improve decoding performance, the room for improvement for Preble is smallerâ on long outputs
- includes Toolbench and ALFWorld agent traces with Poisson arrivals
- NanoFlow, Zhu et al., OSDI 2025, official abstract
- authors: âend-to-end LLM serving is compute bound for most common workloads and LLMsâ
- mechanism: overlap compute, memory, and network work inside one GPU
- inference: contradicts the âdecode is memory-boundâ folk rule for whole-server throughput; measure before optimizing
- Strata, Xie et al., OSDI 2026, paper
- authors: âwithout performance degradation in short-context scenariosâ
- mechanism: GPU-assisted transfers and tier-specific cache layouts; overlap loads with compute
- comparisons use vLLM 0.8.5, LMCache 0.2.1, TensorRT-LLM 0.17.0
- inference: CPU/SSD cache offload is a crowded direction
- LMetric, Zhang et al., OSDI 2026, paper
- authors: âwithout any hyperparameter tuningâ
- mechanism: route by cache-aware new prefill work times the instanceâs batch size
- beats vLLM, NVIDIA Dynamo, llm-d, and a production router on up to 16 H20s
- inference: a learned router must beat this one-line rule
- ServeGen, Xiang et al., NSDI 2026, paper, code
- authors: âWe leave characterizing LLM serving with plugin calls as an important area for future workâ
- mechanism: per-client traffic models that keep demand and request shape correlated
- inference: tool waits, cancellations, and reuse correlations are unmeasured in public data
- Beyond the Buzz, Mitra et al., MLSys 2026 industry, official abstract
- authors: âdisaggregation is most effective for prefill-heavy traffic patterns and larger modelsâ
- inference: the prefill/decode split has a region where it helps; papers that claim it always wins are overclaiming
literature: reusing cache beyond exact prefixes
- the trade: exact prefix reuse is safe but rare; non-prefix reuse is common but changes outputs
- CacheGen, Liu et al., SIGCOMM 2024, paper
- mechanism: compress a stored KV cache into a bitstream, pick quantization per chunk from measured bandwidth
- authors report 3.5â4.3Ă smaller cache and 3.2â3.7Ă lower load delay, 4Ă A40
- lossy: âno more than 2% in accuracy, less than 0.1% in F1 scoreâ
- authors: âfew industry datasets exist to support itâ on context reuse
- CacheBlend, Yao et al., EuroSys 2025, paper
- mechanism: reuse per-chunk caches that are not prefixes; recompute only the ~15% of tokens whose values deviate most
- authors: âTokens with the highest KV deviations on one layer are likely to have the highest KV deviations on the next layerâ
- authors report 2.2â3.3Ă lower first-token time than full recompute, quality within 0.02
- no guarantee; approximate by construction
- Prompt Cache, Gim et al., MLSys 2024, paper
- authors: âLLMs can operate on attention states with discontinuous position IDsâ
- mechanism: precompute marked prompt modules once; concatenate their caches
- quality moves both ways; one retrieval task drops from 7.50 to 4.25
- the paperâs own GPU speedup is stated as 8Ă in the abstract and 1.5â10Ă in the body
- InfiniGen, Lee et al., OSDI 2024, official page
- mechanism: keep the cache in CPU RAM; guess next layerâs important tokens from this layerâs inputs; prefetch only those
- authors: âthe attention inputs of consecutive attention layers are highly similar in LLMsâ
- up to 3Ă over prior offloading on one A6000; 1.52Ă slower than all-on-GPU
- approximate; accuracy âclosely matchesâ above 10% relative cache size
- FlexGen, Sheng et al., ICML 2023, paper
- mechanism: offload weights, activations, and cache across GPU, CPU, disk; linear program picks placement
- OPT-175B on one T4 at 0.69 tokens/s; batch 1 with compression gives 0.052 tokens/s
- authors: âit is possible to trade off latency for higher throughputâ; offline bulk only
- ContextPilot, Jiang et al., MLSys 2026, official abstract
- authors: âcontext alignment and de-duplication techniques to maximize KV-cache reuseâ
- inference: rearranging the prompt to raise hit rate is application-level prior art
- inference across this group
- every non-prefix method is lossy and evaluates quality on question-answering benchmarks, not on agent task completion
- HijackKV (below) turns exactly this approximation into an attack
literature: speculative decoding and kernels
- Leviathan et al., ICML 2023, paper abstract
- a small model proposes tokens; the big model verifies them in one pass, âwithout changing the distributionâ
- exactness holds in exact arithmetic; the determinism papers below show the bits still move
- SpecInfer, Miao et al., ASPLOS 2024, paper
- mechanism: draft a tree of tokens, verify the whole tree with tree attention
- authors: âSpecInferâs performance improvement over existing systems reduces as the batch size ⊠increasesâ
- greedy output is token-identical; stochastic case has a distribution proof
- Medusa, Cai et al., ICML 2024, paper
- mechanism: extra heads on the model predict several tokens ahead
- authors: âit is typically unnecessary to match the distribution of the original modelâ; âtypical acceptanceâ is lossy
- authors: âwhen the batch size exceeds 32, the speedup decreases and may even have a negative effectâ
- EAGLE, Li et al., ICML 2024, paper; EAGLE-3, paper
- mechanism: one-layer draft head that predicts from the target modelâs internal features
- EAGLE-3 authors: âEAGLE-3 improves throughput by 40% at a batch size of 64â in SGLang v0.4.4; EAGLE v1 drops to 0.99Ă there
- vLLM: 1.75Ă at batch 2 to 1.01Ă at batch 56
- authors skip quality evaluation: âTherefore, we do not evaluate generation qualityâ
- acceptance rates under batching are not reported
- FlashInfer, Ye et al., MLSys 2025, paper
- mechanism: one attention library over block-sparse KV layouts; JIT-compiled variants; CPU planner balances work per step
- authors: âbecause LLM serving requires deterministic outputs, we did not incorporate atomic aggregation in Stream-K implementationâ
- 29â69% inter-token latency reduction over SGLang v0.3.4âs Triton backend
- inference: kernel authors already treat determinism as a requirement, but only against atomics, not against batch shape
- Sereno, Xin et al., OSDI 2026, official abstract
- reuse speculative execution to let phone inference yield memory bandwidth to other apps, âwithout hardware modificationâ
- inference across this group
- speculative gains shrink toward 1Ă as batch grows; the honest baseline is a current engine at a production batch size
- both LLM-42 and Vosti exclude speculative decoding from their determinism guarantees
literature: serving agents, not requests
- InferCept, Abhyankar et al., ICML 2024, paper
- authors: recomputing contexts after tool calls âaccounts for 37-40% of total model forwarding timeâ
- mechanism: per request, keep, swap, or drop the cache during a tool call, based on estimated memory waste
- 1.6â2Ă higher sustainable request rate than vLLMâs discard-and-recompute, A100s, GPT-J-6B to Llama3-70B
- tool times in the workload span 9e-5 s to about 29 s; three of six tool types are partly estimated
- authors: âWe leave the comparison of these scheduling policies to future workâ
- Parrot, Lin et al., OSDI 2024, paper
- mechanism: the application marks prompt inputs and outputs as âSemantic Variablesâ so the server sees the call graph
- up to 11.7Ă over LangChain on vLLM for a MetaGPT multi-agent run, A100 80GB, LLaMA-13B
- authors: âParrot only supports cloud-side orchestration of LLM requests without involving dynamic control flow and native functionsâ
- its kernel splits shared-prefix tokens from private tokens; that split is one of the batch-shape dependences the determinism papers flag (our inference)
- Autellix, renamed Agentix at NSDI 2026, Luo et al., arXiv v1, NSDI page
- mechanism: treat the agent as a program; prioritize by the programâs accumulated service, not the callâs
- authors: âWe assume that the LLM invocation pattern of programs emerges only at runtimeâ
- authors: âAutellix improves throughput of programs by 4-15Ă at the same latencyâ over vLLM v0.6.1
- workloads: ShareGPT as programs, BFCL v3 tool use, LATS search on HotpotQA; 8Ă A100
- preempts with KV swaps; nothing on output equality
- KVFlow, NeurIPS 2025, abstract: evict and prefetch by a predefined agent step graph, up to 2.19Ă over SGLang
- FlashAgents, MLSys 2026 oral, page: stream tokens between agents so the next agentâs prefill overlaps the previous oneâs decode
- Murakkab, OSDI 2026, abstract: declarative agent workflows with a profile-guided choice of model and hardware
- Continuum, Li et al., full v7
- authors: âKV cache time-to-live (TTL) mechanismâ plus âProgram-level first-come-first-serve schedulingâ
- authors: âup to 8.18x improvements in both latency and throughputâ on SWE-agent workloads
- authors: âthe current design of Continnum are optimized for ReAct-style, tool-interleaving agentsâ
- authors: âspeculative branches, asynchronous multi-agent coordination, and context foldingâ are future work
- baselines include vLLM 0.10.2 and LMCache 0.3.7
- AgentReplay, Pan et al., full HTML
- authors: âFixed-trace comparisons measure the cost of executing recorded behaviorâ
- mechanism: force recorded tokens through a real engine so scheduling differences stay measurable
- rejects speculative decoding and unsupported histories
- inference: use this for measurement; do not reinvent it
- inference across this group
- agent scheduling already has seven systems and a measurement tool
- InferCept, Autellix, and Continuum all recompute, swap, or retain the cache around tool pauses; those are exactly the execution-path changes Vosti tests
- none checks that the agentâs tokens or task result survive those changes
literature: same input, different output
- this is the newest and least crowded cluster, and the one that fits the humanâs verification work
- Yuan et al., NeurIPS 2025 oral, paper
- measured: DeepSeek-R1-Distill-Qwen-7B in BF16 with greedy decoding shows âup to 9% variation in accuracy and 9,000 tokens difference in response lengthâ from GPU count, GPU type, and batch size
- FP32 is near-zero variance; BF16 is worst because of its 7-bit mantissa
- fix, LayerCast: compute in FP32, store weights in BF16, 34% less memory than full FP32
- authors on batch-invariant kernels: ârobust to continuous batchingâŠbut not to other forms of nondeterminism like changing the TP sizes or GPU typesâ
- Thinking Machines, He et al., Sep 2025, post
- authors: âIf you compose some property under which the kernel is not invariant (i.e. batch-size) with nondeterminism of that property (i.e. the load the server is under), you get a nondeterministic systemâ
- 1000 greedy completions of Qwen3-235B gave 80 unique answers; all agreed for 102 tokens
- batch-invariant kernels: 26 s baseline vs 55 s unoptimized vs 42 s with a better attention kernel
- the post is a blog, not peer reviewed; vLLM and SGLang shipped modes based on it within weeks
- SGLang deterministic mode, Sep 2025, post
- guarantee: same â(inputs, seed)â pair gives the same sample across batch sizes
- overhead 24â55% depending on backend and lengths; average â34.35%â for FlashInfer and FA3
- unsupported then: radix cache on two backends, MoE, tensor parallel above 2, speculative decoding
- LLM-42, Gond et al., SOSP 2026, paper
- authors: âLLM-42 decodes tokens using a non-deterministic fast path and enforces determinism via a lightweight verifyârollback loopâ
- idea: borrow speculative decodingâs verify step; verify under a fixed-shape reduction and roll back flips
- pay only for the traffic that asks for determinism: within 1% of baseline throughput at 10% deterministic traffic
- cost of the alternative: batch-invariant GEMM âpeaks at 194 TFLOPSâ vs cuBLAS 527
- flips are rare: 0.32% of tokens recomputed on ShareGPT, 10.97% on long arXiv prompts
- authors: âLLM-42 currently does not support sharing prefix caches across multiple turns of the same request or sharing across requestsâ
- authors: âdoes not currently integrate with speculative decodingâ
- H100-PCIe, SGLang v0.5.3rc0, Llama-3.1-8B
- MarginGate, Chu et al., May 2026, abstract
- authors: âbatch-induced token flips are sparse ⊠all tested models stay within the 0.3-1.3% rangeâ
- verify only when the top-two logits are close; cuts LLM-42âs extra latency about 2Ă
- TBIK, Zhang et al., ICML 2026 proceedings, v2 methods
- problem: training uses one GPU, serving uses tensor parallelism, so RL rollouts and trainer disagree
- mechanism: âTP-invariant matrix multiplication and reduction primitivesâ with one binary reduction tree for every TP size
- authors: âtotal BIO+TBIK overhead ranging from 22% to 63%â, Qwen3-8B on H20, TP 4
- selected full §§5.2â5.4 independently checked on 8 October 2026
- generated tokens and top-five predictive probabilities match across TP 1/2/4/8 and batch sizes 8/16/32 for three tested models
- Qwen3-32B is tested at TP 2/4/8
- these observations do not establish an engine-level proof of complete raw-logit equality for every supported transition
- end-to-end overhead measurement uses four H20 GPUs with NVLink
- attention is patched, not part of the contribution; quantization and pipeline parallel left open
- Vosti, Qin et al., Duke, Sep 2026, paper, code
- đ§ the human saved this paper on 2026-10-06, so it is already on their radar
- the definition: same model, config, prompt, and sampler state give âbitwise-identical logits at corresponding output positions across executionsâ
- authorsâ tests: âBoth vLLMâs batch-invariant mode and SGLangâs deterministic mode produce logit mismatchesâ
- vLLM 0.28.0 and SGLang 0.5.19 on one H200; Llama3.1-8B mostly passes, Gemma3-4B fails every category
- SGLang deterministic Triton: âonly 3/384 prefillâdecode pairsâ
- authorsâ caution: âa mismatch demonstrates numerical variation, not necessarily a bug for the intended use caseâ
- mechanism: engine in Rust verified with Verus; kernels in a Triton subset checked by their own Z3-based relational verifier
- âVosti chooses kernels independently of runtime engine state and ties cached KV values to their logical token prefixesâ
- the proof splits at the kernel boundary: Verus assumes kernel contracts, the kernel verifier proves them
- effort: 14,043 lines engine, 64,952 lines proof, verified in 52.9 s on 128 threads
- results: passes âall 5,488 bitwise comparisons across seven Llama and Gemma modelsâ
- decode 1.36â3.19Ă faster than vLLMâs invariant modes, prefill slower, âslower than both default modes and SGLangâs deterministic modesâ
- what it does not do, in the authorsâ words
- âVosti currently supports single-GPU executionâ
- âExtending Vosti to MoE would require batch-independent routing policiesâ
- âDispatch among bitwise-equivalent kernels remains a future extensionâ
- âVostiâs verification establishes determinism but not functional correctnessâ
- âAttention assumes input-dependent finiteness, unchecked at runtimeâ
- linear-attention models would need a new cache invariant
- authors cite DeepSeek-V4 as building end-to-end batch-invariant kernels into its training stack
- Volta, Driscoll et al., OOPSLA 2026, abstract
- âthe first equivalence checker for GPU kernelsâ; sound, and complete for a class including attention
- inference: Volta checks equivalence over real numbers, Vosti needs bitwise equality; the two do not compose yet
- DriftBench, Vitale, MLSys 2026, official abstract, slides
- â236,985 prompt-response pairs across 105 configurationsâ: 5 models, 4 GPUs (H100, H200, B200, MI300X), 3 engines, 3 precisions
- slides: answer flip rate by workload is math 16.74%, safety 7.97%, long context 1.55%, code 0.09%
- slides: Llama 3.1 8B moving from H100/FP16 to B200/FP8 flipped 124 of 520 safety prompts, â65 went safe to unsafe and 59 unsafe to safeâ
- their predictor works for unseen hardware (RÂČ 0.909) but not unseen models (0.118)
- full paper is on OpenReview behind a bot check; not read
- inference: this measures drift across stacks; Vosti measures drift inside one stack; this review did not establish a study of drift across a multi-step agent run
- inference across this cluster
- Thinking Machines addresses batch invariance; TBIK combined with batch-invariant operations demonstrates agreement across tested batch and TP sizes
- scheduling-level fixes (LLM-42, MarginGate) are cheap but give up prefix sharing
- the proof-level fix (Vosti) covers one GPU and dense models
- TBIK establishes tested cross-TP reproducibility; Vosti establishes a narrower engine proof
- this review has not established a proof covering distributed recovery, speculation, or MoE
- absence from this reading set does not establish absence from the literature
- multi-step agent outcome sensitivity remains a candidate requiring further prior-work checks
literature: shared cache as a leak
- PromptPeek, Wu et al., NDSS 2025, paper
- authors: âthe KV cache sharing may inadvertently create side channel information, which can be leveraged by the adversary to carefully craft requests ⊠thereby recovering other usersâ promptsâ
- the timing of a prefix hit tells the attacker whether their guess matched another tenantâs prompt
- HijackKV, 2026, paper
- authors: âKV tied to benign tokens may encode an attacker-controlled prefix, silently hijacking model behavior even without adversarial text in the inputâ
- about 94% targeted success against position-independent reuse; recomputation defenses like CacheBlendâs leave 79% at 50% eviction
- inference: Vostiâs rule that a cached value belongs to its exact token prefix blocks HijackKV by construction but not PromptPeekâs timing leak
- a cache contract that covers both correctness and leakage has not been stated
literature: failures and scaling
- GhostServe, Jayakody et al., MLSys 2026, official abstract
- âapplying erasure coding to generate and store the parity shards in host memoryâ so a dead GPUâs cache can be rebuilt
- RaidServe, Xu et al., MLSys 2026, official abstract
- âproactive KVCache backup and on-demand weight recoveryâ to keep tensor-parallel serving alive through a GPU failure
- ServerlessLLM, Fu et al., OSDI 2024, paper
- mechanism: tiered checkpoint cache in RAM and SSD; to move a running request, âthe source server migrates only the tokensâ and the destination rebuilds the cache by prefill
- model start 0.8 s vs 12.1 s for Ray Serve 2.7.0, OPT-6.7B
- authors: âwe do observe instability in CUDA driver callsâ
- BlitzScale, Zhang et al., OSDI 2025, paper
- mechanism: load weights over the GPU network by multicast from running instances; serve with the layers that have arrived
- 47â75% shorter first-token time than ServerlessLLM; 49% less GPU time than static vLLM and DistServe with no SLO misses
- authors: âThe policy depends heavily on workload characteristics, which we leave as future workâ
- DynamoLLM, Stojkovic et al., HPCA 2025, paper
- mechanism: every few minutes, change instance count, tensor-parallel degree, and GPU frequency per request-type pool
- authors: âconserves 53% energy and 38% operational carbon emissionsâ; 8Ă H100 servers, Llama2-70B
- changing TP degree at runtime changes accumulation order, which TBIK and Vosti both flag (our inference)
- Andes, Liu et al., paper: schedule by user-perceived reading experience; preempt by recompute or swap; vLLM v0.6.1 baseline
- VTC, Sheng et al., OSDI 2024, paper: per-client fairness with âa 2Ă tight upper bound on the service differenceâ; fairness is per client, not per agent program
- inference: every recovery, migration, and scaling paper here rebuilds or moves the cache; none asks whether the request continues with the same outputs it would have had
- ServerlessLLM and InferCept rebuild by prefill what decode produced; Vostiâs prefillâdecode tests show production engines disagree on exactly that
literature: hardware beyond GPUs
- WaferLLM, He et al., OSDI 2025, paper, code
- the device: Cerebras WSE-2, â850,000 cores with 40GB of on-chip memoryâ, mesh network, tiny local memories, 5-bit routing addresses
- the model, PLMR: massive Parallelism, non-uniform Latency, small local Memory, limited Routing
- mechanisms: MeshGEMM for prefill, MeshGEMV with K-tree allreduce for decode, cache shifting to balance cores
- headline: â10-20Ă speedups over A100 GPU clusters running SGLang and vLLMâ and â2.5Ă more energy-efficientâ
- the tables say more than the abstract
- decode, Table 8: WSE-2 reaches 2700 tokens/s per request vs 260 on 8 A100s; energy ratio A100/WSE-2 is 0.92â7.02, so the wafer wins on decode energy except against one A100
- prefill, Table 7: energy ratio A100/WSE-2 is 0.05â0.84 in every column, so the A100 uses less energy for prefill in every reported case
- inference: the 2.5Ă energy claim is an aggregate; the phase split is the honest picture
- evaluation limits in the authorsâ words
- âwe evaluate a subset of layers and scale the results proportionallyâ for CodeLLaMA-34B and QWen2-72B, which do not fit
- âTo compare against the H100 fairly, we would need access to the WSE-3 ⊠unavailable to usâ
- âMeshGEMV does not achieve the theoretical 7,000Ă improvementâ; cores âcannot fully overlap memory access and computationâ
- vendor claims for context, not evidence: Cerebras reports 2,100 tokens/s on Llama 3.2 70B and 2,500 on Llama 4 Maverick per user
- press release
- these are batch-size-1 numbers on a system drawing tens of kW; throughput per watt at production batch is not published
- Cerebrasâs own architecture material is a Hot Chips 2024 talk, slides
- âModel layers are mapped to wafer regionsâ; âEach wafer region processes 1 tokenâ
- no peer-reviewed Cerebras-authored WSE-3 inference paper was found; KV capacity, batching, and cost per token are not disclosed
- wafer-scale is now a live academic topic: ISCA 2026 has a fault-tolerant mapping paper for wafer-scale LLM inference (Cui), ASPLOS 2026 has Ouroboros, wafer-scale SRAM compute-in-memory
- DeepSeek-V3 hardware insights, Zhao et al., ISCA 2025 industry, paper
- a retrospective on co-design, not a new system; the inference numbers are analytical
- authors: âDeepSeek-V3 achieves a significant reduction in KV cache size, requiring only 70 KB per token, substantially less than LLaMA-3.1 405Bâs 516 KBâ
- authors: decode on their H800 cluster has a âtheoretical upper limit ⊠approximately 14.76 ms TPOT, equivalent to 67 tokens per secondâ set by expert-parallel communication
- authors: âthis figure is purely theoretical and does not account for the substantial drop in GPU efficiency at small batch sizesâ on the GB200 NVL72 bound
- hardware wishes: unified scale-up and scale-out fabric, hardware multicast and reduce, FP32 or configurable accumulation, system-on-wafer and DRAM-stacked accelerators
- authors: HBM grows under 50% per year while memory demand grows more than 1000% per year
- authors ask for âchecksum and silent-data-corruption diagnosticsâ; correctness is on the hardware wish list too
- MegaScale-Infer, Zhu et al., ByteDance, paper
- authors: âdisaggregates attention and FFN modules within each model layerâ
- the reason: MoE experts stay compute-bound only if many attention replicas feed them; one A100 example leaves 25% expert utilization at batch 156
- decode only; up to 1.90Ă per-GPU throughput over the best of vLLM and TensorRT-LLM; attention on H20, experts on L40S
- authors: ping-pong pipelining âdoes not reduce the per-token latency for an individual micro-batchâ
- production on about 10,000 GPUs, cost cut 1.5â2.0Ă
- Groq TSP, Abts et al., ISCA 2022, DOI, and ISCA 2020
- authors: âguaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches)â
- the chip is statically scheduled, 220 MiB SRAM each, no switches; âthe hardware is disallowed from asserting back pressureâ
- authors: âAs system size grows beyond 264 TSPs, the available global bandwidth flattens to about 14 GB/sec of global bandwidth per TSP endpointâ
- neither paper has LLM decode numbers; Groqâs tokens-per-second claims are not in them
- inference: a fully deterministic chip is the hardware answer to proposal 1âs problem; GPUs are not that
- TPU v4, Jouppi et al., ISCA 2023, paper
- a training and recommender paper; inference appears only by citation
- optical circuit switches: â<5% of system cost and <3% of system powerâ; they route around failed hosts
- 1.2â1.7Ă faster than A100 on MLPerf Training 2.0; âThe newer, 700W H100s were not availableâ for comparison
- inference across hardware
- every non-GPU paper compares against the previous GPU generation
- the two hardware properties the serving papers keep asking for are memory bandwidth per token and a fast reduce across devices
- determinism by construction exists only in Groqâs design; wafer and GPU papers do not discuss it
literature: the 2026 program sweep
- counted by a helper from OSDI, NSDI, MLSys, EuroSys, ASPLOS, ISCA 2026 and SOSP 2025 programs; titles only, approximate
- serving systems, scheduling, multi-tenancy, cold start: about 55
- KV cache, attention, long context: about 38
- inference hardware and accelerators: about 34, mostly ISCA
- edge and on-device: about 25
- MoE serving: about 20
- speculative decoding: about 19
- quantization: about 17
- agents and RAG: about 14
- fault tolerance for inference: about 5
- these are helper estimates without a published title inventory or counting procedure
- do not use them to estimate the size of correctness research or establish a research gap
- the concrete papers above, their mechanisms, and their limits support the proposals
- titles that bear on this reviewâs proposals
- MLSys 2026 âSpeculative Decoding: Performance or Illusion?â, page, unread
- MLSys 2026 âAdaptive Erasure Coding for fault-tolerant LLM servingâ, third recovery paper beside GhostServe and RaidServe
- OSDI 2026 StriaTrace, inference diagnosis, Wu
- SOSP 2025 Orthrus, silent data corruption detection, and PhoenixOS, GPU checkpoint and restore
- ISCA 2026 âFault-tolerant mapping for wafer-scale LLM inferenceâ, Cui
- MLSys 2026 ForeCache and AgenticCache, caching for coding agents
proposal 1: reproducible inference across a specified execution change
- question: can an independently checked engine contract preserve requested numerical outputs when execution changes?
- nearest priors
- TBIK demonstrates cross-TP output/probability agreement
- Vosti proves single-GPU determinism under explicit engine/kernel assumptions
- Vosti §7 motivates multi-GPU contracts
- migration/restart and speculative-protocol proofs below are agent extensions
- first measurement pilot
- pin model, engine release, kernels, precision, prompt, sampler, and random-number state
- begin with one supported change rather than the entire matrix
- compare matched clean execution with changed execution
- one-GPU pilot: supported batching and speculative-decoding configurations
- four-GPU pilot: TP 1/2/4 with TBIK enabled and disabled
- record GPU models and interconnect
- mark Vostiâs unsupported multi-GPU cells unavailable
- recovery pilot: checkpoint and restore on supported configurations
- require complete token history, sampler state, pending drafts, and cache identity
- measure raw logits, log probabilities, generated tokens, task outcomes, latency, throughput, and memory separately
- proof candidates after a measured gap
- distributed execution: extend an engine invariant across devices while respecting TBIKâs established reduction discipline
- restart/migration: state exactly which recoverable state matches uninterrupted execution
- speculation: check acceptance, rollback, numerical outputs, and sampler-state changes
- equal verification logits alone do not prove equal sampled traces
- cost comparison
- compare supported SGLang deterministic modes, TBIK, Vosti, and LLM-42 where code is accessible
- published 34% and 63% overhead figures concern different platforms and workloads
- they are reference observations, not universal acceptance thresholds
- choose a deployment-specific budget before judging a tradeoff
- evidence against the project
- closest systems already supply the precise required guarantee
- observed differences vanish under the documented supported configuration
- checking costs erase the benefit for the chosen use case
- token agreement alone does not reject a use case needing equal logits or probabilities
- feasibility limits
- Vosti artifact independently checked on 8 October 2026
- source and instructions are accessible; verification was not rerun
- multi-GPU access and artifact setup remain prerequisites
- timeline and publication prospects are unknown
- Vosti artifact independently checked on 8 October 2026
proposal 2: when simple routing loses under agent traffic
- status: fallback; its core mechanism is in Continuum and Autellix
- keep as a stress test: shared prefixes across programs, cancellations, correlated bursts, which Continuumâs per-program model may mishandle
- method: replay with AgentReplay through current vLLM and SGLang; compare least-load, cache-first, LMetric, and Continuum
- publishable only if a measured gap appears and the fix is more than a new predictor
proposal 3: recovery that preserves a complete agent execution
- status: fallback; ACRFence, Safe to Resume, and action settlement already cover external effects, per the sibling review
- candidate question: after a supported cache recovery, does the agent preserve the specified numerical outputs and sampler state?
- this is proposal 1âs restart/migration candidate seen from the application side; do them together or not at all
proposal 4: a measured boundary for wafer-scale advantages
- the question: for which request shapes does a mesh chip beat GPUs on time and on energy
- WaferLLMâs own tables show the answer is phase-dependent: decode yes, prefill no
- method: calibrate an analytical model on WaferLLMâs published cycle counts, then sweep batch, context length, and expert sparsity against H100 chunked-prefill baselines
- hard limit: no wafer access means every number is extrapolated; say so, or skip this
selection
- choose proposal 1
- conditional on a precise gap beyond TBIK and Vosti and access to the required hardware
- it uses the humanâs Verus and Rust experience
- measurement can reject the proposal before proof engineering
- the workflow pilot in the group index has fewer hardware dependencies
- proposal 2 and 3 only if their stress tests reveal a gap
- proposal 4 only as a modeling side project
gaps in this review
- DriftBench full paper and the NSDI 2026 Agentix version were not read
- venue sweep is by title; EuroSys, SOSP 2025, and MLSys lists may be incomplete
- Cerebras Hot Chips slides came through a summarizer; quotes from them are not verified against the PDF
- vLLMâs batch-invariance documentation returned HTTP 429 twice
- quotes from helper agents were not re-checked against source text
- LLM-42 code availability unknown; Vosti source/instructions independently checked, not executed
- earlier group consultation concerns cache/routing; a cross-topic update is running on 8 October 2026
related review
Last edited: