hardware and timing side channels (authored by agents unless marked đ§)
literature review and research ideas, checked 7 October 2026
- starts from Wave and Junzhou Heâs WASM timing talk in reading notes
- related security overview
side channel: a program leaks a secret through something other than its output
- e.g. how long it runs, which cache lines it touches, how much power it draws, how big its network packets are
constant-time code: code whose timing and memory addresses do not depend on secrets
hardware performance counter: a CPU or GPU register counting events
- e.g. instructions, cache misses, floating-point operations
takeaways
- the same leak is used two ways now
- attack: steal prompts, tokens, model shape
- oversight: check which model a provider really runs, as Wave does
- recommendation: compare oversight methods by what they certify and whom they trust
- published attacks demonstrate leaks at several LLM-serving layers
- network, shared caches, CPU cache, GPU cache, power
- a candidate project is testing several optimizations against one explicit rule of what observers may learn
- novelty needs checking against current defenses and leakage models
- constant-time code breaks below the source level
- compilers undo it; the WebAssembly spec does not promise a constant-time
select; coverage across current browser tiers is unclear - the repair tool from the talk, WaSCR, trusts the engine and checks its own correctness with random tests
- compilers undo it; the WebAssembly spec does not promise a constant-time
- hardware isolation has explicit limits
- new Spectre variants in 2025; Nvidia GPU partitions leak; a $1000 memory tap breaks Intel and AMD confidential VMs
- physical interposer attacks fall outside Intel and AMDâs stated threat models
- they invalidate an audit only if that audit promises protection against such physical access
- compare each auditâs attacker assumptions before drawing a conclusion
- best ideas for us, in my order; details in the last section
- test whether real WebAssembly engines keep constant-time code constant-time
- leak rules for LLM servers, plus a fuzzer that finds violations
- attack and harden counter-based model oversight (Wave)
- census of secret-dependent branches in WebAssembly crypto shipped on real websites
- branch-removal repair for WebAssembly, proved correct in Verus
seeing which model runs
- Wave, Xu et al., ASPLOS 2026
- the authorâs publication page links the paper and describes GPU performance counters
- the authorsâ artifact lists Nvidia GPUs and Nsight Compute
- existing attack: âattackers split linear layers to evade the upper-bound checkâ
- this rules out a new project whose only contribution is testing layer splitting
- human talk notes, not publication evidence
- âmeasure what model is run only using microarchitectural side channelâ
- âuse hardware counters (PMC)â
- âFLOPS/load/store scale w/ model size due to matrix mathâ
- ârecover periodicity in kernel stats: per token (large), per layer (small)â
- use: âis Modal-aaS service using smaller modelâ; âexport controlâ
- older attacks already read model shape from shared hardware
- Cache Telepathy, Yan et al., 2018
- CPU cache timing reveals matrix multiply calls and sizes
- cuts the VGG search space âfrom more than 10Âłâ” architectures to just 16â
- CSI NN, Batina et al., USENIX Security 2019
- primary author abstract: âa non-invasive and passive attackerâ
- measures timing and electromagnetic emissions from an ARM Cortex-M3 microcontroller
- infers layers, neurons, activation functions, output classes, and weights for tested perceptrons and convolutional networks
- evidence depth: author abstract, not a reproduced attack or GPU result
- Cache Telepathy, Yan et al., 2018
- 2025 to 2026 work moved to GPUs and transformers
- Energon, 2025
- GPU power and temperature readings, no physical access
- âover 89% on average for model family identification and 100% for hyperparameter classificationâ
- Behind Bars, USENIX Security 2026
- Nvidia MIG splits one GPU into isolated instances
- §4â6: âcross-instance L2 cache interference still occursâ
- Membar+Load times loads delayed by memory barriers from another instance
- kernel launches trigger those barriers on tested Hopper GPUs
- this differs from evicting a victimâs cache lines with Prime+Probe
- §6 evaluates 20 models on H100 with an unprivileged neighboring tenant
- trained classifier: mean F1 score 94.28%
- output-length experiment: 97.4% single-token accuracy
- §7.3: âwe assume a batch size of oneâ
- larger batches prevent reliably assigning each observed length to a prompt
- main attacks assume only attacker and victim workloads
- extra-workload experiments measure channel capacity, not the same fingerprinting accuracy
- §7.1 reproduces the mechanism on H200
- tested A30/A100 lack the kernel-launch trigger, so these inference attacks do not transfer directly
- Kraken, SaTML 2026
- âthe GPUâs electromagnetic radiation leaks even 100 cm away through a glass obstacleâ
- Energon, 2025
- model shape also leaks through the API, no hardware needed
- NightVision, July 2026
- §4 builds a token-probability matrix from many prompts and samples
- API exposes each sampled tokenâs log probability; no logit-bias control required
- estimated matrix rank gives hidden dimension
- §4.2 fits prefill latency against prompt length using reference models
- combines the dimension estimate with that fit to estimate depth and parameter count
- ârecovering hidden dimension to within 23% average relative errorâ
- abstract reports about 53% error for depth and parameter count above three billion parameters
- §5 tests 32 open models spanning 135Mâ30B parameters
- §6: ânontrivial estimation error at high costâ
- reducing error needs roughly 100 billion to one trillion tokens under the single-token-logprob interface
- very large sampling cost limits routine commercial-API auditing
- calibration and LLaMA-style architectural assumptions limit transfer to unknown serving stacks
- estimates are not certification of exact weights
- §4 builds a token-probability matrix from many prompts and samples
- Fingerprinting Inference Systems of LLMs, May 2026
- different engines and GPUs round numbers differently
- âthe inference engine, attention backend, and underlying hardware platform can be identified reliablyâ
- Auditing Prompt Caching, ICML 2025
- cache timing showed âOpenAIâs embedding model is a decoder-only Transformer, which was previously not publicly knownâ
- NightVision, July 2026
- audits that prove which model ran, without side channels
- Are You Getting What You Pay For?, Cai et al., 2025
- text-only tests miss quantized or mixed-in substitutes
- proposes trusted enclaves for âprovable cryptographic guarantees of model integrity with only a modest performance overheadâ
- full paper §4.1 compares identity queries, benchmarks, output-distribution tests, greedy decoding, and log probabilities
- framework, GPU, batching, and decoding differences can cause false alarms for the same weights
- §4.3 measures Llama-3-8B through vLLM on one H100 with TLS
- concurrency 1 and 64; first-token latency overhead 9.23% and 16.18%
- throughput loss 15.38% and 2.88%, respectively
- these results do not establish the same overhead for other models or multi-GPU serving
- trust requirement: auditor accepts the hardware attestation and the measured model and code
- recommendation: verify how measurements bind weights, runtime configuration, and responses before claiming end-to-end certification
- TOPLOC, ICML 2025
- provider sends a small hash of internal activations
- â258 bytes of storage per 32 new tokensâ
- needs the provider to cooperate and the auditor to hold the weights
- Bit-Exact AI Inference Verification, Cankaya, ICML 2026 workshop
- auditor recomputes outputs exactly
- âbitwise-precise re-computation does not require access to identical hardwareâ
- hardware governance taxonomy, Ansari, 2026
- âon-chip compute metering, cryptographic proof-of-training, and hardware-embedded enforcement, are also the least matureâ
- Are You Getting What You Pay For?, Cai et al., 2025
- two reasons to doubt counters and enclaves as trust anchors
- SoK on performance counters for security, Das et al., S&P 2019
- âmeasurement imprecisions or incorrect assumptions regarding the measured values can undermine the offered protectionâ
- attacker-controlled counter activity can bypass some defenses
- TEE.fail, S&P 2026
- a memory tap that âshould cost you under $1000 to buildâ
- stolen keys can compromise Nvidiaâs GPU Confidential Computing
- Intel and AMD âconsider interposer attacks to be out of scopeâ
- SoK on performance counters for security, Das et al., S&P 2019
- where Wave sits
- its public artifact already includes adversarial layer splitting
- full paper, §1
- authors: ânot to recover or certify exact model weightsâ
- current platforms lack protected tenant counter access
- proposed trusted collection belongs in the threat model
- selected full methods and evaluation, §§4â7
- authenticated counter delivery bound to a request remains assumed rather than implemented end to end
- tests use customized GPT-2, LLaMA, and Qwen templates without pretrained weights
- 25Mâ10B parameters on RTX 4090, RTX 5080, and H100
- no held-out architecture split specified
- verification requires correct architecture templates, allowed attacks, and bounded noise
- implemented upper-bound attack tests allow each operation to split at most once
- tested on one single-layer GPT-2-like model; arbitrary substitutions and fusion excluded
- collecting required counters adds 1196â5288% overhead over inference without Nsight Compute
- low-cost future collection interfaces are proposed rather than demonstrated here
- quantization, modern engine optimizations, multi-GPU execution, and advanced shared serving remain outside demonstrated scope
- inference: avoiding direct content inspection supplies no quantified privacy bound
- selected sections and tables read; artifact and soundness argument not independently reproduced
- the previous draftâs claims of no provider cooperation, no enclave and no adversarial tests were unsupported
transformed weights can leak without breaching an enclave
- Beijie Liu et al., Hiding Directions, Leaking Structure, August 2026 preprint, selected §§IV/VI/VII and appendix C
- authors: âdoes not exploit TEE vulnerabilities, side channels, or protocol deviationsâ
- ArrowCloak exposes transformed weights to an untrusted GPU while retaining secrets inside a CPU enclave
- main attack assumes the exact public checkpoint used to initialize private fine-tuning
- removes a reused mask subspace, matches weight columns, and estimates their scales
- no victim queries or private training data supplied
- six classification, segmentation, and diffusion configurations tested
- hidden correspondence recovery: 99.92â100%
- classification victim agreement: 94.39â99.54%
- diffusion uses image similarity with matched prompts and seeds
- independently pretrained GPT-2 reference yields only 0.09% correspondence recovery
- exact initialization succeeds; architecture compatibility alone is insufficient here
- raising BERT mask rank from 64 to 512 reduces functional agreement from 95.53% to 48.05%
- correspondence recovery and useful model recovery are different outcomes
- inference: this attacks exposed-weight confidentiality, not counter integrity or fully confidential GPU execution
- selected full access, methods, results, and mitigation sections read
- cryptographic critique, every theorem, and artifact not independently audited
leaks from LLM serving
- over the network, encrypted traffic still shows size and timing
- What Was Your Prompt?, 2024
- each streamed packet shows one tokenâs length
- âaccurately reconstruct[ed] 29% of an AI assistantâs responsesâ
- Whisper Leak, Microsoft, 2025
- guesses the topic of a prompt, âacross 28 popular LLMsâ
- âoften >98% AUPRCâ
- padding, batching, packet injection: ânone provides complete protectionâ
- speculative decoding leak, Wei et al.
- speculative decoding: guess several tokens, keep the right ones
- right and wrong guesses show in packet sizes
- picks the query out of 50 âwith over 75% accuracyâ
- What Was Your Prompt?, 2024
- inside the server, caches shared between users
- Early Bird, Song et al.
- a cache hit answers faster, so timing tells what others asked
- âinfer both confidential system prompts and those issued by other usersâ
- Auditing Prompt Caching
- âseven API providers, including OpenAIâ shared caches across users
- CacheProbe, 2026, asks the same of the OpenRouter gateway
- Early Bird, Song et al.
- on the same machine
- I Know What You Said, USENIX Security 2025
- token embedding lookup touches a cache line per token
- works âwithout privilegeâ; recovered text has âaverage edit distance of 5.2% and 17.3%â for output and input
- Behind Bars, above, does the GPU version
- I Know What You Said, USENIX Security 2025
- related defense that narrows the candidate gap
- KVGov, Addagada, August 2026 preprint
- isolates prefix-cache keys by principal
- author: âthe defense itself is evaluated in simulation calibrated to those measurementsâ
- measured attack timing supports channel existence, not a production defense guarantee
- I Know What You Asked
- primary NDSS paper studies prompt reconstruction through shared KV caches
- Bifrost, Chen et al., June 2026 preprint
- hybrid enclave and encrypted inference with an explicit leakage model
- author summary: âsecrets are provisioned only to an attested CPU TEEâ
- abstract separates projected large-model latency from direct small-model encrypted runs
- narrows any claim that serving systems lack leakage rules
- candidate question: do protections compose when caching, batching and speculation are combined?
- search results alone do not establish that this question is new
- serving systems are covered in systems_ml
- KVGov, Addagada, August 2026 preprint
constant-time code and WebAssembly
- WaSCR is the tool from the humanâs WASM talk note
- WaSCR, Huang, He, Wang, Wang, WWW 2025; free copy
- repairs secret-dependent branches using dependency analysis and selectors
- limits, from evaluation and §7 of the full paper
- covers instruction timing rather than all microarchitectural leaks
- checks correctness on 100,000 random inputs per sample
- dataset contains 20 samples
- measured in the WasmEdge runtime on the GEM5 simulator, not in a browser
- one sample incurs roughly 60Ă overhead
- specification: âdoes not mandate the select instruction to be translated into constant-time machine codeâ
- authors also inspect generated code and report current runtimes usually preserve this property
- calls into JavaScript inside secret branches are not handled
- checkers and typed languages came first
- CT-Wasm, POPL 2019
- a type system marks secrets; proved sound
- provides modified V8 and a rewrite tool for ordinary Wasm engines
- authors: âcode that we experimentally measure to be constant-timeâ
- full paper §6.4.3 tests CT-Wasm and rewritten ordinary Wasm cryptographic routines
- modified dudect compares zero and random keys with Welchâs t-test
- 45 million measurements, each running ten iterations
- §7 explicitly identifies future aggressive V8 optimizations as a risk
- this is existing browser measurement work and an already stated research question
- the WebAssembly group lists its âConstant Timeâ proposal as âInactiveâ in inactive-proposals.md
- current proposal status does not predict when a standard guarantee might appear
- Vivienne, 2021
- runs two copies with different secrets symbolically
- â57 real-world cryptographic implementationsâ
- Constant-Time Wasmtime, 2023
- modified Wasmtime compiles CT-Wasm to ARM and checks the machine code
- one engine, one CPU family
- Swivel, USENIX Security 2021, handles Spectre in WebAssembly
- âbetween 3.3% and 240.2%â overhead for the full fix
- CT-Wasm, POPL 2019
- repair outside WebAssembly
- Pendulum, TOSEM 2024, repairs source code
- DISARM, 2026
- balances branches using timings from the real device
- claims to beat Pendulum and DifFuzzAR on five edge devices
- ZeroLeak, 2023
- GPT-4 writes the patch, leak detectors check it
- âa mere few cents per vulnerability fixedâ
- Fall-Through Semantics, FSTTCS 2025
- language where secret branches run in constant time by definition; proved in Rocq
- compilers break constant-time code
- Fun with flags, Geimer and Maurice, 2025
- âa small set of passes are at the root of most leaksâ
- fix: âdisabling selected optimization passes via compiler flagsâ
- GCC and LLVM only; no JIT engines
- constant-time support in LLVM, Trail of Bits, Dec 2025
- adds âthe
__builtin_ct_selectfamily of intrinsicsâ - âThe Rust compiler team is exploring how to expose these intrinsicsâ
- adds âthe
- Jasminâs ML-KEM, 2025
- proves the compiler keeps the security of the ML-KEM code âused in the popular messenger Signalâ
- Rust verification is covered in formal_verification_rust
- Fun with flags, Geimer and Maurice, 2025
- the CPU under the code keeps changing
- ”Spectre, 2025
- the CPU also guesses branches inside its own microcode
- Branch Privilege Injection, USENIX Security 2025
- âAll intel processors since the 9th generationâ affected
- âleak arbitrary memory at 5.6KiB/s on an up-to-date Ubuntu 24.04â
- Revizor, Microsoft, tests CPUs against a written leak rule
- rule says what a CPU may leak; random programs hunt for more
- started at ASPLOS 2022; its paper list shows it still grows
- S&P 2026 added âtesting of cross-VM and user-kernel leaksâ
- CCS 2024 used it âto test cryptographic libraries against speculation contractsâ
- ”Spectre, 2025
research we could do
- 1 do WebAssembly engines keep constant-time code constant-time?
- candidate gap: changed browser compiler tiers and ordinary Wasm patterns beyond existing CT-Wasm experiments
- CT-Wasm already measures V8 and explicitly raises future JIT optimization risk
- WaSCR checks generated code and GEM5 simulations
- Constant-Time Wasmtime covers one engine on ARM; Fun with flags covers GCC and LLVM
- do: feed
selectpatterns and repaired code to V8, SpiderMonkey, JavaScriptCore, Wasmtime, WasmEdge, Wasmer- every compiler tier, x86 and ARM
- read the machine code; time it on real CPUs
- smallest pilot: one current V8 release on x86, one repaired crypto routine, baseline and optimized tiers
- classify secrets and public inputs explicitly; compare generated branches and memory addresses
- include a known leaking implementation to check that the timing experiment detects it
- distinguish statistical evidence from the checker/compilerâs formal guarantees
- result: counterexamples, or scoped evidence for the tested versions and patterns
- absence of a discovered leak cannot prove all programs or tiers constant-time
- then: a test suite engines can run, and a case for reviving the inactive constant-time proposal
- risk: WaSCR and CT-Wasm already inspect compilation
- continuation requires a new compiler failure or a convincing missing coverage dimension
- a clean result supports the tested cases; publishability is uncertain
- fits: the humanâs Rust and WASM background; Cranelift is Rust
- candidate gap: changed browser compiler tiers and ordinary Wasm patterns beyond existing CT-Wasm experiments
- 2 leak rules for LLM servers, and a fuzzer
- idea: copy Revizorâs method from CPUs to inference servers
- write the rule, e.g. âmay leak total output length, nothing elseâ
- generate prompt pairs equal under the rule; run both; compare timing, packet, cache, and counter traces
- targets: vLLM, llama.cpp, SGLang
- batching, prefix cache, speculative decoding, mixture-of-experts routing, early exit
- smallest pilot: one engine, two tenants, prefix caching with batching enabled and disabled
- measure observer-visible request times before adding privileged counter traces
- public output length, batch composition, and hardware load must be controlled or declared permitted leakage
- repeated independent trials must separate prompt effects from scheduler noise
- the counter tracing tool in the humanâs notes, Hiresperf, gives 10 ”s traces to compare
- why us: it is a serving systems project, not a crypto one
- I have not confirmed that nobody did this
- 3 attack and harden counter-based model oversight
- hypothesis, not a proven bound
- dummy work may imitate a larger modelâs counter signature
- optimized kernels may change the mapping from model size to counters
- inference identity cannot be inferred from total work alone
- do: build the cheating provider
- padding, quantization, batching users together, sparsity, moving work between CPU and GPU
- measure what each trick costs and whether Wave catches it
- also: who reads the counters; compare with TOPLOC and bit-exact recomputation on cost and trust
- feasibility requirement: protected counters or a trusted collector and a clearly bounded malicious provider
- tenant access to an untrusted providerâs reported counters cannot establish identity
- begin with the artifactâs supported GPU and a fixed model family and precision
- closest work: Waveâs released artifact already tests layer splitting
- first reproduce that result and read its assumptions
- mixed precision and multi-device execution are explicitly outside its demonstrated scope
- first test whether normal optimizations cause false alarms before adding malicious padding
- compare end-to-end collection and verification cost, not only solver runtime
- hypothesis, not a proven bound
- 4 census of WebAssembly crypto on real websites
- crawl top sites, keep WebAssembly modules, find the crypto ones
- run Vivienne or WaSCRâs detector; count secret-dependent branches and memory accesses
- check whether the Rust or C source was constant-time before compiling
- WaSCR used 8 real modules; Vivienne used 57 library builds; neither says what the web ships
- fits the crawling work in web_llm_detection
- risk: finding the secret inputs in a stripped binary is manual
- 5 proved branch-removal repair
- gap: WaSCRâs correctness rests on random tests
- do: write the pass in Rust; prove in Verus that output equals input behavior and has no secret-dependent branch
- smaller first step: translation validation, checking each repaired module against its original
- risk: needs a WebAssembly semantics in Verus; large
- weaker ideas
- LLM agent plus a constant-time checker to repair libraries
- ZeroLeak did the 2023 version
- re-measure whether providers fixed Whisper Leak and cache sharing
- cheap, small; CacheProbe started it
- do encrypted agent sessions reveal which tools ran?
- I did not check whether this exists
- LLM agent plus a constant-time checker to repair libraries
review status
- revised 2026-10-07 after opening Waveâs public artifact and WaSCRâs full paper
- Wave is published and already has adversarial experiments
- no universal claim that engines are untested or serving leakage has no defenses
- deeper primary reading: Behind Bars §§4â7, NightVision §§2/4â6, model-audit §§4.1/4.3, CT-Wasm §§6.4.3/7
- CSI NN remains author-abstract evidence
- no attack reproduced and no candidateâs novelty established by these readings
- ChatGPT consultation is tracked in the sweepâs review record when available
- no opinion is attributed before a response is captured
Last edited: