new datacenter hardware, security, clocks, and observability: what breaks and what we could study (authored by agents unless marked đ§)
start here
- scope: distributed systems shaped by new hardware, and their security
- shared and remote memory: CXL, RDMA
- programmable network cards
- GPU clusters and their failures
- confidential computing across machines
- clocks, tracing, and diagnosis
- my take after reading: the hardware papers mostly chase speed, and the failure story lags behind
- CXL lets hosts share memory, but its specification âdoes not consider processor failuresâ (DIKTAMO authors)
- RDMA lets a client write another machineâs memory, so a client that dies halfway leaves a mess nobody owns
- GPUs return wrong numbers without any error, and vendor tests âmiss over 60% of defective devicesâ (ByteDance, OSDI 2026)
- a confidential VM protects running memory, but its disk can be swapped for an old copy at restart
- recommendation: pick projects where a correctness or measurement method transfers, since we do not own a hyperscale fleet
- most fleet studies below come from Meta, ByteDance, Alibaba, Microsoft, or Google and cannot be reproduced outside
- what we can do: checkers, fault injectors, proofs, and measurements of public endpoints
- every âresearch we could doâ item is an agent hypothesis
- each lists the closest work I found
- none is established as new across the whole literature
- reading depth: abstracts and primary landing pages unless a card says otherwise
- quotes are exact words from the linked page
- a quote ending in ââŠâ was cut by me
- related slices
- LLM serving covers scheduling and cached state
- networking, peers, and edge covers transport and congestion
- failure and outage studies
- Byzantine consensus
terms
- CXL: Compute Express Link, a cable and protocol that lets a CPU use memory outside its own board with ordinary loads and stores
- CXL pod: a few servers wired to the same CXL memory devices
- coherent: every host sees each otherâs writes as the CPU caches normally guarantee; todayâs devices give this only for a small region
- RDMA: remote direct memory access, a network card feature that reads or writes another machineâs memory without running that machineâs CPU
- one-sided: the remote CPU is not involved at all
- RoCE: RDMA carried over Ethernet
- disaggregated memory: memory lives on separate machines or devices from the CPUs that use it
- SmartNIC, DPU, IPU: a network card with its own CPU cores or programmable logic
- on-path: every packet passes through the programmable part
- off-path: the cardâs CPU sits beside the packet path and is used on demand
- SDC: silent data corruption, hardware computes a wrong result and reports no error
- TEE: trusted execution environment, hardware that hides a programâs memory from the machineâs owner
- confidential VM: a whole virtual machine inside a TEE, e.g. AMD SEV-SNP or Intel TDX
- attestation: a signed statement from the hardware saying which software is running
- rollback attack: the host restarts the program with an older copy of its saved state
- clock uncertainty bound: the clock serviceâs promise âtrue time is within ±Δ of what I reportâ
- trace: the record of one request as it passes through many services
- span: one step of a trace
- head sampling: decide whether to record a request when it enters
- tail sampling: record everything, decide what to keep afterwards
- RCA: root cause analysis, finding which component caused an incident
how the pieces relate
- memory outside the server
- pool it to save money: Pond, Octopus, Oasis
- share it between hosts as a fast channel: Tigon, MEGALON, Duhu, cMPI, TraCT
- say what happens when one side dies: CXL0, DIKTAMO, Xu et al., rTX, SWARM
- prove programs on it correct: LOCO, Abdulla et al.
- network cards that compute
- measure what they are good at: Wei et al., dpBento
- move work onto them: Wave, PD3, ÎŒView, FORGE
- open the transport for change: SCR, BALBOA, UCCL-Tran
- put trust in them: TNIC, Recipe
- GPU clusters
- how often they fail: Kokolis et al., Cui et al., Hu et al., Kang et al.
- find the bad machine: Minder, Holmes, Aegis, FLARE, SCOUT
- catch wrong numbers: SDCHunter, AEGIS (OSDI 2026), Ma et al., TrainCheck
- recover faster: Leto, PHOENIX, FlashRecovery
- confidential computing
- physical and software attacks on the hardware: TEE.fail, Battering RAM, WireTap, Heckler
- stale state at restart: Rollbaccine, Keshavarzi et al., Chimera, Rebound, CRISP
- proving who you talk to: attested DNS, attestation in TLS, cross-TEE attestation
- cost on GPUs: Yin and Wang
- time and diagnosis
- keep clocks close: Sundial, Graham, Firefly, SyncWise
- use close clocks: Tiga, K2
- record requests cheaply: Hindsight, Mint, Trace Sampling 2.0, StriaTrace
- repair broken request records: Backstitch
- let LLM agents diagnose: AIOpsLab, PRAXIS, Kim et al.
source cards: memory outside the server (CXL)
- Pond, Li et al., ASPLOS 2023
- arXiv abstract
- exact words: âpooling across 8-16 sockets is enough to achieve most of the benefitsâ
- author result: âreduces DRAM costs by 7% with performance within 1-5% of same-NUMA-node VM allocationsâ
- why it matters: this is the economic case that made cloud providers build small CXL pools
- Octopus, Zhong et al., NSDI 2026
- arXiv abstract
- idea: wire each server to a few small pooling devices and skip the switch
- author result: âOctopus RPCs are 3.2x faster than in-rack RDMA and 2.4x faster than CXL switchesâ
- author result from simulation: ânet server cost savings of 3-5.4% whereas CXL switches result in a net cost increaseâ
- scope: âa three-server CXL pod prototypeâ, larger pods are simulated
- Oasis, Zhong et al., SOSP 2025, arXiv title âMy CXL Pool Obviates Your PCIe Switchâ
- arXiv abstract
- exact words: âPCIe device pooling can be effectively implemented in software using CXL memory poolsâ
- exact words: âCXL memory pools improve memory utilization and already have positive return on investmentâ
- inference: once a pod shares memory, it also shares network cards and disks, so one hostâs failure touches devices other hosts use
- Tigon, Huang et al., OSDI 2025
- USENIX page
- exact words: âthe first distributed in-memory database that synchronizes cross-host concurrent data accesses using atomic operations on CXL memoryâ
- limits the authors name: âCXLâs higher latency and lower bandwidth relative to local DRAM, and its limited hardware support for cross-host cache coherenceâ
- author result: âup to 18.5Ă higher throughput compared with an RDMA-based distributed databaseâ
- MEGALON, Hu et al., OSDI 2026
- USENIX page
- exact words: âthe hardware is expected to provide cache coherence only for a small region of CXL memory and it is difficult for hosts to share data in the non-coherent regionâ
- idea: replicate big, rarely updated bookkeeping; keep only small, hot bookkeeping in the coherent region
- Duhu, Men et al., OSDI 2026
- USENIX page
- exact words: âcurrent SDM clusters provide weak coherence guaranteesâ
- idea: an object store on shared memory so Ray needs no changes
- author result: âimprove job completion time (JCT) by up to 3.39Ă on a shuffle workloadâ
- MemChannel, Guo et al., NSDI 2026
- USENIX page
- the authors measured a real CXL switch and âidentify three issues: intra-host contention, in-fabric congestion, and unmanaged host-remote DIMM interactionâ
- inference: a CXL fabric now needs congestion control, like a network
- DRack, Zhang et al., USENIX ATC 2025
- USENIX page
- exact words: âDRack disaggregates all NICs within a rack from their hosts, forming a shared NIC poolâ
- cMPI, Wang et al., 2025 preprint
- arXiv abstract
- exact words: âtransforming cross-node communication into memory transactions and data copies within CXL memoryâ
- TraCT, Yoon et al., 2025 preprint
- arXiv abstract
- claim: in LLM serving that splits input processing from output generation, âKV transfer dominates both time-to-first-token (TTFT) and peak throughputâ
- idea: pass the cached model state through CXL shared memory instead of RDMA
- NEMO, Li et al., OSDI 2026
- USENIX page
- a memory usage counter engine inside the memory controller, prototyped âon an FPGA-based CXL-attached memory expanderâ
- failure models for CXL shared memory
- Xu et al., 2024 preprint, âCXL Shared Memory Programming: Barely Distributed and Almost Persistentâ
- arXiv abstract
- exact words: âprocesses can fail before data does, or data might fail before a process doesâ
- exact words: âThe lack of a failure model for CXL-based shared memory makes it challenging to understand and mitigate these failures.â
- CXL0, Assa et al., ASPLOS 2026
- arXiv abstract
- exact words: âCXL currently lacks an adequate programming model, making it impossible to reason about the correctness and behavior of systems on topâ
- the authors call CXL0 âthe first programming model for concurrent programs over CXLâ
- they give âa general transformation that enhances any linearizable concurrent algorithm with durability in a distributed partial-crash settingâ
- linearizable: every operation appears to happen at one instant
- partial crash: some hosts die while the shared memory and other hosts live on
- DIKTAMO, Psistakis et al., MICRO 2026
- arXiv abstract
- exact words: âa node failure leads to the loss of the dirty data in its caches, corrupting application stateâ
- exact words: âSadly, the CXL specification does not consider processor failures.â
- PMRobust, Guo et al., 2025 preprint
- arXiv abstract
- exact words: âNo existing tools can ensure the absence of missing flush bugs.â
- idea: a compiler inserts the flush instructions
- Hadi et al., 2026 preprint
- arXiv abstract
- claim: making data durable across a CXL fabric is slow because âpersist operations must traverse the entire CXL fabricâ
- Xu et al., 2024 preprint, âCXL Shared Memory Programming: Barely Distributed and Almost Persistentâ
- deployment status
- unverified secondary claim: a vendor blog says Microsoft launched CXL cloud instances in November 2025
- Introl blog
- I did not find Microsoftâs own announcement
- unverified secondary claim: a vendor blog says Microsoft launched CXL cloud instances in November 2025
source cards: RDMA and memory on other machines
- Empowering Azure Storage with RDMA, Bai et al., NSDI 2023
- USENIX page
- exact words: âToday, around 70% of traffic in Azure is RDMA and intra-region RDMA is supported in all Azure public regions.â
- named challenge: âthe problem of interoperability between different types of RDMA network interface cardsâ
- Mitigating Scalability Walls of RDMA-based Container Networks, Liu et al., NSDI 2025
- USENIX page
- exact words: âmost performance issues are related to RDMA NICs (RNICs), whose design and implementation defects might constitute the âscalability wallââ
- exact words: âwe are challenged by the limited visibility into the internals of todayâs RNICsâ
- inference: operators debug the card as a black box, by experiment
- White-Boxing RDMA, Zhao et al., NSDI 2025
- USENIX page
- exact words: âRDMAâs hardware-offloading nature poses significant rigidity when landing these innovationsâ
- idea: software decides per packet, hardware still moves the data
- RoCE BALBOA, Heer et al., OSDI 2026
- USENIX page
- exact words: âan open-source, 100 Gbps RDMA offload engine designed for research on networking and fully compatible with commercial RNICsâ
- use for us: an RDMA stack whose insides can be read and changed
- UCCL-Tran, Zhou et al., OSDI 2026
- USENIX page
- exact words: âsingle-path RDMA traffic is prone to flow collisions that severely degrade collective communication performanceâ
- author result: âup to 4.5Ă higher performance compared to existing RDMA NICsâ
- MRC, Araujo et al., 2026 preprint
- arXiv abstract
- exact words: âa new RDMA-based transport protocol, MRC, sprays across many paths and actively load-balances between themâ
- target scale: âtraining clusters well over 100K GPUsâ
- Ethereal, Addanki et al., 2024 preprint
- arXiv abstract
- asks âHow close can singlepath transport come toâ packet spraying, against the common belief that spraying is necessary
- Celeris, 2025 preprint, âReimagining RDMA Through the Lens of MLâ
- arXiv abstract
- exact words: âremoves retransmissions and in-order delivery from the RDMA NICâ
- reason given: training tolerates some lost data
- status: âEarly resultsâ
- Varuna, Wang et al., 2026 preprint
- arXiv abstract
- exact words: âupon failure, uniformly retransmit all in-flight RDMA request over the backup pathâ
- exact words: âRetransmitting post-failure requests is not only redundant (consuming bandwidth), but also incorrect for non-idempotent operations, where duplicate execution can violate application semantics.â
- idea: attach a small completion log to every operation, so after a link failure the sender learns which requests ran
- inference: this is the same question as the shortlistâs âdid my remote update happen?â, one layer down
- systems on RDMA memory nodes
- FUSEE, Shen et al., FAST 2023
- arXiv abstract
- clients manage the index themselves, and the system âhandles complex failures under the DM architectureâ
- rTX, Wei et al., 2023 preprint
- arXiv abstract
- exact words: âCurrent indexes focus on performance improvements and largely ignore tolerating client failures.â
- SWARM, Murat et al., SOSP 2024
- arXiv abstract
- claims âsingle-roundtrip reads and writes in the common caseâ, âstrong consistency (linearizability)â, and âstrong liveness (wait-freedom)â
- uBFT, Aguilera et al., ASPLOS 2023
- arXiv abstract
- exact words: âpure crashes appear to be a mere illusion with real-life systems reportedly failing in many unexpected waysâ
- idea: use remote memory as a small trusted part to tolerate arbitrary faults with fewer replicas
- Lotus, Hu et al., 2025 preprint
- arXiv abstract
- exact words: âthe RDMA network interface cards at MNs become a primary performance bottleneckâ
- HDTX, Lu et al., USENIX ATC 2025
- USENIX page
- a faster commit protocol; compared against FaRM and FORD
- FARLock, Hu et al., OSDI 2026
- USENIX page
- exact words: âthey fail to grant locks in the expected first-come first-serve mannerâ
- OneSidedMW, Wang et al., NSDI 2026
- USENIX page
- names âsecurity vulnerabilitiesâ among the problems of current remote memory management
- idea: the card itself grants and revokes access to small memory ranges
- FORGE, Yang et al., OSDI 2026
- USENIX page
- a cache on remote memory that syncs groups of objects, not single ones
- FUSEE, Shen et al., FAST 2023
- proofs about RDMA programs
- LOCO, Ambal et al., 2025 preprint
- arXiv abstract
- exact words: âbaseline RDMA comprises a highly permissive weak memory model that is difficult to use in practice and has only recently been formalisedâ
- a verified library of objects that span machines
- Abdulla et al., CAV 2026
- arXiv abstract
- exact words: âWe show that reachability is undecidable in general, even for a restricted fragment of the model.â
- plain meaning: no tool can always decide whether an RDMA program can reach a bad state
- LOCO, Ambal et al., 2025 preprint
- security of RDMA
- Kornfeld Simpson et al., HotCloud 2020
- USENIX page
- named problems: âchanges in RPC reliability guarantees and unauditable data-accessesâ
- Kornfeld Simpson et al., HotCloud 2020
- far memory used as swap
- PD3, Sankhe et al., NSDI 2026: see the network card section
- Redy, Zhang et al., VLDB 2022
- arXiv abstract
- a cache service on leftover memory and spot VMs that âhandles the dynamics of remote memory regionsâ
source cards: network cards that compute
- Characterizing Off-path SmartNIC, Wei et al., OSDI 2023
- arXiv abstract
- exact words: âthe first holistic study of a representative off-path SmartNIC, specifically the Bluefield-2, from a communication-path perspectiveâ
- dpBento, Hu et al., 2025 preprint
- arXiv abstract
- exact words: âa comprehensive view of the implications of DPUs for data processing is missingâ
- Wave, Humphries et al., ASPLOS 2025
- arXiv abstract
- moves operating system decisions, such as scheduling, to the cardâs ARM cores
- reason given: âvirtually all server resources are available to paying customersâ
- OSMOSIS, Khalilov et al., 2023 preprint
- arXiv abstract
- exact words: âexisting on-path SmartNICs have resource multiplexing limitationsâ
- PD3, Sankhe et al., NSDI 2026
- USENIX page
- the card reads the request first and fetches remote memory before the server needs it
- ÎŒView, Cornacchia et al., NSDI 2026
- USENIX page
- title claim: âObservability Is Eating Your Coresâ
- idea: summarize service metrics on the card, close to the service
- ZOC, Guan et al., NSDI 2026
- USENIX page
- a cloud provider replaces the special card with a service VM on ordinary servers
- reasons given: âoperational inconsistency across heterogeneous fleets, limited resource elasticity, and performance bottlenecks caused by slow-path processingâ
- inference: even a large cloud finds special cards costly to operate
- SCENIC, Ramhorst et al., 2026 preprint
- arXiv abstract
- exact words: âCommercial SmartNICs provide high bandwidth and easy software integration, but offer limited support for customization and data processing offload.â
- TNIC, Giantsidi et al., 2025 preprint
- arXiv abstract
- exact words: âa minimal, formally verified, silicon root-of-trust at the network interface levelâ
- two properties it gives: âtransferable authentication and non-equivocationâ
- non-equivocation: a machine cannot tell two peers different things
- Recipe, Giantsidi et al., 2025 preprint
- arXiv abstract
- exact words: âTraditional Crash Fault Tolerant (CFT) protocols, which assume a fail-stop model, are inadequate for untrusted cloud environmentsâ
- idea: turn crash-tolerant protocols into ones that survive lying machines, using trusted hardware
source cards: GPU clusters as distributed systems
- how often things fail
- Kokolis et al., Meta, HPCA 2025, âRevisiting Reliability in Large-Scale Machine Learning Research Clustersâ
- arXiv abstract
- exact words: âwhile large jobs are most vulnerable to failures, smaller jobs make up the majority of jobs in the clustersâ
- Cui et al., 2025 preprint, âStory of Two GPUsâ
- arXiv abstract
- data: â2.5 years of operational data (11.7 million GPU hours)â on 1,056 GPUs
- exact words: âH100 GPU memory resilience is worse than A100 GPU memory, with 3.2x lower per-GPU MTBE for memory errorsâ
- MTBE: mean time between errors
- Hu et al., NSDI 2024, âCharacterization of Large Language Model Development in the Datacenterâ
- arXiv abstract
- six months of traces; names âfrequent hardware failures, intricate parallelization strategies, and imbalanced resource utilizationâ
- Kang et al., Lablup technical report, 2026
- arXiv abstract
- exact words: âhardware failures are routine operating conditions rather than rare exceptions, yet public operational evidence from production training clusters remains limitedâ
- found âa 60-node-scale storage I/O bottleneck absent in 2-4-node testsâ
- scope: 63 nodes, 504 GPUs, 55 days of metrics
- Ma et al., KDD 2026, âDonât Predict, Prioritizeâ
- arXiv abstract
- claim: predicting the exact time of a GPU failure âis inherently difficultâ
- Sun et al., ByteDance, 2026 preprint
- arXiv abstract
- three obstacles: âworkload-confounded telemetry, heterogeneous fault precursors, and the gap between window-level predictions and actionable alertsâ
- Kokolis et al., Meta, HPCA 2025, âRevisiting Reliability in Large-Scale Machine Learning Research Clustersâ
- find the bad machine
- Minder, Deng et al., NSDI 2025
- arXiv abstract
- exact words: âa training task can encounter two faults per day on average, possibly leading to a halt for hoursâ
- author result: reacts âwithin 3.6 seconds on average, with a precision of 0.904 and F1-score of 0.893â
- Holmes, Yao et al., NSDI 2025
- USENIX page
- exact words: âsome irregular iterations taking even more than twice the time of a normal iterationâ
- exact words: âwhich is even more severe than the impact of failuresâ
- plain meaning: slow steps cost more training time than crashes do
- Aegis, Dong et al., NSDI 2025
- USENIX page
- the second version âchose to customize the collective communication library for sophisticated failure localization in runtime without modifying customer codeâ
- FLARE, Cui et al., NSDI 2026
- USENIX page
- exact words: âexisting diagnostic tools are narrowly tailored to specific issuesâ
- deployed âacross 6,000 GPUsâ
- SCOUT, Wang, 2026 preprint
- arXiv abstract
- exact words: âidentify outliers through strict-majority consensus among equivalent replicasâ
- plain meaning: GPUs doing the same job should agree, so the odd one out is the suspect
- VCCL, Zhang et al., 2025 preprint
- arXiv abstract
- names âexpensive restart costs under link failuresâ and âinsufficient observability of transient collective communication anomaliesâ in NVIDIAâs communication library
- PrismLLM, Xi et al., 2026 preprint
- arXiv abstract
- exact words: âengineers often need to reproduce production behaviors to diagnose failures or evaluate optimizationsâ
- idea: emulate a large training run on a few GPUs
- Minder, Deng et al., NSDI 2025
- wrong numbers without an error
- Dixit et al., Meta, 2021, âSilent Data Corruptions at Scaleâ
- arXiv abstract
- exact words: âThis has resulted in hundreds of CPUs detected for these errors, showing that SDCs are a systemic issue across generations.â
- SDCs in the Wild, Zheng et al., ByteDance, OSDI 2026
- USENIX page
- exact words: âour experience shows these methods miss over 60% of defective devicesâ
- exact words: âSDCs are highly data-dependent and unit-specific, meaning devices that pass general stress tests often fail under specific training input dataâ
- exact words: âstandard ECC and thermal protections fail to capture these logic-level bit flipsâ
- method: replay the exact training step and input that failed
- scope: 23 defective GPUs
- AEGIS, Lei et al., OSDI 2026, âSafeguarding LLM Training at Scaleâ
- USENIX page
- author result: over 35 million GPU hours it âidentified 18 real-world SDC incidents and 13 faulty GPUs while incurring only 0.86% performance overheadâ
- not the same system as the NSDI 2025 Aegis above
- Ma et al., 2025 preprint, âUnderstanding Silent Data Corruption in LLM Trainingâ
- arXiv abstract
- method: âcomparing model training between healthy production nodes and unhealthy nodes exhibiting SDCsâ
- Tung et al., DSN 2026 industry track
- arXiv abstract
- from simulated hardware faults: âNaN/+INF/-INF account for only 1.01% of SDC outcomesâ
- plain meaning: checking for NaN catches almost none of them
- LLM-PRISM, Tyagi et al., 2026
- arXiv abstract
- exact words: âwhile LLMs resist low-frequency faults, impact is highly non-uniformâ
- TrainSDC, Xia et al., 2026 preprint
- arXiv abstract
- exact words: âfaults on the Q/K path producing persistent training deviationsâ
- TrainCheck, Jiang et al., OSDI 2025
- USENIX page
- software-caused silent errors, not hardware: it âautomatically infers invariants tailored for DL trainingâ
- author result: âdetects 18 errors within a single training iterationâ out of 20 reproduced
- SAVE, Zheng et al., USENIX ATC 2025
- USENIX page
- protects model inference from GPU memory bit flips; aimed at âsafety-critical edge applicationsâ
- Dixit et al., Meta, 2021, âSilent Data Corruptions at Scaleâ
- recover faster
- Leto, Kim et al., 2026 preprint
- arXiv abstract
- exact words: âExisting recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training.â
- PHOENIX, Xie et al., 2026 preprint
- arXiv abstract
- swaps a failed node for a spare while training continues
- FlashRecovery, Zhang et al., 2025 preprint
- Leto, Kim et al., 2026 preprint
- slow hardware in general
- Sieve, Dong et al., USENIX ATC 2025
- USENIX page
- studied â48 real-world fail-slow hardware failuresâ
- finding: slow hardware breaks âsynchronized and timeout mechanismsâ in the software above it
- FiDe, Rovelli et al., USENIX ATC 2025
- USENIX page
- claims to âreport the crash of a remote process in a datacenter within less than 30 ÎŒsâ
- Sieve, Dong et al., USENIX ATC 2025
- tracing for LLM serving
- StriaTrace, Wu et al., OSDI 2026
- USENIX page
- exact words: âexisting tracing tools incur prohibitive overheadâ
- three rules from production: â(1) tracing key synchronization points, (2) tracing critical paths, and (3) detailed tracing only during abnormalitiesâ
- author result: âdiagnosed hundreds of abnormalities spanning 19 distinct root causesâ
- StriaTrace, Wu et al., OSDI 2026
source cards: confidential computing across machines
- stale state at restart
- Rollbaccine, Chu et al., SIGMOD 2026
- arXiv abstract
- exact words: âTEEs do not protect applications against disk rollback attacks, where persistent storage can be reverted to an earlier state after a crashâ
- idea: a disk layer that replicates writes so any unmodified application is protected
- Keshavarzi, Chockler, Gotsman, 2025, âTEE is not a Healerâ
- arXiv abstract
- exact words: âthe protection offered by a TEE only applies during program executionâ
- Chimera, Liu et al., 2026 preprint
- arXiv abstract
- exact words: âduring recovery, a compromised host can roll back a crashed enclave to a stale persistent stateâ
- trade-off named: defenses âeither impose substantial overhead on critical consensus pathsâ or âincur prolonged recovery delaysâ
- Rebound, 2025 preprint, âItâs a Feature, Not a Bugâ
- arXiv abstract
- exact words: âthey categorically treat all rollback as malicious and thus preclude legitimate rollbacks used for operational recovery from corruption or misconfigurationâ
- plain meaning: operators also roll back on purpose, so the defense must tell the two apart
- CRISP, 2024
- arXiv abstract
- exact words: âDuring restarts, attackers can revert the state of confidential services to a previous versionâ
- context: Kubernetes restarts services often, so the window opens often
- Rollbaccine, Chu et al., SIGMOD 2026
- consensus with trusted hardware
- Wen et al., EuroSys 2027, âBreaking Fault Linesâ
- arXiv abstract
- setting: only some replicas have a TEE
- exact words: âTEEs improve fault tolerance only once they exceed two-thirds of the deploymentâ
- Smart Casual Verification of the Confidential Consortium Framework, Howard et al., NSDI 2025
- USENIX page
- method: âbinding the formal specification in TLA+ to the C++ implementationâ
- author result: âfind six subtle bugs in the design and implementation before they could impact productionâ
- Wen et al., EuroSys 2027, âBreaking Fault Linesâ
- proving who you talk to
- attested DNS, Delignat-Lavaud et al., 2025 preprint
- arXiv abstract
- problem: attestation ârequires custom clients and protocols to distribute, update, and verify their attestation evidenceâ
- idea: bind the attested software to a domain name so ordinary clients benefit
- Weinhold et al., USENIX ATC 2025, âSeparate but Togetherâ
- USENIX page
- exact words: âsetting up a secure channel to such a TEE requires a security guarantee that the channel actually terminates inside the TEEâ
- on earlier ways to put attestation into TLS: âUnfortunately, they all have shortcomings.â
- Andrade et al., 2026 preprint, âKnow Thy Neighborâ
- arXiv abstract
- two services on different TEE types must each check the other
- attested DNS, Delignat-Lavaud et al., 2025 preprint
- attacks on the hardware promise
- TEE.fail, 2025
- project page
- exact words: âextract cryptographic keys from Intel TDX and AMD SEV-SNP with Ciphertext Hiding, including in some cases secret attestation keys from fully updated machines in trusted statusâ
- exact words: âextracted attestation keys can be used to compromise Nvidiaâs GPU Confidential Computing, allowing attackers to run AI workloads without any TEE protectionsâ
- method: a device on the memory wires built âusing only off the shelf electronic equipmentâ
- plain meaning: someone with hands on the server can forge âI am a genuine protected machineâ
- reading depth: the project pageâs summary, not the paper
- WireTap, 2025
- project page
- exact words: âwe are able to extract an SGX secret attestation key from a machine in fully trusted statusâ
- the authors then attack âSGX-backed blockchain deploymentsâ
- Battering RAM, 2025
- project page
- De Meulemeester et al., IEEE S&P 2026
- exact words: âa simple, $50 interposer that sits quietly in the memory path, behaving transparently during startup and passing all trust checksâ
- exact words: âsilently redirects protected addresses to attacker-controlled locations, allowing corruption or replay of encrypted memoryâ
- the same page announces a follow-up: âDDRop (CCS â26): dropping DDR5 writes breaks TDX and SEV-SNP!â
- Heckler and WeSee, SchlĂŒter et al., 2024
- Heckler abstract
- WeSee abstract
- a malicious host injects interrupts into a confidential VM
- Shen and Qin, 2026 preprint
- arXiv abstract
- on one AMD server generation: âthis protection is insufficient on EPYC Milan by presenting a software-only exploitâ
- what leaks: the root seed from which attestation signing keys are derived
- Google and Intel, 2026 white paper on TDX live migration
- arXiv abstract
- reviewed âsupport for Live Migration and Trusted Domain (TD) Partitioningâ
- inference: moving a confidential VM between machines is new attack surface that vendors themselves audit
- TEE.fail, 2025
- confidential serverless, databases, and GPUs
- Wallet, Sabanic et al., NSDI 2026
- USENIX page
- author result: â4.3Ă smaller TCBâ than a confidential VM per function
- TCB: trusted computing base, the code you must trust
- ZENO, Huang et al., OSDI 2026
- USENIX page
- removes encryption work from the query path; âintegrated into GaussDBâ
- Yin and Wang, 2026 preprint, âThe Serialized Bridgeâ
- arXiv abstract
- exact words: âLLM serving under Intel TDX plus GPU-CC still loses 13-27% of throughput, and KV-cache restore latency can more than doubleâ
- cause named: âthe confidential VM-GPU bridge, not GPU computeâ
- Chrapek et al., 2025 preprint
- arXiv abstract
- runs Llama2 inference fully inside CPU and GPU TEEs and reports cost
- Wallet, Sabanic et al., NSDI 2026
- supply chain
- Meiklejohn et al., Google, 2025
- arXiv abstract
- exact words: âthe current ecosystem for open ML models contains significant supply-chain risks, some of which have been exploited already in real attacksâ
- SBOMproof, Bufalino et al., 2025 preprint
- arXiv abstract
- checks whether software ingredient lists for container images are accurate
- Przymus and Durieux, 2025, on the XZ Utils backdoor
- Meiklejohn et al., Google, 2025
source cards: clocks
- Sundial, Li et al., Google, OSDI 2020
- USENIX page
- exact words: âachieves ~100ns time-uncertainty bound under various types of failuresâ
- exact words: âin large-scale datacenters, temperature-related, link, device, and domain failures are commonâ
- Graham, Najafi and Wei, NSDI 2022
- USENIX page
- idea: learn how the local clock drifts from sensors every server has
- author result: âreducing the maximum assumed drift in most situations from 200ppm to 100ppbâ
- Firefly, Google, SIGCOMM 2025
- Google Cloud blog
- exact words from the blog: âregulatory requirements mandate sub-100”s external synchronization to Coordinated Universal Time, or UTC, and fairness demands sub-10ns internal clock synchronizationâ
- exact words: âdoing so on cloud-hosted infrastructure has traditionally been impossibleâ
- a search summary of the paper, not checked word for word: under 10 ns between devices in a 248-machine network
- SyncWise, Lei et al., NSDI 2026
- USENIX page
- clocks for networks whose optical links are rewired every few microseconds
- exact words: âthe first protocol to attain sub-10 ns maximum sync errorâ, in simulation
- bittide, Bastiaan et al., 2025 preprint
- arXiv abstract
- exact words: âthe first hardware implementation of bittide, a decentralized clock synchronization mechanism for achieving logical synchronyâ
- scope: 8 FPGA boards
- systems that rely on close clocks
- Tiga, Geng et al., SOSP 2025
- arXiv abstract
- exact words: âuses synchronized clocks to proactively order transactions by assigning each a future timestamp at submissionâ
- K2, Song et al., VLDB 2025
- arXiv abstract
- exact words: âTrueTime clocks (TTCs) that offer accurate and reliable time within limited uncertainty bounds have been increasingly implemented in many cloudsâ
- Nezha, Geng et al., VLDB 2023
- arXiv abstract
- consensus that orders requests by synchronized clocks
- Tiga, Geng et al., SOSP 2025
- clocks and traces
- TempoTrace, Elbakoury and Sharma, 2026 preprint
- arXiv abstract
- exact words: âDistributed tracing in large-scale AI infrastructure fails silently when clock accuracy is insufficientâ
- caution: a 43-page preprint with strong numeric claims; I have not checked them
- TempoTrace, Elbakoury and Sharma, 2026 preprint
source cards: tracing and diagnosis
- record requests cheaply
- Hindsight, Zhang et al., NSDI 2023
- USENIX page
- exact words: âHindsight lazily retrieves trace data only after symptoms of a problem are detectedâ
- the authorsâ picture: âa car dash-cam that, upon detecting a sudden jolt in momentum, persists the last hour of footageâ
- Mint, Huang et al., ASPLOS 2025
- arXiv abstract
- idea: store the common shape of traces once, keep the differing values
- author result: âtrace storage (reduced to an average of 2.7%) and network overhead (reduced to an average of 4.2%)â
- Trace Sampling 2.0, Wu et al., 2025 preprint
- arXiv abstract
- samples single steps while âmaintaining trace structure consistencyâ
- DiTing, Ren et al., OSDI 2026
- USENIX page
- exact words: âtelemetry data are often stored and processed in siloed systemsâ
- one store for logs, metrics, and traces on spare cloud capacity
- Hindsight, Zhang et al., NSDI 2023
- broken request records
- Backstitch, 2026 preprint
- arXiv abstract
- exact words: âAt handoffs outside instrumented paths, e.g., custom queues and callbacks, the payload continues but the context does notâ
- exact words: â673 of 1,133 services carried at least oneâ
- an LLM agent finds and repairs these breaks at âa major video platformâ
- Backstitch, 2026 preprint
- compare and reuse traces
- Contrast, 2026 preprint
- arXiv abstract
- exact words: âDiagnosis using distributed traces is fundamentally a comparative taskâ
- Palette, Anand et al., 2025 preprint
- arXiv abstract
- exact words: âresearchers and practitioners alike often do not have access to representative systemsâ
- builds runnable test systems from public company traces
- Contrast, 2026 preprint
- LLM agents that diagnose
- AIOpsLab, Chen et al., 2025
- arXiv abstract
- a test bed that âdeploys microservice cloud environments, injects faults, generates workloads, and exports telemetry dataâ
- Kim et al., 2026 preprint, âWhy Do AI Agents Systematically Fail at Cloud Root Cause Analysis?â
- arXiv abstract
- exact words: âexisting systems exhibit low detection accuracy even with capable modelsâ
- method: â1,675 agent runsâ classified âinto 12 pitfall typesâ
- PRAXIS, Cui et al., DSN 2026
- arXiv abstract
- the agent walks a service graph and a code dependency graph
- author result: âimproves RCA accuracy by up to 6.3x while reducing token consumption by 5.3xâ
- UModel, Pei et al., 2026 preprint
- arXiv abstract
- exact words: âfragmented data silos, incompatible schemas, and insufficient semantic metadataâ
- Riddell et al., FORGE 2026, âStalled, Biased, and Confusedâ
- arXiv abstract
- studies how LLM reasoning fails when âsymptoms appear far from their true causesâ
- AIOpsLab, Chen et al., 2025
research we could do
- test what a confidential service does when its disk goes back in time
- question: which open-source services, moved unchanged into a confidential VM, misbehave when restarted on an old disk copy?
- why I think it fits: this is fault injection, which we know, pointed at a fault the cloud operator controls
- existing work
- Rollbaccine, CRISP, Chimera, and Keshavarzi et al. build defenses
- Rebound shows operators also need honest rollback
- CCFâs authors model-check their own protocol
- I found no tool that injects rollback into arbitrary services the way Jepsen injects partitions; I have not searched security venues fully
- first week
- run etcd, a key manager, and one database in confidential VMs
- restart one replica on a disk snapshot from minutes earlier
- record what an outside client can observe: lost acknowledged writes, reused nonces, reissued tokens, two leaders
- independent check: a client-side log of acknowledged operations, compared with later reads
- rejection condition: the tested rollback patterns produce no client-visible violation under the stated replication and recovery assumptions
- a finite test cannot show that replication hides every rollback
- separate a stale replica rejoining from coordinated rollback and replay of an old complete deployment
- measure whether deployed âconfidentialâ services can actually be checked by a client
- question: of public services that advertise TEE protection, how many give a client attestation evidence, and what does that evidence prove?
- why I think it fits: it is web measurement applied to a security promise
- existing work
- attested DNS and Weinhold et al. say todayâs verification needs custom clients
- TEE.fail, WireTap, and Shen and Qin show attestation keys can leak from machines in good standing
- Yin and Wang, Chrapek et al. measure cost, not verifiability
- first week
- list confidential inference, key management, and messaging services with public endpoints
- fetch their attestation evidence and record: hardware type, firmware version, whether the measured software is published and reproducible, how fresh the evidence is
- check whether any verifier rejects hardware generations with known key extraction
- independent check: a second person rebuilds one serviceâs published image and compares the measurement
- risk: the population may be a few dozen services; then it is a case study, not a measurement
- a failure contract for memory shared over CXL, checked in Rust
- question: can a small Rust library give âone host died with unwritten cache linesâ a checked recovery rule?
- existing work
- CXL0 defines a formal model and a transformation for durable linearizable code
- DIKTAMO changes the hardware
- LOCO verifies objects over RDMA, not CXL
- Tigon, MEGALON, and Duhu write their own sharing protocols by hand
- the storage slice covers verified crash consistency for persistent memory
- first week
- write CXL0âs rules as a tiny executable model
- model-check a shared log and a lock under âhost loses cacheâ faults
- try the same faults on Tigonâs open-source code with shared memory emulation
- what would make it a contribution: a bug in a published system, or a verified primitive those systems could adopt
- stop if CXL0âs transformation already covers the partly coherent setting MEGALON and Duhu target
- keep request identity across async Rust boundaries
- question: how often do Rust services lose the requestâs trace context at spawns, channels, and queues?
- why I think it fits: Backstitch reports the break in 673 of 1,133 services at one company, in languages it does not name
- existing work
- Backstitch repairs breaks with an agent on a private fleet
- Hindsight and Mint assume context arrives intact
- the Rust slice studies cancellation across boundaries
- first week
- pick twenty open-source Rust services that use the tracing or OpenTelemetry crates
- tag requests at entry and count where the tag is missing downstream
- classify breaks: detached task, channel hop, thread pool, external queue
- possible outcome: a lint or a type-level rule, plus the first public count
- stop if breaks are rare in Rust because the common libraries already carry context
- a public test bed for wrong-number detectors in multi-GPU jobs
- question: do the published detection ideas still work outside the fleet they were tuned on?
- existing work
- SDCHunter and AEGIS report fleet results nobody else can rerun
- SCOUT compares replicas; TrainSDC, LLM-PRISM, and Tung et al. inject simulated faults
- Tung et al. report that NaN checks see about 1% of corruptions
- PrismLLM emulates large runs on few GPUs
- first week
- inject bit patterns from Tung et al.âs published statistics into a small multi-GPU run
- compare three cheap checks: NaN check, replica agreement, checksum on matrix multiply
- report detection rate and overhead
- gap I am less sure of: serving, where no gradient averages the error away and one wrong token reaches a user
- SAVE targets edge devices; I found no fleet study for LLM serving
- risk: without real defective GPUs the injected faults may not look like real ones; ByteDance says real ones depend on the exact input
- check the clockâs promise from the outside
- question: when a cloud clock service says âwithin ±Δâ, how often is that false, and do systems built on it notice?
- existing work
- Sundial and Graham design for failures; Firefly and SyncWise push accuracy
- Tiga, K2, and Nezha assume the bound holds
- TempoTrace claims trace ordering breaks under loose clocks, unverified
- first week
- first identify what each bound promises: relative agreement, UTC accuracy, or both
- compare reported intervals across VM pairs while measuring network delay
- inconsistent intervals can refute joint promises under explicit delay assumptions
- agreement alone cannot establish UTC accuracy because both clocks can share the same error
- use an independent time reference with a stated error bound before claiming absolute accuracy
- inject a bound violation into one clock-ordered open-source system and check its behavior
- this controlled experiment does not measure the frequency of real cloud violations
- weak point: I have not yet searched for existing cloud clock measurements, so this may be done
weaker ideas, listed so nobody repeats the thinking
- another LLM diagnosis agent: crowded; AIOpsLab, PRAXIS, UModel, and Kim et al. already cover building and criticizing them
- another trace sampler: Hindsight, Mint, and Trace Sampling 2.0 cover the main design choices
- anything needing a hyperscale GPU fleet or custom CXL hardware
Last edited: