cloud resource management: serverless, fast scaling, fine-grained scheduling (authored by agents unless marked đ§)
source scope
- original review retained, with targeted primary-source checks on 7 Oct 2026
- performance numbers below are author reports under their evaluation conditions
- checked abstracts establish claims, not independent replication
- items marked unverified need a primary-source check before citation
- research ideas are agent recommendations, with novelty limits stated below
the problem in one paragraph
- providers share machines among customers with changing demand
- reserving each customerâs peak wastes capacity
- allocating on demand can make requests wait
- the mechanisms below reduce allocation delay
- sandbox startup
- orchestration and autoscaling
- live CPU and memory resizing
- agent inference: fast mechanisms do not settle end-to-end behavior
- compatibility, interference, failures, and recovery remain separate questions
words used below
- sandbox: the isolated box one customerâs code runs in
- container, lightweight virtual machine (microVM), or WebAssembly
- cold start: a request arrives and no sandbox is ready, so one must be created first
- warm start: a ready sandbox takes the request
- horizontal scaling: add more sandboxes
- vertical scaling: give one running sandbox more cores or memory
- snapshot: saved memory of an initialized sandbox, restored later to skip startup work
- control plane: the software that decides how many sandboxes run and where
- tail latency: the slowest few percent of requests
- oversubscription: selling more resources than the machine has, betting customers do not peak together
- durable execution: the platform saves a workflowâs progress so that it resumes recorded progress after a crash
- steps may replay if their completion record was not saved
- duplicate external effects require separate handling
- metastable failure: an overload that keeps itself going after its cause is gone
what the human already has
- talk note âTowards Microsecond-Scale vm Core Provisioning Agility on Serverless Platformsâ, Yibo Yan, NSL meeting
- in the reading notes
- now on arXiv as HyperFlux, see below
- correction: HyperFlux names its guest operating system FluxOS
- the overall system and its guest component have different names
- talk note âGranula Resource Demand Heterogeneityâ, Coulson Liang
- the HotOS 2025 talk on Hiresperf, see below
- paper collection: almost nothing on this topic
- âServerless Runtime - Database Co-Design With Asynchronous I-Oâ, EdgeSys 2024
- âCluster-Based Scalable Network Servicesâ, SOSP 1997
- âVeriSMo: A Verified Security Module for Confidential VMsâ, OSDI 2024
- âDifference Engine: Harnessing Memory Redundancy in Virtual Machinesâ
- the rest is datacenter networking
literature: cold start
some prototypes approach warm-start latency on their tested functions
- comparison depends on language initialization, storage, networking, and the isolation boundary
Firecracker, NSDI 2020, AWS
- the lightweight virtual machine monitor under Lambda; written in Rust
- why it exists: âpublic infrastructure providers, who need both strong security and minimal overheadâ
- author report: millions of production workloads and trillions of monthly requests
Serverless in the Wild, ATC 2020, Azure
- production Azure trace used as an evaluation workload in serverless research
- âmost functions are invoked very infrequently, but there is an 8-order-of-magnitude range of invocation frequenciesâ
- consequence: keeping every function warm is hopeless, so cold starts must be cheap
Serverless Cold Starts and Where to Find Them, EuroSys 2025, Huawei
- âa month-long trace of 85 billion user requests and 11.9 million cold starts from Huaweiâs serverless cloud platformâ
- full primary methods, §§2 and 4
- request and pod records cover five regions and 20 clusters over 31 days
- not every function in those regions is included
- function metadata covers one region
- timestamps use milliseconds; recorded execution and cold-start durations use microseconds
- component times distinguish pod selection or creation, code download/extraction, dependency loading, and networking/routing/scheduling
- existing pod pools have popular runtimes and libraries preinstalled
- default keep-alive is one minute, renewed on each request
- Figure 11 plots hourly mean component times
- these are not per-request tail percentiles or universal maximum delays
- region and hourly workload composition affect the observed bottleneck
- implication: preserve event-level distributions and separate pool selection from creation in comparisons
- read-depth limit: selected full dataset, lifecycle, and component methods inspected
- raw trace not independently analyzed
How Does It Function?, SoCC 2023, Huawei; selected full §§2â5
- authors: âover 7 months with over 1.4 trillion function invocations combinedâ
- private: 200 internal functions, 141 observed days within 234 calendar days
- second-resolution arrival counts and mean delays
- minute-resolution resource measures and instance counts
- public: 5019 functions, one availability zone, 26 consecutive days
- released arrivals have minute resolution
- cold-start, package-size, and language data withheld
- pod-averaged utilization/delays cannot recover individual-request tails
- platform delay includes scheduling and networking
- correlations with instance allocation do not identify causal bottlenecks or cache behavior
- forecasting selects ten functions with highest median request counts
- private 38 training days in separated chunks, then seven validation and seven test days
- public 20/3/3 days
- inference: aggregate replay and forecasting need rare-function and within-bin burst controls
- selected full characterization/forecasting methods read; artifact not executed
Fork in the Road (AFaaS), OSDI 2025, Ant Group and Tsinghua
- authors: âcontrol path latencyâ and âresource contention latencyâ
- user-code initialization is the third source
- authors report 18 months in production
- full primary methods, §§2, 4â6 and 8
- latency starts when a node receives a scaling request and ends with the functionâs response
- earlier scheduler waiting is outside this boundary
- replaces local runtime communication, pools setup resources, and clones progressively specialized initialized parents
- unlike disk restoration, a suitable parent remains in memory
- controlled tests use FunctionBench/SeBS functions on one 24-core server
- compares progressively optimized Catalyzer variants, Kata cloning a prepared VM, and gVisor starting from scratch
- these baselines do not begin with identical reusable state
- production comparison covers eight Node.js functions over one day
- CataOnly runs separately on matching hardware with mocked peer-service responses
- not a randomized comparison across all production functions
- limitations include runtime-specific integration and duplicated random state after cloning
- random state must be reinitialized to preserve instance uniqueness
- extreme concurrency still causes contention
- latency starts when a node receives a scaling request and ends with the functionâs response
- read-depth limit: selected full methods read; traces and security not independently tested
Spice: Rethinking Process Snapshots for Near-Warm Serverless Cold Starts, OSDI 2026, MIT
- âthese limitations stem from a lack of OS support for snapshot restorationâ
- author report: restore and execution add 0.6â18 ms over warm invocation
- compared systems add 3.6â1197 ms in those experiments
- full primary methods, §§2â5 and 7
- reported invocation latency includes restoration and function execution
- excludes container/isolation setup and platform orchestration
- differs from AFaaSâs local startup boundary and customer network latency
- compares tuned CRIU, FaaSnap, and REAP from a cold host page cache
- warm comparison retains state with cold CPU caches
- FunctionBench plus Java/Node.js ports run on one server
- mixed-load test scales Azure trace durations and maps them to benchmark functions
- not original traced applications
- pauses threads at a safe point and profiles later page accesses before preparing snapshots
- unchanged files must remain available at restoration
- prototype uses static paths; production needs content identity
- process metadata restoration uses Junctionâs library operating system
- unmodified application binaries run inside its sandbox
- native-Linux metadata restoration and VM integration remain work
- reported invocation latency includes restoration and function execution
- read-depth limit: selected methods/evaluation read; artifact and security not independently tested
Dandelion, SOSP 2025, ETH Zurich
- drops the guest operating system: user code is pure computation, the platform does all input and output
- sandboxes âcold start in hundreds of microsecondsâ
- author reports: two to three orders less variability and 96% less committed memory
- comparison is against Firecracker in the evaluated applications
- cost: you must rewrite the application as a graph of pure functions
Fast end-to-end cloud application cold-start with initscripts, arXiv Aug 2026, MIT
- the developer hands the platform a script of initialization work, so it runs while the platform still downloads the binary
- full primary methods, §§4â6
- authors: âthe costs of binary download and container creation are includedâ
- start latency runs from spawning an instance until it begins handling its first request
- excludes completion of that request
- cold means the first application run on that machine
- fetches initialization state while binaries and isolation are prepared
- modified RPC client libraries return the prefetched results
- library compatibility and result naming are part of deployment work
- compares applications with and without initscripts on the same sigmaOS platform
- eight 16-vCPU AWS machines for application experiments
- separate bare-metal machine for snapshot-system comparison
- reports average 1.67Ă start speedup for ServerlessBench functions fetching S3 inputs and weights
- functions without that external fetching do not improve
- unmodified memcached still reconstructs its data structures
- compatibility copying through the filesystem costs about 85 ms
- implication: compare first-request completion, transfer copies, and initialization work as well as start time
- read-depth limit: selected full implementation and evaluation inspected
- no artifact execution or generalized benefit claim
Sabre, OSDI 2024; selected full §§3â5
- authors: âhardware-accelerated (de)compressionâ
- Intel IAA losslessly compresses Firecracker snapshots
- restoration overlaps decompression and disk reads
- contiguous copying or scattered placement depends on page sparsity
- buffer pools trade memory density for restoration speed
- tested Sapphire Rapids, DDR5, 550-MB/s storage; end-to-end evaluation uses one IAA engine
- dirty-page snapshots and working sets recorded during first invocation evaluated
- vSwarm, FunctionBench, and SeBS workloads include changed DNA inputs and image sizes
- synthetic Python-list illustrates an upper bound
- working-set prefetching improves 25â55%; corresponding cold-start improvement reaches 20%
- restoration speed is not first-request completion speed
- held-out workloads, P99 estimates, and concurrent accelerator contention unspecified in selected methods
- snapshot creation is outside the restoration path but retains deployment cost
- inference: measure page coverage, accelerator occupancy, storage bandwidth, and request completion separately
- selected full methods/evaluation read; artifact not executed
TrEnv-X, March 2026 v2; selected full §§3â5/8â9
- offline snapshots form deduplicated remote read-only templates
- restoration copies metadata; CXL mappings and RDMA faults use copy-on-write
- reuse kills processes and purges filesystem changes; network state persists
- authors: âwe need fallback to create absolutely clean network environment from scratch as usualâ
- fallback applies to network customization or stricter security
- ten Python/Node functions on dual-32-core Xeon, experimental CXL, and Soft-RoCE
- synthetic and Azure/Huawei patterns; restoration baselines enhanced with pooling and remote snapshots
- up to 7Ă P99 end-to-end Azure improvement versus REAP+
- six agents replay recorded LLM outputs and timing; 200 instances share 20 cores
- startup improves 35â49% versus E2B; lazy faults follow startup
- shared-browser P99 end-to-end improvement: 2â58%
- peak-memory improvement: 10â61% versus E2B, up to 48% versus E2B+
- cost model uses memoryĂduration, not measured energy or deployment expense
- selected methods read; implementation, SLO guarantees, and snapshot-creation costs unvalidated
- source CXL latency unit appears inconsistent; unresolved
- offline snapshots form deduplicated remote read-only templates
Restoring Uniqueness in MicroVM Snapshots, 2021, AWS
- snapshots can clone random state unless restoration supplies fresh state
- authors: ârestoring the uniqueness of the VMsâ
- examples include identifiers, secrets, and nonces
- full paper, §§3â4: reseeding alone is insufficient if requests resume before reset hooks finish
- proposed interfaces notify guests of a new instance generation or clear marked memory pages
- experiment requirement: gate external connections and writes until required reset hooks complete
older work I know but did not reopen: REAP and vHive (ASPLOS 2021), Catalyzer (ASPLOS 2020), FaaSnap (EuroSys 2022), Mitosis remote fork (OSDI 2023; checked below), SEUSS, SOCK
literature: control plane and autoscaling
orchestration can dominate startup after sandbox creation becomes cheap
Dirigent, SOSP 2024, ETH Zurich
- authors identify âthe latency to schedule functionsâ as potentially much larger than sandbox startup
- cause: Kubernetes writes state to its store on every scale step
- clean rewrite; author reports 2500 sandboxes per second, 1250Ă Knative under its comparison
KubeDirect, NSDI 2026, Peking University
- keeps Kubernetes, but lets controllers message each other and skip the central store
- ephemeral controller state requires coordinated recovery
- its recovery design and TLA+ model are covered below
- âreduces serving latency by 26.7x over Knativeâ
The High Cost of Keeping Warm, arXiv 2025
- rebuilt the scaling behavior of AWS Lambda and Google Cloud Run in an open system
- creating and destroying instances burns â10-40% of the CPU cycles spent on request handlingâ
- allocated memory exceeds used memory by â2-10 timesâ
Jiagu, ATC 2024; selected full §§4â7
- authors: âdecoupling prediction and schedulingâ
- authors: âwe simulate how Jiagu can optimize the cold startâ
- cached capacity tables support placement; misses require inference and later asynchronous updates
- scaling stops routing before eviction, retaining instances for reactivation or migration
- OpenFaaS prototype: 24 homogeneous Xeon machines, six benchmark functions, four Huawei-derived pattern sets
- compares Kubernetes, reimplemented Gsight/Owl, and scaling-stage ablation
- 81.0â93.7% scheduling-cost reduction versus Gsight
- 57.4â69.3% cold-start reductions combine measured scheduling with simulated initialization
- these are not deployed end-to-end request-latency measurements
- new-function prediction errors converge after approximately 5â30 measurements
- this does not establish temporal-shift robustness
- inference: update work moves off placementâs critical path rather than disappearing
- include profiling, retraining, migration, and stale-table costs
- selected full methods/results read; artifact not executed
Coach, ASPLOS 2025, Microsoft Azure
- introduces CoachVMs with guaranteed and oversubscribed resource portions
- pairs workloads whose resource peaks occur at different times
- âCoach enables platforms to host up to ~26% more VMs with minimal performance degradationâ
- full paper, §§3â4: the capacity result comes from trace-driven simulation using a production scheduler
- separate real-server experiments measure workload performance
- does not establish a deployed fleet-wide 26% gain
- characterization emphasizes VMs lasting over one day
- monitoring every 20 seconds predicts demand over the next five minutes
- inference: short agent bursts require separate measurements
- compare burst duration with monitoring delay before reusing the predictor
- the abstract singles out memory as the hard resource to take back
Harvest VMs, OSDI 2020, Azure; selected full §§4â7
- authors: âgrows and shrinks according to the amount of unallocated resourcesâ
- survival and available-core predictions support service objectives
- population predictions rather than individual guarantees
- fixed virtual-core count receives changing physical CPU allocations
- regular VMs taking the minimum allocation trigger eviction
- harvests unallocated CPU rather than reclaiming allocated idle CPU or memory
- production-arrival simulation: 25 clusters, 14 regions, 45 training days and 45 evaluation days
- ten-minute allocation-state resolution
- separate real Hadoop deployments adapt task admission and give shrinking workers a 30-second grace period
- masters remain on regular VMs; remote storage and replicas support recovery
- 91% useful-core cost reduction versus one-core regular VMs assumes half-price additional evictable cores and includes recovery cost
- average saving versus standard evictable VMs: 42%
- not a universal cloud-bill reduction
- inference: agent bursts need separate latency, checkpoint, retry, and external-effect tests
- selected full simulation/deployment methods read; current pricing and artifact unchecked
literature: fine-grained scheduling
move cores as demand changes instead of reserving each tenantâs peak
Shenango, NSDI 2019; selected full §§4â7
- authors: âevery 5 ”sâ
- a dedicated IOKernel core coordinates allocation and packets
- five-microsecond congestion polling is not core-transfer completion or request latency
- guaranteed cores are not oversubscribed; critical sections can defer preemption
- malicious refusal mitigation remains unimplemented
- principal comparison budgets eight physical cores and sixteen hyperthreads
- IOKernel consumes two hyperthreads; competing systems differ in networking capabilities
- Poisson TCP arrivals, memcached, synthetic tasks, and background swaptions
- sudden-load experiment uses synthetic one-microsecond work rather than production arrivals
- measures median and 99.9th-percentile response latency alongside foreground/background throughput
- selected full methods/evaluation read; runtime and isolation not independently verified
Caladan, OSDI 2020, MIT
- âfast core allocation instead of resource partitioningâ
- author-reported tail latency under shifting demand: 580 ms to 52 ÎŒs
- limit: tenants are processes on one Linux host, not isolated virtual machines
Concord, SOSP 2023; selected full §§3â6
- authors: âclosely approximatingâ strict preemption and a single queue
- authors: âdevelopers must modify their codeâ
- compiler checks shared signals; workers yield at checks rather than precise interrupt points
- central queue and bounded local queues; idle dispatcher executes requests
- external calls and annotated locks defer preemption
- protected sections can raise tails; source and LLVM-compatible compilation required
- principal evaluation: two connected Xeon machines, 14 workers, Poisson arrivals
- synthetic tasks and memory-resident LevelDB, matched tail-slowdown targets
- 83% greater throughput for one high-dispersion LevelDB workload with two-microsecond quantum
- not a universal latency improvement; lighter-load tails can worsen
- separate instrumentation averages about 1.04% overhead
- quantum variation is empirical, not a worst-case bound
- inference: test deferred sections, dispatcher saturation, and changing task mixtures
- selected full design/evaluation read; artifact and production deployment unverified
Efficient Scheduling Policies for Microsecond-Scale Tasks, NSDI 2022; selected full §§3â6
- authors: âstatic core allocations often outperform reallocation with small tasksâ
- inference: faster reallocation alone does not establish a better policy
- static-allocation comparison assumes average load is âconstant and known a prioriâ
- changing means and longer tasks alter the result
- simulation defaults: 32 cores, one-microsecond exponential tasks, Poisson arrivals, 50% mean load
- models 100-nanosecond communication and five-microsecond allocation
- excludes preemption, allocation-related cache effects, and statistics-reporting cost
- Caladan implementation revokes cores between tasks
- polling interval does not guarantee reclaim latency
- real tests use memcached, Poisson UDP arrivals, and background swaptions on 32 hyperthreads
- compares allocation policies on the same runtime, not whole Shenango/Caladan systems
- similar median/P99 latency with higher background throughput
- selected full model/implementation/evaluation read; production arrivals and independent execution unchecked
HyperFlux, arXiv Aug 2026, Yibo Yan and Seo Jin Park, USC
- brings Caladan-speed core moves to hardware-isolated virtual machines
- avoids guest vCPU hot-plug by changing physical backing of existing vCPU threads
- author-reported core movement: 13 ÎŒs, including forced reclaim
- author-reported footprint: 3.2 MB; cold boot: 1.37 ms
- cores only; memory is not elastic here
Quicksand, NSDI 2025, MIT, Brown, USC, VMware
- splits an application into pieces that each use mostly one resource, then moves the pieces between machines in milliseconds
- âresource proclets, granular units that each primarily consume resources of one typeâ
- cost: the application must be written against Quicksandâs data structures
Granny, NSDI 2025, Imperial
- same goal for OpenMP and MPI programs, using WebAssembly units that snapshot and migrate
- âreduces the makespan for OpenMP workloads by up to 60% and the fragmentation for MPI workloads by up to 25%â
Granular Resource Demand Heterogeneity (Hiresperf), HotOS 2025, Liang, Govindan, Park, USC
- one request needs different hardware at different moments
- slide example: a search requestâs two phases use 117 vs 218 bytes per microsecond of memory bandwidth
- Hiresperf profiles each invocationâs resource usage at 10 ”s resolution
- authors report 7â15% overhead or lower
- the authorsâ own to-do list: âBatched Processingâ, âQoS with High Utilizationâ, âFunction as a serviceâ
- authors call for a âProactive Approachâ, contrasting Caladanâs reactive allocation
- one request needs different hardware at different moments
literature: the new workload, AI agent sandboxes
AI code execution adds persistent sessions, idle intervals, and checkpointing
- agent behavior belongs to the AI agents group
- this page covers the execution platform
AWS Lambda MicroVMs, AWS blog, 10 Jul 2026
- snapshots as a product: âMicroVMs are launched from MicroVM images, which are pre-initialized Firecracker snapshotsâ
- aimed at âAI coding assistants and agentsâ
- official guide: âvertically scale to 4x of configured baselineâ
- advertised capability, not independently measured scaling latency or availability
- official pricing: âYou pay for baseline compute resources while your MicroVM is runningâ
- additional consumed memory and CPU billed for active duration, per second
- snapshot storage, reads, writes, and data transfer add costs
- scaling description includes âup to 8GB / 4vCPUâ
- this may describe four times the default rather than a universal maximum
- its second example scales an 8-GB baseline to 32 GB
- effective maximum remains unchecked against service limits
- read-depth limit: official announcement, guide, and pricing passages checked on 7 October 2026
- no API execution, deployment, billing verification, or latency measurement
AgentCgroup, submitted February 2026, revised July 2026, §§3 and 5â6
- authors: âImage pull time is excluded from all reported initialization measurementsâ
- 111 GLM tasks and 33 Haiku tasks, with the latter drawn from the former set
- 144 executions do not mean 144 distinct tasks
- one 24-core, 128-GB machine; initially unlimited Podman containers
- resource characterization samples CPU and memory once per second
- tool time includes dispatch, subprocess execution, and collecting the result
- initialization includes image-layer ID remapping and agent startup
- 55â60% lifecycle share combines initialization with tool execution
- useful computation is included, not just avoidable OS overhead
- maximum reported memory ratio is 4060 MB peak divided by 264 MB execution average
- 15.4Ă belongs to one extreme task, not the typical task
- one-second samples do not establish sub-millisecond burst timing
- controller evaluation replays three memory traces at 50Ă speed on a different 16-GB machine
- compares priority enforcement with no isolation under induced memory pressure
- not a complete end-to-end coding-task or production-concurrency evaluation
- tool-call hints and per-tool cgroups already overlap a proposed hint-driven memory-lending controller
- read-depth limit: selected full characterization, controller design, and preliminary evaluation inspected
- prototype not executed; hint accuracy and total usable-memory latency remain unverified
DeltaBox, arXiv May 2026, Shanghai Jiao Tong
- agents checkpoint often and âsubsequent checkpoints in AI agents are highly similarâ, so save only the difference
- full paper, §§4â6: earlier abstract latency figures differ from the inspected evaluation
- Table 2 reports weighted checkpoint call-to-return latency 10.83 ms
- excludes asynchronous dump completion
- restore takes 1.86 ms with a retained process template, 9.29 ms after eviction
- experiment requirement: measure when the recovery point becomes usable and crash before asynchronous work finishes
- filesystem layers and process checkpoints form a consistent pair
- pauses the agent while saving state
- moves model-network requests outside the checkpointed process
- external network effects cannot be rolled back
- evaluation replays 24 SWE-bench trajectories on 4-vCPU, 8-GB VMs
- baselines include a self-hosted E2B implementation, not the hosted product
- raw fork measurements omit the complete checkpoint machinery
Crab, Apr 2026, §§4â7
- saves different amounts of state according to each agent turnâs effects
- schedules checkpoint work across tenants and overlaps it with model requests
- §7.2 crash-recovery evaluation injects one randomly positioned crash per run
- separate preemption experiments inject 1â5 preemptions per task
- 100 tasks per configuration selected for success without failures
- tests recovered task results or final patches, not arbitrary external effects
- inference: replaying an outstanding remote command may repeat an external effect
- test this hypothesis with a service that records every operation
SpecBox, Jul 2026 preprint; selected full §§3â5
- starts the sandbox while the model is still writing the tool call
- âcuts P99 end-to-end latency by up to 2.9Ă relative to the on-demand sandbox baselineâ
- keyword/semantic routing and transition counts predict environment preparation
- preparation does not execute side-effecting tool calls
- authors: âunused warmups are discardedâ
- cancellation and cleanup timing unmeasured in selected sections
- one 16-core Docker host uses Qwen3.5-Max cloud API
- 200 generated trajectories from 32 MCP servers use actual execution, then stepwise replay
- reported trajectory-length distributions need reconciliation
- at 20 QPS, session P99 is 88.7 versus 257.2 seconds
- session completion, not sandbox startup
- replay does not measure independently evolving live-agent trajectories
- local CPU/memory measured; provider GPU use, token equality, and total cost unestablished
- shared transfer trusts host access controls; malicious agents/infrastructure excluded
- semantic cache reports 84.8% bypass ratio and validation fallback
- validated hits can skip sandbox setup; others return to normal execution
- this is not correct-result precision
- similarity does not prove equivalent results
- selected full methods read; implementation and artifacts unverified
Agentic AI Workload Characterization, arXiv May 2026
- agents move âfrom read/explore behavior early in execution to execute/write behavior laterâ
Aries, Jul 2026 preprint; selected full §§3â4
- experimentation framework, not a microVM fork mechanism
- authors: âtool sandboxes alternate between long idle periods and short resource burstsâ
- framework separates task, harness, model, and sandbox choices with correlated telemetry
- commercial evidence: one eight-hour day, ten model-serving instances, minute samples from 100 sandboxes
- synchronized private step timing unavailable
- controlled experiments: 20 tasks each from three benchmarks, chosen for 50% OpenClaw success, five repetitions
- selection limits transfer to all agent workloads
- tool-resource experiments use DeepSeek V4 Flash API and second-resolution sampling
- commercial samples combine harness and tool costs; controlled telemetry separates them
- AWS-pricing replay uses recorded tool timelines and varied keepalive
- cost multipliers are analytical results, not deployed billing
- possible 3Ă saving is an ideal-system proposal, not a tested controller
- selected full methods read; artifact, synchronization, and privacy not independently verified
forkd and mitos
- earlier search leads remain unverified
- do not confuse mitos with the RDMA-based Mitosis paper
Mitosis, OSDI 2023; abstract checked
- authors: âfast remote forkâ
- Linux containers share previously initialized state through RDMA reads
literature: staying correct when things fail
scaling-specific failures and durable execution
- general distributed verification belongs to the distributed systems group
Metastable Failures in the Wild, OSDI 2022
- âat least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failuresâ
Formal Analysis of Metastable Failures in Software Systems, arXiv 2025, Alvaro, Isaacs, Majumdar and others
- models request and response servers as Markov chains
- full paper, §§3â4: models bounded queues, worker pools, retries, timeouts, and dropped requests
- continuous-time Markov chains approximate service and queue timing
- averages service rates across request types
- simulation retains details that the mathematical approximation discards
- §6.3 varies arrival and processing rates to study recovery
- experiment requirement: compare predicted recovery against simulation and measured histories before using the approximation as a safety bound
Characterizing Metastable Faults and Failures, arXiv 2026, Cornell and Penn
- full paper, §7 already models a capacity-management feedback loop
- authors: âwakeup commands take longer to execute than sleep commandsâ
- overload delays heartbeats, causing a manager to remove workers before replacement workers start
- oscillation persists after the initial disturbance ends
- network loss is unnecessary in the example
- model uses 40 workers and abstracts overload as a function of active worker count
- omits the external load balancer
- assumptions include eventual execution of component actions
- proposed repair delays worker removal until replacement capacity is ready
Mutiny!, DSN 2024
- corrupts one value in the Kubernetes store and watches
- âservice under/overprovisioning (24%)â of injections
Quantifying Autoscaler Vulnerabilities, arXiv 2026
- simulation only; âhorizontal autoscaling exhibits greater susceptibility to transient anomalies, particularly near threshold boundariesâ
Anvil, OSDI 2024
- Kubernetes controllers written in Rust and proved in Verus to always reach the desired state
- authors: âreconciliation is fundamentally not a safety propertyâ
- proves eventual progress under its stated assumptions
- covers controllers for ZooKeeper, RabbitMQ, FluentBit; not the scheduler or the autoscaler
AWS Nitro Isolation Engine, announced 2025; production account checked
- authors: âguest memory allocations are always scrubbed before reuseâ
- Isabelle/HOL proves properties of the isolation component
- scheduling and allocation policy remain outside that component
Distributed Speculative Execution, OSDI 2026, MIT and Microsoft
- durable execution âusually forces frequent and synchronous persistence, resulting in significant latency overheadsâ
- the runtime skips the saves and repairs state after a failure, hidden from the developer
Consistency and Correctness in Data-Oriented Workflow Systems, CIDR 2026, Stonebraker, Zhou, Kraft, Li
- âAlthough many developers can write and test a saga, few get it right when the server crashesâ
- âdurability alone is not sufficientâ
Beldi, OSDI 2020; abstract checked
- authors: âfault-tolerant and transactional stateful serverless functionsâ
Netherite, VLDB 2022; abstract checked
- authors: âworkflow steps to group commit, even if causally dependentâ
Boki, Temporal, and Cloudflare Durable Objects
- retained background leads, not reopened here
Peeking Behind the Curtains of Serverless Platforms, ATC 2018
- the last big study from the customerâs side that I found
- âlaunching more than 50,000 function instances across these three servicesâ
- found âthat Google had bugs that allow customers to use resources for freeâ
what I conclude from the literature
- several mechanisms are fast in evaluated prototypes
- the results do not imply that startup or scheduling is solved
- the NSDI 2022 policy study above shows why mechanism speed can mislead
- faster designs change where correctness obligations sit
- KubeDirect supplies recovery protocols and a TLA+ model
- HyperFlux relies on host-controlled parking and a guest resume protocol
- snapshots require restoration of unique state
- agent recommendation: audit those boundaries rather than presume lost safety
- memory elasticity already has substantial prior work
- see the checked papers added below
- a new project needs a workload or guarantee they do not cover
- independent platform measurement exists
- Denoâs 2024 Lambda benchmark provides âConfiguration, benchmarking scripts, and raw dataâ
- a runtime vendorâs benchmark has a narrower scope and possible commercial bias
- a broad 2026 sandbox study is a candidate, not an established empty field
- application compatibility remains a tradeoff
- Dandelion, Quicksand, and Granny expose different programming constraints
- compare their supported interfaces before judging deployment effort
research we could do
- ordering is an agent recommendation
- absence from this review is not evidence of novelty
- prioritize a small reproduction before choosing a full project
- prove the core hand-off of an elastic virtual machine layer
- what: take HyperFluxâs protocol for moving a core between virtual machines and prove it in Verus
- no core ever runs two virtual machines
- a guest whose core is taken by force loses no work and never sees a torn state
- a bursting high-priority guest gets a core within a stated bound
- why it could work
- the paper exposes a six-state vCPU protocol suitable for an explicit specification
- hypothesis: interrupted handoffs can expose ordering errors that the current tests do not exercise
- Nitro and VeriSMo establish close verification precedents
- novelty would lie in the specific handoff and implementation guarantees, not cloud isolation verification itself
- risk
- correction: HyperFlux §4.10 says âHyperflux is implemented in Rustâ
- Rust implementation does not remove the trusted KVM, Linux, interrupt, and hardware assumptions
- first reproduce the implementation and specify which interruption points the proof covers
- bounded allocation progress requires stated scheduling and capacity assumptions
- preserving guest work requires an honest, compatible guest substrate
- malicious guest isolation and correct guest continuation are separate properties
- overlap: formal verification group for the Verus side
- check the gap between a fast control planeâs model and its implementation
- correction: KubeDirect already models end-to-end correctness
- §4.4: âwe use TLA+ to verify the end-to-end properties of Kubedirectâ
- source: KubeDirect full paper
- its liveness assumption requires repeated sufficiently long connected intervals
- what: obtain the model, reproduce its checks, then test implementation traces against its transitions
- delayed old messages, controller replacement, overlapping deletion and creation, external updates
- evaluate how long surplus or missing sandboxes persist after recovery
- closest existing work
- KubeDirectâs own model covers convergence
- Dirigent §4.2 and §5.4 cover recovery design and failure experiments
- Anvil covers verified reconciliation implementations
- possible contribution
- a demonstrated model/implementation mismatch or a checked connection between them
- a second model alone offers little differentiation
- stop condition: the released model and existing tests already cover the proposed fault cases
- overlap: distributed systems and formal verification groups
- overloads that autoscaling causes itself
- what: cold start delay causes timeouts, timeouts cause retries, retries cause more cold starts. Find when this loop sustains itself, on Knative and on the open Lambda-like system from âThe High Cost of Keeping Warmâ, and on scale-to-zero databases
- why
- the Markov chain paper §6.3 already studies recovery by changing processing rates
- authors: âa recovery policy can throttle the arrival rate or increase the processing rateâ
- the June 2026 paper already combines delayed provisioning and heartbeat-driven capacity loss
- agent recommendation: reproduce that case before studying a real Knative or database controller
- candidate distinction remains unestablished
- measure interactions among retries, cold starts, and memory reclamation in the implementation
- compare measured recovery boundaries with both published models
- test whether a guard requiring ready replacement capacity prevents the observed failure
- the general retry feedback mechanism is established prior work
- the 2026 autoscaler fault study is simulation only
- fits ârobust ⊠websites and databasesâ directly
- stop condition: the implementation only reproduces an existing model without exposing a mismatch or a useful new control rule
- measure serverless and sandbox platforms from the customerâs side, 2026 edition
- what: rerun the 2018 âPeeking Behind the Curtainsâ questions on todayâs products: Lambda, Lambda MicroVMs, Cloud Run, Cloudflare Workers, E2B, Modal, Fly, plus Neon and Aurora scale to zero
- cold start, resume from snapshot, how fast vertical scaling reacts, billed versus used, are clones from one snapshot really different
- why
- you do measurement; this needs careful method more than a new system
- billing and snapshot semantics are measurable questions
- prior audits of each product must be checked before claiming a new finding
- gives a public dataset the academic systems lack for agent sandboxes
- risk: reads as âjust a benchmarkâ unless a finding surprises; the snapshot clone check (idea 6) could be that finding
- cold starts that real website visitors see
- what: crawl to find websites served by serverless hosts (response headers give Vercel, Netlify, Cloudflare, Lambda URLs away), then visit each gently after different idle gaps and measure the extra delay
- why
- joins your web measurement work with this topic
- candidate distinction: user-visible delays across independently hosted websites
- provider attribution from headers can be incomplete or misleading
- independent benchmarks already exist, as noted above
- answers a plain question: how many sites make the first visitor after a quiet hour wait, and how long
- risk
- separating cold start from cache miss and network noise needs a careful design
- keep the request rate tiny; check with the web measurement groupâs ethics notes
- clones that should differ but do not
- what: restore one snapshot twice, run both, and report state that is equal but should be unique (random numbers, identifiers, session keys, timestamps, open connections). Survey popular images and language runtimes
- why
- AWSâs snapshot uniqueness paper is direct prior work
- candidate distinction: current runtime/image combinations and documented reset hooks
- check the race between hook completion and the first external operation
- define expected identity separately for a clone and a rollback within one logical instance
- a duplicated key or nonce is a security bug, so findings are reportable
- risk: Linux and the big runtimes may have fixed most cases already (the virtual machine generation ID work); then the result is a short negative paper
- check or prove durable workflow code
- what, two sizes
- a tester that crashes a workflow at every step and replays it, reporting steps that ran twice or differently (Temporal, DBOS, Restate, Durable Functions)
- a small durable execution library in Rust proved in Verus: given steps that are safe to repeat, the workflowâs effect equals one clean run
- why: CIDR 2026 says âfew get it right when the server crashesâ; speculative execution (OSDI 2026) makes the runtime harder to trust
- risk: deterministic simulation testing companies and the vendorsâ own replay tests cover part of the first; I could not check how much
- newer local-state competitors: DeltaBox and Crab
- compare filesystem/process consistency with external-effect recovery
- crash during open-file updates, child-process work, pending I/O, and failed checkpoints
- successful final-task recovery alone does not establish that remote effects occurred once
- overlap: distributed systems group
- fast memory lending between agent sandboxes
- what: the memory twin of HyperFlux, driven by hints of what tool the agent is about to run
- closest work: AgentCgroup, Pond, RamRyder, Memoryless, and Squeezy
- a claim that nobody has built elastic memory would be false
- candidate distinction: tool-call hints plus a measured bound on safe reclaim latency
- risk: pure systems building, far from your strengths, and SJTU and Tsinghua move fast here. I would only do it as a partner to the lab
- place functions by their measured per-call demand
- what: the âFunction as a serviceâ direction on the Hiresperf slides
- risk: it is the Hiresperf authorsâ own next step. Offer to help, do not compete
additional checked memory work
- Pond, ASPLOS 2023
- authors: âpooling across 8-16 sockets is enough to achieve most of the benefitsâ
- CXL permits load/store access to pooled memory
- evaluate agent memory lending against pooling, not just static reservations
- RamRyder, OSDI 2026
- authors: âallocate memory bandwidth and capacity (mostly) independentlyâ
- separates resource quantities often bundled together
- Memoryless, preprint Sep 2026
- authors: âmemory locality as a serverless control-plane primitiveâ
- varies local/remote memory mix when placing functions
- abstract checked; latest preprint needs full evaluation scrutiny
- Squeezy
- full paper, §§4â6, revised Dec 2025, EuroSys 2026
- authors: âsub-second reclamation of multiple GiBs of memoryâ
- partitions a shared VMâs memory by function instance, with a separate shared-runtime partition
- returns an instanceâs partition after that instance terminates
- does not establish arbitrary reclamation from a live persistent agent
- Linux 6.6.30 implementation excludes ordinary allocations from these partitions
- reclaiming 2 GiB takes 127 ms in its controlled test, versus 2.5 s for virtio-mem
- eliminates guest-page migration and defers zeroing to the host
- this does not eliminate sanitization
- §6.2.1 replays four Azure invocation traces under abundant memory
- §6.2.2 uses four functions and synthetic trace-like load on one host under scarce memory
- comparison implements HarvestVM-style optimizations rather than running the unavailable original prototype
- experiment requirement: distinguish instance termination from live-memory lending and include host zeroing under contention
- Puffer
- author-posted full text, §§IVâVI, IEEE TPDS Jan 2026
- publisher access unavailable; inspected the paper text, not a secondary summary
- reserves memory blocks for each function and returns them at completion
- prepopulates page mappings or populates them lazily according to workload sensitivity
- controlled resize of 512 MB takes under 1 ms
- separate page-population tests take 46â251 ms for 512 MBâ3 GB
- resizing a reservation is not the same as making all memory usable
- tests eight Python functions on a fixed-NUMA machine with swap disabled
- additional synthetic workload draws memory and runtime from distributions rather than replaying production traces
- linked artifacts were unavailable when checked
- closest competitor for memory lending; compare total usable-memory latency and live-agent reclamation
- author-posted full text, §§IVâVI, IEEE TPDS Jan 2026
cross-review and related coverage
- cross-topic ChatGPT review
- energy, edge platforms, isolation, and cluster scheduling
- GPU sharing and model serving belong to systems for machine learning
- consensus, storage, databases, and durable workflow verification belong to distributed systems
- Hydro and Service Weaver belong to the distributed systems group
remaining evidence limits
- Spice and AFaaS full methods now separate restoration latency from local platform startup and qualify benchmark/production comparisons
- Huawei component measurements, initscript start-time accounting, and AgentCgroup peak and replay definitions now checked in selected full primary methods
- Lambda MicroVM scaling and billing descriptions now checked against official documentation
- pricing parenthetical and larger example do not establish a universal resource maximum
- older cold-start papers and forkd/mitos remain leads
- novelty checks are targeted, not exhaustive
- no platform measurements or proofs were independently reproduced
Last edited: