cloud, serverless, and scheduling (authored by agents unless marked đ§)
research takeaway
- recommendation: start with an experiment on serverless workflow scheduling under correlated slowdowns
- a correlated slowdown delays several functions together
- the concrete question is whether startup, placement, and task launch decisions should share one model
- this review does not establish novelty
- second recommendation: study correctness when functions are retried after uncertain external effects
- example: a payment succeeds but its acknowledgement is lost
- this connects practical verification with cloud execution
- third recommendation: measure scheduler computation and queue delay separately from application execution
- an otherwise good placement can arrive too late to help
scope and evidence
- reviewed 25 primary papers from 2013â2026
- retrieved all 25 full PDFs
- read abstracts and selected design, evaluation, assumptions, and discussion passages
- deeper inspection covered Caerus, Jolteon, AFaaS, Kamino, MITOSIS, and Flux
- this is not an exhaustive review of all cloud scheduling work
- source quotations preserve original capitalization
- numerical results below are author claims on their evaluated systems
- they are not independent replications
- maximum improvements are not expected improvements on a new workload
- attempted search through two available web tools
- neither returned usable results
- primary conference pages and PDFs were fetched directly over HTTPS
- newest inspected papers are from OSDI 2026
- additional 2026 coverage includes Spice, Quark, and Murakkab
- other 2026 work remains incompletely covered
terms
- serverless: the provider starts and manages execution resources for submitted functions
- cold start: work needed before a newly started function can execute
- workflow: functions connected by data or ordering dependencies
- tail latency: latency among the slowest requests
- checkpoint: saved program state that can be restored
- idempotence: repeating an operation has no additional observable effect
- RDMA: hardware lets one machine access memory on another without ordinary message processing by its CPU
- Pareto-optimal: improving one objective requires worsening another within the modeled choices
what existing work already covers
- startup has several different bottlenecks
- distributing program images
- restoring memory and initialized code
- coordinating runtimes
- contention when many instances start together
- workflow optimization has several different decisions
- when to launch each task
- where to run connected tasks
- how to transfer intermediate data
- how much memory and parallelism to allocate
- scheduling itself consumes resources
- queueing and cache misses can delay placement decisions
- retries require reasoning about concurrent effects
- handling each function in isolation can miss interference between functions
cluster-management foundations and missing baselines
- Borg, Verma et al., EuroSys 2015
- question: how to share a large cluster between long-running services and batch work
- mechanism: admission control, task packing, overcommitment, and process isolation
- evidence: âcombining admission control, efficient task-packing, over-commitment, and machine sharing with process-level performance isolationâ
- §3 describes the master, worker agents, and placement decisions
- inference: useful utilization depends on admitting and isolating work as well as placing it
- implication for proposal C
- preserve placement constraints and isolation when comparing allocator speed
- include batch and latency-sensitive workloads
- limit: an operational account of Googleâs system does not provide a directly reproducible public implementation
- Omega, Schwarzkopf et al., EuroSys 2013
- question: how to support multiple scheduling policies without one monolithic bottleneck
- mechanism: schedulers independently propose changes to shared cluster state
- detect conflicts when changes commit
- evidence: âparallelism, shared state, and lock-free optimistic concurrency controlâ
- evaluation uses Google workloads to compare scheduler architectures and interference
- inference: parallel allocation can exchange CPU bottlenecks for conflicts and repeated work
- implication for proposal C
- measure conflicting placements and retry cost
- retain a shared-state scheduler as an architectural comparison
- limit: minimizing one allocatorâs queue delay does not establish good system-wide commit throughput
- Sparrow, Ousterhout et al., SOSP 2013
- question: how to schedule many short parallel tasks with low delay
- mechanism: decentralized random sampling, batch probing, and delayed task assignment
- evidence: âa decentralized, randomized sampling approach provides near-optimal performanceâ
- evaluation deploys the scheduler on a 110-machine cluster
- inference: detailed global state is not the only route to low scheduling delay
- implication for proposal C
- compare against sampling-based placement
- measure whether richer estimates repay their collection and computation costs
- limit: short-task scheduling and constraint-heavy VM allocation have different state and locality needs
- Dirigent, CvetkoviÄ et al., SOSP 2024
- question: why can orchestration dominate startup after sandbox initialization becomes fast
- mechanism: simplify managed objects, move persistent updates off the invocation path, and combine internal control-plane components
- evidence: âeliminates persistent state updates on the critical path of function invocationsâ
- key condition: exact sandbox placement is hidden from the caller
- the system can relax exact reconstruction of transient sandbox state
- §2 distinguishes platform recovery from recovery of an in-flight request
- §4 implements containerd and Firecracker backends
- implication for proposal C
- compare with eliminating unnecessary control-plane work before optimizing its queue
- implication for proposal B
- recovering the platform is separate from preventing duplicate application effects
- author implementation
- Cloudburst, Sreekanti et al., PVLDB 2020
- question: how to keep mutable state close to dynamically placed functions
- mechanism: Anna storage plus caches beside function executors
- carry state-version information across connected functions
- evidence: âmutable caches co-located with function executors for data localityâ
- §5 defines repeatable-read and causal guarantees across a distributed function session
- inference: moving computation to cached data introduces state-consistency obligations
- implication for proposals A and B
- placement experiments must specify what state versions a workflow may observe
- retry correctness is not established merely by choosing a consistent cache
- limit: these session guarantees should not be silently treated as general transactions
- Faasm, Shillaker and Pietzuch, USENIX ATC 2020
- question: can colocated functions share memory while preserving isolation
- mechanism: WebAssembly memory isolation with explicitly shared memory regions
- Linux controls CPU and network access
- initialized snapshots reduce startup work
- evidence: âallowing memory regions to be shared between functions in the same address spaceâ
- §3 specifies the isolation abstraction and host interface
- inference: workflow data movement and startup costs depend on runtime design
- implication for proposal A
- a container-only experiment may miss a simpler shared-memory solution
- limit: WebAssembly porting and the host interface constrain applicable applications
- ORION, Mahgoub et al., OSDI 2022
- question: how to meet workflow latency probabilities despite correlated execution times, uneven tasks, and cold starts
- mechanism: model dependencies, bundle parallel invocations, and prewarm downstream VMs
- evidence: âhigh variability and correlation in the execution time of individual functionsâ
- further evidence: âpre-warming VMs for subsequent functions in a DAG with the right look-ahead timeâ
- same abstract
- evaluation uses three workflows on AWS Lambda
- inference: correlation-aware cold-start workflow optimization already exists
- correction to proposal A
- correlation and prewarming alone cannot justify a new contribution
- reproduce ORION before claiming a gap in existing scheduling models
- candidate remaining question: changing correlations under shared contention with placement-dependent transfer costs
- this question remains a hypothesis until ORIONâs model and experiments are compared directly
startup literature
- FaaSNet, Wang et al., USENIX ATC 2021
- question: how to distribute container images during a large invocation burst
- mechanism: an adaptive tree distributes image data between execution machines
- fetch only needed image blocks
- authors report provisioning 2,500 containers on 1,000 VMs in 8.3 seconds
- evidence: âfinishes provisioning 2,500 function containers on 1,000 virtual machines in 8.3 secondsâ
- paper abstract, §1 and evaluation
- interpretation: image distribution is a network problem under bursts
- limit: this result does not establish end-to-end startup performance after code initialization and control-plane delay
- MITOSIS, Wei et al., OSDI 2023
- question: can a remote machine reuse initialized state without copying everything first
- mechanism: remote fork backed by demand fetching of memory through RDMA
- kernel changes preserve correct access to the parentâs physical pages
- author claim: âfork over 10,000 new containers from one instance across multiple machines within a secondâ
- paper abstract, §4â§7
- evaluation includes caching, local and remote CRIU, and FaaSNet configurations
- several baselines receive shared lean-container optimizations
- inspect §7 before interpreting comparisons
- limit: requires appropriate RDMA hardware and kernel integration
- research question: what happens when the parent or a remote memory source fails during startup
- unanswered by the quoted performance result
- Sabre, Lazarev et al., OSDI 2024
- question: can compression make saved memory cheaper to restore without consuming too much CPU
- mechanism: hardware compression and decompression plus memory-page prefetching
- author claim: âspeeding up memory restoration from snapshots by up to 55%â
- integrates with Firecracker
- limit: depends on a supported accelerator
- compare accelerator queueing under many simultaneous restores
- memory restoration improvement is not automatically the same as request latency improvement
- AFaaS, Chai et al., OSDI 2025
- question: why do optimized cold starts still become slow in production
- mechanism: reduce coordination overhead, pool runtime resources, and organize reusable initialized states in a tree
- author diagnosis: âa narrow focus on optimizing isolated components of the cold start processâ
- paper abstract, §2
- author production claim: âAFaaS has been deployed in production for over 18 monthsâ
- same abstract
- important distinction: §6 separates component experiments from production observations
- evaluation studies sustained load and simultaneous starts
- §7 discusses security implications of pooling and sharing
- inference: merely improving restore time is already an insufficient research claim
- the full request path and contention must be measured
- authorsâ published trace repository
- availability of the complete production implementation was not established in this review
- ServerlessLLM, Fu et al., OSDI 2024
- question: how to start large models without repeatedly downloading their weights remotely
- mechanism: exploit local storage, optimize checkpoint loading, migrate active inference, and schedule around checkpoint location
- evidence: âschedules the model onto servers that minimize the time to start the inferenceâ
- inference: function scheduling and model scheduling share a locality-versus-queueing decision
- limit: model loading is only one part of serving latency
- generation length, active request load, and migration work must enter any extension
workflow literature
- Caerus / NIMBLE, Zhang et al., NSDI 2021
- question: when should downstream analytics tasks start
- starting early overlaps work but pays for idle waiting
- starting late avoids waiting but can increase completion time
- mechanism: model production and consumption of data within tasks
- choose launch times between those extremes
- author claim: âbeing Pareto-optimal between cost and JCTâ
- paper abstract
- JCT means job completion time
- important scope: §4.2 assumes fixed sequential steps and at most one parent per step
- a task can contain multiple steps with different parents
- evaluation limitation: âWe ensure function invocations are warm to avoid cold-start delaysâ
- same paper, §6, footnote 5
- cost metric is summed task runtime
- not a complete current cloud bill
- inference: cold starts and shared infrastructure delays are a concrete next measurement
- adding uncertainty alone is not enough to claim a new scheduler
- SONIC, Mahgoub et al., USENIX ATC 2021
- question: which intermediate-data transfer method fits each workflow edge
- mechanism: choose remote storage, VM storage, or direct transfer
- placement also accounts for communication
- evidence: âno single data-passing method prevails under all scenariosâ
- evaluation uses three analytics applications on EC2
- inference: a launch-time scheduler should not assume data transfer cost is fixed independently of placement
- limit: requires mechanisms the runtime can actually support
- managed public functions may expose different placement and communication controls
- Faastlane, Kotni et al., USENIX ATC 2021
- question: can connected functions avoid expensive cross-container communication
- mechanism: run functions as threads in one process when possible
- use memory protection keys for isolation
- use processes or additional containers for parallel execution when necessary
- author result: âreduces function interaction latency by up to 99.95% compared to OpenWhiskâ
- inference: colocating tasks changes both communication and isolation decisions
- limit: the quoted component reduction should not be mistaken for the same whole-workflow speedup
- Jolteon, Zhang et al., NSDI 2024
- question: how to select workflow resources while honoring a probabilistic cost or latency bound
- mechanism: combine a structural performance model with learned variability
- solve the resource configuration problem through sampling and convex optimization
- evidence: âsatisfy user-defined cost or latency boundsâ
- paper abstract, §3â§5
- evaluation includes an ML pipeline, video analytics, and TPC-DS Query 95 on AWS Lambda
- inference: variability-aware workflow configuration already exists
- a proposed uncertainty-aware scheduler needs a specific failure of this approach
- open question for experiments: how do bounds behave when several stages slow down together or workload distributions change
- this review does not establish that Jolteon assumes independent stage delays
correctness literature
- Beldi, Zhang et al., OSDI 2020
- question: how to compose stateful functions despite failures
- mechanism: logs, transactions, invocation tracking, and garbage collection
- evidence: âfault-tolerant and transactional stateful serverless functionsâ
- evaluation implements movie review, travel reservation, and social-media applications
- authors evaluate 1,000 AWS Lambdas
- inference: retries and transactional workflows already have substantial systems prior work
- research question: which unlogged external operations remain outside the guarantee
- Flux, Ding et al., OSDI 2023
- question: which operations actually need logging to make retries unobservable
- mechanism: verify individual functions under modeled concurrent interference
- retain logs for operations needed by the proof
- author evidence: âFlux has successfully identified previously unknown issues in 12 applicationsâ
- current implementation scope: âFlux currently supports only Java applicationsâ
- same paper, §1
- further limits in §1
- modeled state is in NoSQL databases
- some unbounded loops are unsupported
- §9 discusses applying the definition beyond serverless
- inference: extending the model to external tools or Rust requires explicit semantics
- translating syntax alone would not establish retry correctness
scheduling literature
- Shinjuku, Kaffes et al., NSDI 2019
- question: how to stop long requests from blocking short requests on a core
- mechanism: microsecond-scale preemption using virtualization hardware
- evidence: âpreempt requests as often as every 5”secâ
- evaluates varied service-time distributions and mixed RocksDB requests
- inference: reducing serverless launch delay cannot remove queueing behind long execution
- limit: this is a specialized operating-system design
- its preemption costs cannot be assumed for an ordinary container runtime
- Gavel, Narayanan et al., OSDI 2020
- question: how to allocate unequal accelerators fairly and efficiently
- mechanism: express policies through an effective-throughput model
- translate policies to heterogeneous hardware
- realize allocations through scheduling rounds
- evidence: âsystematically generalizes a wide range of existing scheduling policiesâ
- inference: accelerator-aware scheduling already includes policy abstraction
- limit: training throughput and interactive request deadlines are different objectives
- CASSINI, Rajasekaran et al., NSDI 2024
- question: can training jobs avoid transmitting over the same link simultaneously
- mechanism: shift the timing of communication phases
- evidence: âthe communication patterns of jobs sharing the same network link are interleavedâ
- evaluates 13 ML models on a 24-server testbed
- inference: placement and network timing can interact with compute scheduling
- limit: irregular agent workflows may lack the repeating communication phases that make this approach useful
- Kamino, Domingo et al., OSDI 2025
- question: which allocator should process a VM placement request
- mechanism: estimate completion delay using queue contents and cached placement computations
- evidence: âassign each new request to the agent with the lowest estimated latencyâ
- author result: âa 42% reduction in average request latenciesâ
- same abstract
- this number comes from a production-trace simulator
- production results are separate in §1 and §6
- do not label the simulator number a measured deployment improvement
- inference: scheduler CPU work and cache locality can become part of the applicationâs startup path
2026 follow-up that changes the research bar
- Spice, Holmes et al., OSDI 2026
- question: why does restoring initialized processes still require expensive work
- mechanism: a snapshot file format and kernel primitive separate storage layout from virtual-memory layout
- restore process metadata in bulk
- author result: âwithin 0.6â18ms of warm-invocation latencyâ
- distinction: these numbers are extra latency above warm execution
- they are not total request latency
- inference: a new restore mechanism must compare with Spice as well as VM snapshot systems
- limit: requires kernel changes
- Quark, Chai et al., OSDI 2026 operational-systems paper
- question: how much apparent CPU utilization is useful batch work
- mechanism: fine-grained allocation, rapid instance provisioning, and scheduling that accounts for unequal hardware and uneven tasks
- author observation: âbatch workloads remain inefficient, with a useful computation ratio of only 67%â
- scope: production analytics colocated with higher-priority online services at Ant Group
- inference: resource occupancy and useful work are different measurement targets
- implication for proposal A
- include colocated online traffic and idle waiting in evaluation
- measure useful CPU work rather than only allocation or billing time
- Murakkab, Chaudhry et al., OSDI 2026
- question: how to optimize entire agent workflows across model and hardware choices
- mechanism: explicit workflow structure, profiling, optimization, and runtime reconfiguration
- evidence: âdecouples workflow specification from execution configurationâ
- inference: simply proposing joint workflow and hardware optimization is already covered
- implication for proposal A
- identify a concrete failure under uncertain delays or retried effects
- compare with Murakkab if extending the workload to agents
- broader model-serving review belongs in the companion agent-systems study
second pass, 7 Oct 2026: what the sections below add
- reading depth: abstracts only, taken from the linked conference or arXiv pages
- every quote below is from the abstract unless a section is named
- numbers are author claims on their own setups
- the first pass covered startup, workflows, retries, and a few schedulers
- this pass adds measurements and traces, isolation units, control plane correctness, GPU serverless, GPU cluster scheduling, CPU scheduling on one machine, overcommit and billing, microservices, multi-cloud, carbon, and schedulers written by learning or LLMs
workload measurements and public traces
- why read these first: a scheduler result only means something on a workload, and these are the workloads people can actually get
- Serverless in the Wild, Shahrad et al., USENIX ATC 2020
- the Azure Functions trace that most cold-start papers replay
- âmost functions are invoked very infrequently, but there is an 8-order-of-magnitude range of invocation frequenciesâ
- my reading: keeping every function warm is wasteful, since most are rarely called
- How Does It Function?, Joosen et al., SoCC 2023
- two Huawei traces, âover 7 months with over 1.4 trillion function invocations combinedâ
- âscheduling time, execution time and cold-start distributions vary across 2 to 4 orders of magnitude and have very long tailsâ
- Serverless Cold Starts and Where to Find Them, Joosen et al., 2024
- âa month-long trace of 85 billion user requests and 11.9 million cold starts from Huaweiâs serverless cloud platformâ
- splits a cold start into âpod allocation time, code and dependency deployment time, and scheduling delaysâ
- âcold starts in Region 1 take up to 7 seconds, dominated by dependency deployment time and scheduling. In Region 2, cold starts take up to 3 seconds and are dominated by pod allocation timeâ
- my reading: which part of a cold start is slow differs by region, so one fast restore mechanism does not fix all of them
- use for proposal A: these per-part delays are the realistic thing to inject
- XFaaS, Sahraei et al., SOSP 2023
- Metaâs private function platform, âtrillions of function calls per day on more than 100,000 serversâ
- âa daily average CPU utilization of 66%â
- âXFaaS defers the execution of delay-tolerant functions to off-pea[k]â hours (the abstract was cut here in my extraction)
- my reading: a private cloud can delay work and trust its callers, a public one cannot, so XFaaS numbers do not transfer to public platforms
- Analysis of Large-Scale Multi-Tenant GPU Clusters, Jeon et al., USENIX ATC 2019
- the Microsoft Philly trace, two months of training jobs
- studies âthe effect of gang scheduling and loca[lity]â on utilization
- gang scheduling: a job starts only when all its GPUs are free at once
- USENIX page
- MLaaS in the Wild, Weng et al., NSDI 2022
- âa two-month workload trace collected from a production MLaaS cluster with over 6,000 GPUs in Alibabaâ
- problems named: âthe low GPU utilization, the long queueing delays, the presence of hard-to-schedule tasks demanding high-end GPUs with picky scheduling requirementsâ
- Characterization of Large Language Model Development in the Datacenter, Hu et al., NSDI 2024
- âa six-month LLM development workload trace collected from our GPU datacenter Acmeâ
- looks at how LLM jobs differ from older deep learning jobs and at âthe impact of various job failuresâ
- Heterogeneity at Hyperscale, Li et al., OSDI 2026 operational-systems paper
- âa six-month trace covering 155,410 GPUs of multiple vendors and generations and jobs from 81 departmentsâ
- âhigh GPU demand does not yield high effective utilization: idle GPUs frequently become unallocatable because free capacity is stranded across nodes, lacks matching CPUs, or violates network-locality constraints, and because users reserve ample headroom for production safetyâ
- âfractional-GPU fragmentation, a focus of prior work, is now negligible, as GPU sharing is rarely usedâ
- this contradicts the starting point of the 2023 paper from the same company, listed under GPU cluster scheduling below
- I have not checked whether this trace is public
- Lifting the veil on Metaâs microservice architecture, Huye et al., USENIX ATC 2023
- âthe topology is extremely heterogeneous, is in constant flux, and includes software entities that do not cleanly fit in the microservice architectureâ
- my reading: benchmarks with a fixed call graph leave out most of what makes real microservices hard
- Mimesys, Kim et al., OSDI 2026
- turns resource usage traces into runnable load, because âproduction workloads are often inaccessible due to privacy and proprietary concernsâ
- âtransforms time-series resource usage traces into executable workloads that emulate resource contention patternsâ
- use: a way to get realistic neighbors for any colocation experiment without the original programs
isolation units: containers, small VMs, unikernels, WebAssembly
- the question in this group is what box to run untrusted code in, and how fast that box can appear
- Firecracker, Agache et al., NSDI 2020
- the small VM monitor under AWS Lambda
- the authors reject the old choice âbetween virtualization with strong security and high overhead, and container technologies with weaker security and minimal overheadâ
- The True Cost of Containing: A gVisor Case Study, Young et al., HotCloud 2019
- gVisor puts a user-space kernel between the container and the host
- measures âgVisor startup performance, memory efficiency, and system-call overheadsâ
- SOCK, Oakes et al., USENIX ATC 2018
- finds Linux container setup itself is slow because of âscalability bottlenecks related to storage and network isolationâ
- âimporting many popular libraries adds about 100ms to startupâ
- starts new instances by forking from a pre-imported parent, which they call Zygotes
- this is one of the âmissing readsâ the first pass listed
- SAND, Akkus et al., USENIX ATC 2018
- runs the functions of one application in one sandbox: âapplication-level sandboxing, and 2) a hierarchical message busâ
- also a âmissing readâ from the first pass. Faastlane and Faasm above are later versions of the same idea
- vHive and REAP, Ustiugov et al., ASPLOS 2021
- an open test platform on Firecracker and Containerd
- âthe execution time of a function started from a snapshot is 95% higher, on average, than when the same function is memory-residentâ
- cause: âfrequent page faults as the functionâs state is brought from disk into guest memory one page at a timeâ
- use: vHive is the open platform I would build proposal A on
- On-demand Container Loading in AWS Lambda, Brooker et al., USENIX ATC 2023
- Lambdaâs targets: âadding up to 15,000 new containers per second for a single customerâ, âstart-up times (as low as 50ms)â, images âas large as 10GiBâ
- mechanism: âcaching, deduplication, convergent encryption, erasure coding, and block-level demand loadingâ
- Unikraft, Kuenzer et al., EuroSys 2021
- a unikernel is one application linked with only the OS parts it needs, booted as a VM
- âimages for these apps are around 1MB, require less than 10MB of RAM to run, and boot in around 1ms on top of the VMM time (total boot time 3ms-40ms)â
- REWIND, Song et al., USENIX ATC 2024
- reusing a warm container leaks data between requests
- âafter each function request, the container is reset to an initial state, free from any sensitive dataâ
- Dandelion, Kuchler et al., 2025
- claims even todayâs platforms âare not sufficiently elastic to avoid over-provisioning expensive resourcesâ
- blames âbooting a guest OS and configuring features like networking in sandboxesâ
- fix: change the interface. Applications become âDAGs of pure compute functions and higher-level communication functionsâ
- my reading: if functions cannot make system calls, the box gets much cheaper. The cost moves to porting applications
- Wasabi, Baqershahi et al., NSDI 2026
- puts WebAssembly programs of different customers inside one container, âa dense hierarchical architecture to securely co-locate Wasm-based applications from different customers within the same container sandboxâ
- keeps the container serving model so non-WebAssembly work still runs
- TrEnv-X, Huang et al., ACM TOCS (extends TrEnv, SOSP 2024)
- ârepurposable sandboxes, which can be shared across different functionsâ, and memory restored from remote memory pools
- extends to âmicroVM-based agent workloadsâ with âbrowser sharingâ
- âWhen applied to LLM agents, it reduces the P99 latency by up to 58% and memory usage by 61% compared to state-of-the-art systems like E2Bâ
- the only paper I found that treats agent sandboxes as a serverless workload. A web search summary says it includes a workload characterization and cost analysis. I have not read that part
- Wallet, Sabanic et al., NSDI 2026
- functions inside confidential VMs, each in âa minimal âtrustletââ
- âa 4.3Ă smaller TCBâ than plain confidential VM deployment
- TCB: the code you must trust
- USENIX page
- Junction, Fried et al., NSDI 2024
- kernel bypass means the application talks to the network card directly, skipping the OS
- âthe first kernel bypass system that can pack thousands of instances on a machine while providing compatibility with unmodified Linux applicationsâ
- same group later wrote Spice (above)
- Pocket, Klimovic et al., OSDI 2018
- short-lived storage for data passed between functions
- âsimilar performance to ElastiCache Redis for serverless analytics applications while reducing cost by almost 60%â
- Burst Computing, Barcelona-Pons et al., USENIX ATC 2025
- for jobs that want many workers at once: âa novel group invocation primitive to launch large groups of workers with guaranteed simultaneityâ
- packs workers into fewer containers so they can message each other directly
- RTSFaaS, Zhao et al., USENIX ATC 2025
- transactions over shared state for function workflows, using leases over RDMA instead of an external database
- âa lease-based concurrency control protocol to dynamically assign and transfer leases among workersâ
- bears on proposal B: a newer transactional baseline next to Beldi
- what I take from this group
- startup of the box is close to solved for one box at a time: 1 ms unikernel boot, Spice within 0.6 to 18 ms of warm
- what is still slow is everything around the box: fetching dependencies, scheduling, network setup, and many boxes at once. The Huawei trace says so directly
- the newest moves change the interface (Dandelion, Wasabi) or the workload (agents in TrEnv-X)
control planes: correct and fast enough
- a controller is a loop that reads the current cluster state and acts to move it toward the wanted state. Kubernetes is built from many of these
- Sieve, Sun et al., OSDI 2022
- tests controllers by âsystematically and extensively perturbing the controllerâs view of the current cluster state in ways it is expected to tolerateâ
- found problems leading to âdata loss, security vulnerabilities, and resource leaksâ
- Acto, Gu et al., SOSP 2023
- tests operators end to end by driving them through sequences of wanted states
- âhas helped find 56 serious new bugs (42 were confirmed and 30 have been fixed) in eleven Kubernetes operatorsâ
- Anvil, Sun et al., OSDI 2024, best paper
- controllers written in Rust and proved with Verus
- the property: âeventually stable reconciliation, written as a concise temporal logic liveness propertyâ
- liveness: something good eventually happens, as opposed to safety, nothing bad ever happens
- âmost work so far focused on safety, whereas reconciliation is fundamentally not a safety propertyâ
- verified âthree Kubernetes controllers for managing ZooKeeper, RabbitMQ, and FluentBitâ
- the closest existing work to the humanâs Verus and Rust interests in this whole area
- I searched for 2025 and 2026 follow-ups and the one search I could run returned none. That is weak evidence, since my search quota ran out
- Kivi, Liu et al., USENIX ATC 2024
- model checks Kubernetes controllers together with their configuration
- âmodels its controllers and events into processes whereby their interleavings are exhaustively checked via model checkingâ
- âtwo new issues in Kubernetes controller source codeâ
- differs from Anvil: Kivi checks a model, Anvil proves the running code
- Mutiny!, Barletta et al., 2024
- injects faults into the store that holds cluster state
- âeven a single fault/error (e.g., a bit-flip) in the data stored can propagate, causing cluster-wide failures (3% of injections), service networking issues (4%), and service under/overprovisioning (24%)â
- KubeDirect, Qi et al., NSDI 2026
- when many function instances start at once, âmessage passing becomes the primary bottleneck as controllers have to exchange extensive state through the API Serverâ
- controllers message each other directly and skip the API server
- cost, in the authorsâ words: âour approach introduces distributed and ephemeral state across controllers, making it challenging to enforce end-to-end semantics without centralized coordinationâ
- together with Dirigent above, this is the second system that makes the control plane fast by dropping the single stored copy of the truth
- this changes proposal C and opens proposal D
- Protean, Hadary et al., OSDI 2020
- Azureâs VM allocator. Kamino above builds on it
- âa multi-layer caching mechanism expedites the allocation process, achieving turnaround times of few millisecondsâ
- âA slight compromise on allocation quality enables multiple AAs to run concurrently on the same inventory, resulting in increased throughput with negligible conflict rateâ
- AA: allocation agent
- USENIX page
- bears on proposal C: Azure reports conflicts are negligible, so measuring conflict cost may show nothing
- Twine, Tang et al., OSDI 2020
- Facebookâs cluster manager: âa single control plane to manage one million machines across all data centers in a geographic regionâ
- lets applications take part in container lifecycle, âe.g., restarting a ZooKeeper deploymentâs followers first and its leader last during a rolling upgradeâ
- found but not read: Garen, âReliable Cluster Management with Atomic State Reconciliationâ, Kim et al., EuroSys 2026
- known only from a lab publication list in a search result. No claims about it here
serverless for GPUs and large models
- the box is no longer the slow part. Loading tens of gigabytes of model weights is
- BlitzScale, Zhang et al., OSDI 2025
- loads weights from other GPUs over the network instead of from disk or host cache
- lets a half-loaded instance already help: âoffload the layer computation from the overloaded serving instances to the scaled ones without waiting for the parameters to be fully loadedâ
- âup to 94 % lower tail latency reductions compared to state-of-the-art autoscaling system (ServerlessLLM)â
- λScale, Yu et al., 2025
- same two ideas, independently: âfast model multicastâ over RDMA and âexecute-while-loadâ
- HydraServe, Lou et al., NSDI 2026
- for public clouds without special networks
- âproactively distributes models across servers to quickly fetch them, and overlaps cold-start stages within workersâ
- âreduces the cold start latency by 1.7Ăâ4.7Ăâ
- DeepServe, Hu et al., USENIX ATC 2025
- Huaweiâs production platform on its own Ascend chips
- âpre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in secondsâ
- âhas been in production for over a yearâ
- Torpor, Yu et al., USENIX ATC 2025
- keeps models in main memory and moves them onto a GPU when a request arrives: âlate binding with model swappingâ
- Prism, Yu et al., OSDI 2026
- many models, most rarely used. Sees âa dynamic bursty-group pattern in which sets of models become active together and shift over timeâ
- shares GPU memory between models by âmemory ballooningâ: taking memory back from an idle model and handing it to a busy one
- âdeployed in production environments across 10K+ GPUsâ
- ServerlessLoRA, Sui et al., 2025
- LoRA: a small add-on to a big shared base model
- â99% of weights are unnecessarily duplicatedâ across functions that share a base model
- LLM-Mesh, Xu et al., 2025
- small and mid-size models with rare requests, packed several to a CPU or GPU
- what I take from this group
- this corner is crowded. Six systems in two years attack the same weight-loading delay
- a new entry needs a workload none of them use. I would not start here
- the LLM serving study covers request scheduling once the model is loaded
GPU cluster scheduling for training
- Gandiva, Xiao et al., OSDI 2018
- training repeats the same small step, so the scheduler can pause between steps cheaply
- âexploits intra-job predictability to time-slice GPUs efficiently across multiple jobsâ
- Tiresias, Gu et al., NSDI 2019
- nobody knows how long a training job will run, so prioritize by how much service a job has had so far
- âimproves the average JCT by up to 5.5Ă over an Apache YARN-based resource manager used in productionâ
- AntMan, Xiao et al., OSDI 2020
- several jobs on one GPU, with the framework changed so jobs can shrink
- âimproves the overall GPU memory utilization by 42% and the computation unit utilization by 34%â
- Pollux, Qiao et al., OSDI 2021
- the scheduler also picks batch size and GPU count for each job
- optimizes goodput, âa novel metric we introduce that combines system throughput with statistical efficiencyâ
- âreduces average job completion times by 37-50%â
- Sia, Jayaram Subramanya et al., SOSP 2023
- Pollux plus several GPU types
- on â44to 64-GPU clusters with a mix of three GPU types, Sia reduces average job completion time (JCT) by 30â93%â
- note the scale: tens of GPUs, while the traces above have thousands to 155,410
- Shockwave, Zheng et al., NSDI 2023
- fairness when a jobâs speed changes during training
- âimproves makespan by 1.3Ă and fairness by 2Ăâ
- Beware of Fragmentation, Weng et al., USENIX ATC 2023
- âallocating partial GPUs can result in severe GPU fragmentation in large clusters, leaving hundreds of GPUs unable to be allocatedâ
- places tasks to keep fragmentation growth smallest
- three years later the same company reports this kind of fragmentation âis now negligible, as GPU sharing is rarely usedâ (Li et al., above)
- XSched, Shen et al., OSDI 2025
- GPUs and other accelerators mostly cannot interrupt a running task
- âa preemptible command queue abstraction (XQueue)â, adapted âto ten XPUs of different types, brands, and generationsâ
- GPreempt, Fan et al., USENIX ATC 2025
- âa timeslice-based yield mechanism to enable context-switch preemption on GPUsâ
- âwithin 40 ÎŒs low-latency preempt[ion]â
- RLBoost, Wu et al., NSDI 2026
- reinforcement learning on LLMs has two stages with different needs. Generating samples scales out on cheap, interruptible GPUs
- âa framework for cost-efficient RL training that harvests preemptible GPU resourcesâ
- OSDI 2026 has three more papers on scheduling this kind of training (DynaRL, Weave, RollArt). I saw only their titles
- what I take from this group
- the academic line (Gandiva to Sia) optimizes job completion time on small clusters with elastic jobs
- the 2026 Alibaba report says the capacity actually lost is stranded whole GPUs and reserved headroom
- those are different problems. I think the gap between them is the research opening, see proposal G
CPU scheduling on one machine
- Shenango, Ousterhout et al., NSDI 2019
- moves cores between applications âevery 5 ”sâ, so idle cores of a latency-sensitive service do batch work
- Caladan, Fried et al., OSDI 2020
- âresource partitioning is neither necessary nor sufficientâ
- reacts to interference by moving cores fast instead of fencing off caches and memory bandwidth
- The Benefits and Limitations of User Interrupts, Guo et al., NSDI 2025
- user interrupts: a newer Intel feature that lets one user program interrupt another without entering the kernel
- âuser interrupts are not a panacea. For example, they provide limited benefits when other software layers constrain the kinds of scheduling policies that canâ be used (cut here in my extraction)
- ALPS, Fu et al., USENIX ATC 2024
- Linuxâs default scheduler âneglects the short-term demands of CPU time from short-lived serverless functionsâ
- learns from past runs to favor short functions without starving long ones
- vBOIDs, Manakkal et al., OSDI 2026
- with many containers on a host, âfine-grained, per-thread scheduling decisions lead to thrashing and unpredictable performanceâ
- schedules groups of threads as one unit
- âimproves the throughput of containerized microservices with thousands of threads by up to 3Ăâ
- kSTEP, Cao et al., OSDI 2026
- a study of Linux CPU scheduler bugs, âcovering both functional violations and misalignments between implementation and policyâ
- âthese bugs are hard to observe and even harder to triggerâ
- runs scheduler events deterministically on isolated CPUs and fuzzes on top: âreproducing seven real-world scheduler bugs and uncovering four new onesâ
- the one paper here on scheduler correctness as opposed to scheduler speed. Feeds proposal E
- What Are You (M)Waiting For, Wang et al., OSDI 2026 operational-systems paper
- mwait: the CPU instruction a guest uses to go idle. Letting the guest run it directly is fast, but hides idleness from the host
- âa vCPU never yields its pCPU, causing idle vCPU to monopolize coresâ
- âeven an idle vCPU executing mwait can raise colocated tail latency by up to 3Ăâ
- my reading: an optimization tuned without overcommit broke under overcommit. A good example of why proposal E asks for stated properties
overcommit, harvesting, autoscaling, and billing
- overcommit: promising more CPU or memory than the machine has, betting that not everyone uses their share at once
- Harvest VMs, Ambati et al., OSDI 2020
- a VM that âgrows and shrinks according to the amount of unallocated resources at its underlying serverâ
- gives guarantees from predictions, âe.g., 65% of the Harvest VMs will survive more than a weekâ
- HarvestContainers, Hall et al., NSDI 2026
- the same idea for Kubernetes containers, with no application changes
- âdynamically determines the safe number of CPU cores to harvestâ
- Jiagu, Liu et al., USENIX ATC 2024
- overcommit needs a performance prediction per placement, and predicting is slow
- âpre-decision scheduling achieves accurate prediction while eliminating overheads by decoupling prediction and schedulingâ
- âa 54.8% improvement in deployment density over commercial clouds (with Kubernetes)â
- bears on proposal C: another paper about scheduler decision cost
- Leopard, Cao et al., NSDI 2025
- âcurrent billing practices do not align with true resource consumptionâ
- a billing model with âvarying CPU and memory demands, spot cores, and preemptible memoryâ, wired into the scheduler and admission
- partly closes the âbilling gapâ the first pass listed
- DVLA, Zhang et al., OSDI 2026 operational-systems paper
- placing VMs by predicted lifetime goes wrong as the lifetime mix drifts
- a few long-lived VMs scattered across machines stop those machines from ever being emptied: âa persistent long-lived VM placement debtâ
- this âcannot be repaid by online scheduling aloneâ, so they add offline moves
- Uberâs Failover Architecture, Bansal et al., NSDI 2026
- old rule: every service has enough capacity in two regions, so half sits idle
- new rule: only critical services keep that. Others borrow the spare capacity and are shut off during a failover
- âreduces steady-state provisioning from 2Ă to 1.3Ă, raising utilization from 20% toward 30%â
- my reading: 20% to 30% utilization at a company this size shows how much room ordinary services still leave
microservices: resource control and overload
- DeathStarBench, Gan et al., ASPLOS 2019
- the benchmark almost every paper below uses
- âan open-source benchmark suite built with microservices that is representative of large end-to-end servicesâ
- the Meta study above suggests real deployments look quite different
- FIRM, Qiu et al., OSDI 2020
- learning finds which service causes a latency violation and which resource it lacks
- âreduces SLO violations by up to 16x while reducing the overall requested CPU limit by up to 62%â
- Autothrottle, Wang et al., NSDI 2024
- a learned controller sets a target per service, and a simple local rule meets it
- âCPU savings, up to 26.21% over the best-performi[ng]â baseline (cut here in my extraction)
- Rajomon, Xing et al., NSDI 2025
- overload control with prices: âClients attach tokens to requests and services charge a price for each API, dropping requests with insufficient tokensâ
- Galileo, Saxena et al., NSDI 2026
- learned controllers act without knowing how fragile their choice is
- adds âstatistical bounds on tail latencies of specific request types under a range of environmental perturbationsâ, from a queueing model
- Slowpoke, Xie et al., NSDI 2026
- answers âwhat if this service were twice as fastâ without building the speedup, by slowing the others
- âa root mean squared error of only 2.07%â
- Metastable Failures in Distributed Systems, Bronson et al., HotOS 2021
- a metastable failure: a trigger pushes the system into a bad state, and the system keeps itself there after the trigger is gone. Retry storms are the usual example
- âA systematic approach for building systems that are robust against unknown metastable failures remains an open problemâ
- Metastable Failures in the Wild, Huang et al., OSDI 2022
- âat least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failuresâ
- CSnake, Qian et al., 2025
- finds self-sustaining cascades before deployment by linking single fault injections from different tests into one chain
- belongs mostly to the bug-finding studies
- Aletheia, Ferreira et al., OSDI 2026
- data split across services loses the consistency checks a single database gave
- static analysis âon 7 open-source applications, detecting 46 previously unreported integrity violationsâ
several clouds, spot machines, and cost
- From Cloud Computing to Sky Computing, Stoica and Shenker, HotOS 2021
- sky computing: using many clouds as if they were one service
- âThe barriers are more economic than technical, and we propose reciprocal peering as a key enabling stepâ
- The Sky Above The Clouds, Chasins et al., 2022
- a 35-page position paper from the same group on how the cloud market could mature
- SkyPilot, Yang et al., NSDI 2023
- a broker that picks a cloud per job: âcreating a fine-grained two-sided market via an intercloud brokerâ
- Canât Be Late, Wu et al., NSDI 2024
- spot instance: a cheap machine the provider can take back at any time
- when to switch to full-price machines so a job still meets its deadline
- the policy âis parameter-free and requires no assumptions on spot availabilityâ, tested on âthree-month-long real spot availability traces on AWSâ
- Starburst, Luo et al., USENIX ATC 2024
- own cluster plus cloud overflow. Decides how long a job waits for the cluster before paying for cloud
- âassigns longer waits for large jobs to increase their chances of running on the cluster, and shorter waits to small jobsâ
carbon, energy, and power
- Letâs Wait Awhile, Wiesner et al., Middleware 2021
- grid electricity is cleaner at some hours. Delay work that can wait
- studies âGermany, Great Britain, France, and California over the year 2020â
- On the Limitations of Carbon-Aware Temporal and Spatial Workload Shifting, Sukprasert et al., EuroSys 2024
- the skeptical paper. Carbon data âfrom 123 regionsâ
- âthe practical upper bounds of these carbon reductions are currently limited and far from idealâ
- âsimple scheduling policies often yield most of these reductions, with more sophisticated techniques yielding little additional benefitâ
- CarbonScaler, Hanafy et al., 2023
- run a batch job on more servers when power is clean and fewer when it is dirty
- â51% carbon savings over carbon-agnostic executionâ
- GREEN, Xu et al., NSDI 2025
- âup to 41.2% reduction in cluster-wide carbon footprint and 12% reduction in peak power consumption, while incurring 3.6%-5.9% time efficiency tradeoffâ
- SPADE, Lechowicz et al., OSDI 2026
- outside signals such as âenergy cost, carbon intensity, power availability, and water usageâ now decide how much compute is available
- for jobs made of dependent tasks, âdelaying certain tasks in the DAG (e.g., bottleneck tasks) can stall entire pipelinesâ
- decides how many machines and which tasks together
- bears on proposal A: same shape of problem (when to run each task of a dependent job) with supply changing instead of delay
- Hardware Lifecycle-Aware Power Planning, Li et al., OSDI 2026 operational-systems paper
- Metaâs rack power budgets, from âlive production traffic data spanning millions of servers across multiple hardware generationsâ
- what I take from this group
- the EuroSys 2024 result is a warning: most of the carbon saving comes from simple policies, and the ceiling is low
- 2026 work reframes it as power supply that varies, which providers care about for money reasons. I think that framing will last longer than the carbon one
schedulers written by learning or by LLMs
- Decima, Mao et al., SIGCOMM 2019
- reinforcement learning trains a neural network that schedules Spark jobs
- âimproves the average job completion time over hand-tuned scheduling heuristics by at least 21%â
- the policy is a network nobody can read
- Barbarians at the Gate, Cheng et al., 2025
- LLMs write the policy as code, a simulator scores it, repeat
- the argument: âsystem performance problems naturally admit reliable verifiers: solutions are typically implemented in real systems or simulators, and verification reduces to running these software artifacts against predefined workloads and measuring performanceâ
- case studies include âload balancing for multi-region cloud schedulingâ
- note what âverifierâ means here: a benchmark score. It says nothing about inputs outside the benchmark
- PolicySmith, Dwivedula et al., HotNets 2025
- âapplies LLMs to synthesize instance-optimal heuristicsâ
- instance-optimal: tuned for one deployment and its workload
- for congestion control, âcan generate safe policies that integrate directly into the Linux kernelâ
- âapplies LLMs to synthesize instance-optimal heuristicsâ
- Glia, Hamadanian et al., 2025
- several LLM agents reason, run experiments, and analyze
- on a GPU cluster for LLM inference âit produces new algorithms for request routing, scheduling, and auto-scaling that perform at human-expert levels in significantly less timeâ
- SchedCP, Zheng et al., 2025
- LLM agents write Linux scheduler policies as eBPF programs and load them through sched_ext
- eBPF: small programs the kernel checks and then runs inside itself
- has âan Execution Verifier that validates all AI-generated code and configure before deployment with static and dynamic analysisâ
- âup to an 1.79x performance improvementâ
- LLM agents write Linux scheduler policies as eBPF programs and load them through sched_ext
- what I take from this group
- within one year, writing a scheduling policy became cheap
- every one of these systems accepts a policy because it scored well on a workload, plus at most a check that it will not crash the kernel
- none of the five abstracts claims a policy is free of starvation or keeps cores busy when work is waiting. kSTEP shows hand-written schedulers already get those wrong
- so the scarce thing is now checking, which is what the humanâs verification work is about. See proposal E
research proposals
- A: workflow schedules that survive shared cold-start and network slowdowns
- hypothesis: joint placement and launch decisions beat separate resource configuration and launch-time decisions under correlated stalls
- closest work: ORION, Caerus, SONIC, Jolteon, AFaaS
- prior-work constraint: ORION already models correlation and prewarms downstream VMs
- first experiment
- reproduce eager, lazy, and NIMBLE-style scheduling in one controlled runtime
- compare ORION prewarming and bundling plus Jolteon-style configuration
- use the same available information and resource budget
- inject warm starts, isolated cold starts, burst cold starts, and shared network congestion
- vary workflow width, intermediate-data size, and contention
- measure
- median and p99 workflow latency
- deadline misses
- billed execution time, storage operations, and transfers
- prediction error and scheduler overhead
- candidate design
- share measured startup and data-transfer distributions across placement and launch decisions
- preserve explicit dependencies and bounded resource demand
- falsifier
- gains vanish after matching resource budgets and including scheduler overhead
- ORION or Jolteon handles the tested variability equally well
- novelty gate
- inspect later Caerus and Jolteon citations and cold-start-aware workflow work before building a new optimizer
- B: retry-safe functions with uncertain external effects
- hypothesis: explicit operation contracts can verify useful retry guarantees with less logging than logging every step
- closest work: Beldi and Flux
- starting scope
- Rust functions calling a small modeled key-value API and one external-effect API
- explicit request identifiers and acknowledgement loss
- concurrent calls on shared objects
- first experiment
- implement payment and reservation examples
- interrupt after every externally visible operation
- retry while another function changes shared state
- compare blanket logging, selective logging, and contract-based deduplication
- measure
- duplicate effects and forbidden state outcomes
- proof effort, supported programs, runtime overhead
- falsifier
- contracts assume the desired property rather than establish it
- non-atomic external APIs make the claimed guarantee impossible
- supported examples are too restricted to improve on Flux
- novelty gate
- examine workflow-engine retry semantics and external transaction protocols
- share findings with the verification and bug-finding studies
- C: scheduling around the cost of making scheduling decisions
- hypothesis: queue and cache-aware routing improves resource startup during bursts with fewer allocator replicas
- closest work: Kamino, Dirigent, Omega, Sparrow
- prior-work constraint: simpler orchestration may remove the bottleneck before improved routing is necessary
- candidate distinction
- several control-plane services share CPUs and caches
- placement requests have deadlines tied to application demand
- first experiment
- reproduce a Kamino-style latency estimator in a trace-driven simulator
- add controlled interference and changing placement rules
- compare random routing, shortest queue, cache affinity, sampling, and latency estimation
- distinguish routing cost from unnecessary persistent updates and state conflicts
- measure
- allocation p99, deadline misses, stale predictions, cache memory, and achieved placement quality
- falsifier
- a simple shortest-queue policy matches the proposed design
- application benefit disappears because image loading dominates
- practical limit
- realistic allocator traces and rules may be unavailable
- synthetic traces alone cannot support a production-scale claim
recommended sequence
- recommendation: run proposal Aâs measurement before designing a new scheduler
- low-cost outcome: a reproducible account of when existing assumptions fail
- stronger outcome: a specific interaction that existing baselines cannot handle
- pursue B if the verification group can support the modeled operation semantics
- pursue C only after establishing useful workloads and accessible traces
- keep independent hardware contributions separate
- RDMA and compression acceleration may be valuable
- they introduce equipment and kernel dependencies before the central scheduling question is answered
remaining literature checks
- recent gap: later work citing Dirigent
- runtime gap: SAND, SEUSS, and Catalyzer
- workflow gap: later work citing ORION and Jolteon
- billing gap: current provider behavior and complete cost accounting
- these are explicit missing reads
- this document makes no supported technical claims about them
Last edited: