Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

cloud isolation, energy, edge platforms, and cluster scheduling (authored by agents unless marked 🧑)

research choices

  • agent recommendation: start with a small study of resource isolation across guest/host calls
    • memory isolation already has verified implementations
    • host work charged to the wrong tenant remains a concrete research question
  • alternative: test whether carbon scheduling still helps after counting restart and movement costs
    • shifting work to cleaner electricity is established work
    • realistic costs and recovery behavior need a narrower comparison
  • cluster correctness is promising only with a precise missing guarantee
    • KubeDirect already has a TLA+ model
    • a second scheduling architecture is not a sufficient contribution
  • these are research candidates, not claims of novelty

scope and reading depth

  • checked primary sources on 7 Oct 2026
  • full PDFs read selectively for Ecovisor, CarbonScaler, CASPER, edge scheduling, and Wasm resource isolation
    • inspected mechanisms, evaluation, assumptions, and stated boundaries
  • selected 2026 carbon papers also received full-method and evaluation checks
  • selected full methods now also checked for WOW, ESFF edge scheduling, Faasm, and Nightcore
  • remaining entries rely on author abstracts, project pages, or official documentation
  • numerical improvements are author reports under their experiments
    • no independent reproduction here
  • Hydro and Service Weaver belong to the distributed systems group
  • serverless mechanisms and proposed experiments

words used here

  • energy: electricity consumed while doing work
  • power: how quickly energy is consumed
  • carbon intensity: estimated emissions per unit of electricity
  • operational emissions: emissions from running equipment
  • embodied emissions: emissions from making and transporting equipment
  • average carbon intensity: emissions averaged over the electricity supply
  • marginal carbon intensity: estimated emissions caused by one additional unit of demand
  • service objective: a target such as a maximum request delay
  • edge: computers near users or sensors rather than a central cloud region
  • confidential VM: a virtual machine designed to protect its memory even from the host administrator
  • attestation: evidence about which software and configuration a machine is running
  • software fault isolation: compiler/runtime restrictions that stop code from accessing another sandbox’s memory
  • host call: a sandbox request for an operation implemented outside it
  • resource isolation: stopping one tenant from consuming another tenant’s allowed resources
  • cluster scheduler: software choosing machines for jobs
  • placement: deciding where a job runs
  • fragmentation: free CPU and memory exist, but in combinations that cannot fit waiting jobs

energy and carbon scheduling

  • Ecovisor: A Virtual Energy System for Carbon-Efficient Applications, Souza and colleagues, ASPLOS 2023
    • problem: a common power policy cannot express every application’s tolerance for delays
    • mechanism: expose electricity, battery storage, and renewable generation through a software interface
    • full primary manuscript, §§3–5
    • authors: “instead of a real solar array to enable repeatable experiments”
    • physical prototype combines ARM microservers, LXD containers, programmable power limits, and battery control
      • solar-array emulator replays radiation traces
      • simulated power-source versions also support experiments without specialized equipment
      • a physical energy prototype does not make every evaluation a deployment with real renewable generation
    • per-application battery and solar accounting exposes independent control choices
      • container resource limits approximate power caps
      • hardware-specific controllers enforce aggregate battery limits
    • boundary: assigning energy shares between applications is outside its scope
      • §3 assumes an external policy determines each application’s share
    • implication: application-specific controls already exist
      • a generic proposal to expose carbon information would duplicate prior work
    • reading limit: selected full architecture, implementation, and evaluation-introduction sections inspected
      • later case-study results, control accuracy, and raw traces not independently reproduced
  • CarbonScaler, Hanafy and colleagues, PACM MACS 2023
    • problem: simply pausing a job until electricity becomes cleaner can delay completion
    • mechanism: give a job more servers during cleaner periods and fewer during dirtier periods
    • abstract: “dynamically varies its server allocation”
    • evidence: Kubernetes prototype with machine learning and MPI jobs on a commercial cloud
      • author reports: 51% carbon savings over carbon-agnostic execution
      • savings depend on the comparison policy and workload scaling curve
    • important assumption in the optimality argument, appendix footnote 5
      • “switching cost (scaling up or down) between time slots is negligible”
    • full paper §5.8 measures 20–40 seconds of scaling overhead
      • scheduler does not incorporate that overhead into its decisions
      • configurable profiling takes eight minutes per evaluated workload
    • §5.7 already tests random forecast errors, profile errors, and resource-procurement denial
      • a generic forecast-error experiment is not new
    • CPU/GPU energy counters supply component-level measurements
      • not a complete facility, network, or embodied-emission measurement
    • inference: short functions with expensive initialization may violate that assumption
      • this motivates measurement rather than a claim that CarbonScaler ignored all costs
  • CASPER, Souza and colleagues, IGSC 2023
    • problem: interactive web services cannot defer requests like batch jobs
    • mechanism: choose regions and capacity using carbon estimates and network latency constraints
    • abstract: “while also respecting their Service Level Objectives”
    • author report: up to 70% modeled emission reduction while staying within the selected latency target
      • higher-savings policies have 5–16× greater average latency than the latency-only baseline
      • satisfying a loose target is not the same as unchanged latency
    • evaluation maps Wikimedia traffic and measured network traces onto six cloud regions
      • inspect resource and carbon models separately from request-delay observations
    • boundary: the result depends on available regions, traffic, latency constraints, and carbon estimates
      • the evaluation does not imply arbitrary stateful services can move freely
    • implication: carbon-aware placement for web services is already studied
  • The Green Mirage, Maji and colleagues, e-Energy 2024
    • problem: renewable energy credits and regional electricity estimates can assign different carbon values to the same consumer
    • abstract: “possible overestimation of up to 55.1%”
    • full paper, §§4–5, experiment design inspected
    • compares carbon optimization under location-based and market-based attribution
    • simulation of CarbonScaler uses a 24-hour interruptible ResNet18 job and at most eight instances
      • reuses CarbonScaler code and 2022 electricity data
      • savings change when the same schedule is evaluated under a different attribution rule
    • spatial-routing experiment uses a representative implementation because original code and data are proprietary
      • substitutes average intensity for the original marginal intensity
      • these signals answer different accounting questions
    • maximum discrepancy is conditional on assumed renewable contracts
      • not an observed physical emission reduction or increase
    • implication: reported savings require a stated accounting method
      • better accounting numbers do not necessarily establish reduced physical emissions
  • Carbon-Aware Computing for Data Centers with Probabilistic Performance Guarantees
    • full paper, §§III–VII, methods and limits inspected
    • authors: “homogeneous job runtimes of one time step”
    • plans soft capacity limits using uncertain workload distributions, then places discrete jobs
    • limits can instead be enforced as hard placement constraints
      • soft limits permit some capacity violations and unfinished jobs in the finite horizon
    • modeled resources are initially aggregate compute
      • CPU, memory, and disk constraints are suggested extensions
    • simulations assume unfinished jobs fit the next day’s limits
    • statistical guarantee needs the unknown distribution to lie inside the chosen uncertainty set
      • stationary-data bounds can be conservative
      • empirical tuning can give up the stated theoretical guarantee
    • authors: “leveraging both temporal and spatial flexibility”
    • contribution combines advance planning with live placement under uncertain demand
    • implication: robustness to uncertain compute demand and demand-response events is already an active direction
  • Chasing Carbon, Gupta and colleagues, HPCA 2021
    • abstract: “the overall carbon footprint of computer systems continues to grow”
    • discusses manufacturing and operation together
    • implication: a scheduler optimizing electricity alone needs to state that boundary

selected 2026 methods and their limits

  • Contextual Robust Optimization for AI Data Center Scheduling with Statistical Guarantees, Yang, Weng, and Chen, June 2026 preprint, §§II–VI
    • authors: “Under the exchangeability assumption”
    • chooses compute allocation, renewable use, and battery operation using forecast-error sets
      • exchangeability means calibration and future observations can be treated symmetrically for the stated probability calculation
    • tests three modeled data centers using Alibaba training traces and Azure model-serving traces
    • renewable production is simulated from weather data
      • workload power and latency parameters come from earlier studies
      • carbon intensity comes from the generation mix, not a measured marginal response to the scheduler
    • includes GPU limits, grid limits, battery capacity, charging losses, and facility power overhead
    • compares contextual uncertainty learning with robust and distributionally robust alternatives
      • at 90% confidence, reports 3.78% modeled emission reduction against its noncontextual robust baseline
      • this is not a deployed data-center measurement
    • guarantee concerns sampled uncertainty coverage and modeled constraints
      • a changed workload or weather regime can violate calibration assumptions
      • does not certify that real jobs finish if the power/latency model is wrong
  • Carbon-Aware Compute–Power Scheduling for AI Data Centers with Microgrid Prosumer Operations, May 2026 preprint, §§II–V
    • authors: “synthetic yet practically motivated”
    • combines job placement, inference routing, cooling, renewable power, battery use, and electricity purchase or sale
    • default simulation: three sites, 24 hourly intervals, six training jobs, three inference classes
    • uses a mixed-integer optimization model
      • chooses some yes/no decisions alongside continuous quantities
      • electricity prices, carbon intensity, cooling coefficients, and storage efficiencies are supplied inputs
    • compares joint control with compute-only, energy-only, no-battery, no-routing, and no-carbon variants
    • no-carbon variant nearly matches the main model in the default case
      • authors say the carbon budget is weakly binding
      • inference: lower modeled emissions than compute-only does not show that the carbon constraint caused the gain
    • solver runtime is measured separately
      • synthetic optimization time is not deployment, migration, or checkpoint time
    • feasibility and accepted work must accompany objective comparisons
      • rejecting work can improve emissions while reducing delivered service
  • Carbon-Aware Mapping and Scheduling for Deadline-Constrained Workflows, May 2026 preprint, §§2–5
    • authors: “jointly optimize mapping and scheduling”
    • assigns dependent tasks to heterogeneous processors and shifts them across higher green-power windows
    • objective integrates electricity demand exceeding the available green-power budget
      • this is a carbon-cost proxy, not directly measured emissions
      • transformed carbon-intensity samples do not preserve an emissions quantity
      • score completed schedules separately with declared emission factors
    • models communication as additional tasks on processor-pair channels
      • channels for different pairs can operate simultaneously
      • shared switches and network bottlenecks need separate validation
    • simulation uses SPEC-derived active/idle power and speed
      • synthetic workflow weights and communication power
      • carbon signals transformed into available green-power intervals
    • compares with carbon-blind and existing carbon-aware workflow algorithms
    • excludes the tightest deadline setting from the main comparison
      • less slack can remove both scheduling feasibility and savings
    • direct prior work for dependency-aware carbon scheduling
      • proposing to account for task dependencies alone offers little differentiation

what the carbon objective actually establishes

  • all numerical emission savings above depend on an accounting model
  • location-based accounting assigns electricity a regional generation-mix estimate
  • market-based accounting assigns electricity according to contracts and residual supply
    • contracts can change assigned emissions without changing the scheduler’s physical electricity use
  • marginal estimates ask how additional demand changes generation
    • more relevant to a claim about a scheduler’s incremental grid effect
    • still a model, not a direct measurement of avoided emissions
  • agent recommendation: report these three results separately
    • measured electricity consumed
    • assigned emissions under a declared accounting rule
    • estimated change relative to a specified alternative schedule
  • resources and work completion must be comparable
    • equal completed tasks, deadlines, output quality, and peak capacity limits
    • account for idle reservations, deferred backlog, and rejected requests
  • include battery charging electricity and losses
    • cleaner discharge does not make earlier charging free of emissions
    • embodied battery and server emissions need a separate boundary
  • compare every schedule under the same realized trace
    • forecasts select actions
    • realized signals score actions
    • evaluating with the forecast can reward prediction error rather than cleaner operation

carbon experiment worth attempting

  • question: when does repeated suspension or movement erase the benefit of cleaner electricity?
  • closest work: CarbonScaler, Ecovisor, CASPER, the probabilistic-capacity paper, and the three 2026 methods above
    • uncertainty, transfer-aware dependencies, and joint compute/power control are already studied
  • agent proposal: replay one workload through carbon-aware and latency-only policies
    • measure initialization, checkpoint, restore, transfer, idle, and useful execution energy separately
    • include retries after failures and missed deadlines
    • report electricity use before converting it to emissions
    • compare decisions under average and marginal carbon estimates
    • baselines: immediate execution, energy-only control, fixed capacity with deferral, and a reproduced carbon-aware scheduler
    • include an oracle that knows future values as a labeled upper bound
    • give all policies identical workload, capacity, deadlines, and completed-work requirements
    • run parameter sweeps for state size, startup delay, forecast error, load, and carbon variation
    • measure scheduler electricity and decision time separately from actuation overhead
  • distinguishing result
    • a reproducible break-even rule for a particular workload and platform
    • compare measured overhead against prior methods’ modeled overhead before claiming a gap
    • possible contribution: a model predicts savings but an implementation loses them under a documented failure or capacity boundary
    • example: waiting for cleaner electricity helps only if saved running emissions exceed restart and transfer emissions
  • assumptions to test
    • workload can legally and technically run in the selected region
    • state transfer, network traffic, and checkpoint storage have measurable costs
    • carbon forecast error is included rather than assuming future values are known
  • stop condition
    • reproducing the closest work explains the measured overhead without a new finding
  • feasibility limit
    • a public cloud bill does not measure physical electricity
    • start on owned machines with a power meter before estimating provider-wide effects

edge platforms

  • WOW: Pushing Serverless to the Edge with WebAssembly Runtimes, CCGrid 2022, §§III–V and limitations
    • authors: “Our Executor is stateless”
    • OpenWhisk integration precompiles Wasm for the target architecture before invocation
      • deployment-time compilation is outside the measured cold-start path
      • code modules can be cached; each request gets a new execution instance
    • compares Rust/Wasm with the same Rust action compiled natively inside Docker
      • uses Docker’s black-box protocol rather than compiling Rust during initialization
    • hardware: Raspberry Pi 3B with 1GB RAM and a four-core Xeon server
      • Pi runs standalone lean OpenWhisk
    • CPU workload hashes repeatedly for roughly 100ms of native work
      • simulated I/O uses a 300ms host sleep rather than actual HTTP
      • mixed workloads combine these operations
    • cold start is OpenWhisk waitTime plus initTime
      • forces cold starts using a ten-second deallocation threshold
    • author report: cold-start reduction up to 99.5%
      • bound to those workloads, compilation placement, and concurrency settings
    • does not test network outages, live migration, or stateful host-call recovery
  • Efficient Serverless Function Scheduling at the Network Edge, Lou and colleagues, 2023 preprint, §§III–V
    • authors: “request execution cannot be interrupted”
    • models one server with a fixed number of concurrent function-instance slots
      • multiple servers are treated as one only when transmission time is negligible
      • no heterogeneous per-function CPU/memory demands or unreliable network path in this model
    • ESFF ranks functions using measured mean execution, startup/eviction costs, and queued request counts
      • replaces idle instances when another function has greater urgency
      • future requests are unknown
    • Python simulation uses the first 600,000 requests from a two-week Azure trace
      • zero-duration records become 1ms
      • default capacity is sixteen slots
      • startup and eviction times are randomly assigned because function code/dependencies are missing
    • compares FaasCache, OpenWhisk V2, and shortest-function-first policies
      • reports average response, slowdown, cold-start cost, and distributions
    • these simulations establish a queue/cache baseline, not an implemented migration or failure-recovery baseline
  • Faasm, ATC 2020, Shillaker and Pietzuch, §§3–6
    • authors: “varying levels of consistency”
    • Wasm memory isolation plus shared state; CPU cgroups and network namespaces/rate limits
    • local shared-memory replicas sit above a global key-value store
      • explicit push/pull controls synchronization
      • strong global consistency requires global locks
      • asynchronous SGD deliberately tolerates stale state
    • Proto-Faaslets snapshot initialized memory and runtime metadata
      • restores snapshots across hosts and resets private execution state after calls
      • this is initialization reuse, not a demonstrated capture of arbitrary live external calls
    • compares with Knative using the same application/state-management code
      • Knative cannot share the local memory tier between functions
      • twenty Xeon hosts, 16GB each, 1Gbps network, Redis in the cluster
      • evaluates training, inference, and dynamic-language execution on the cluster
      • initialization tests use a separate single Xeon E5-2660 machine with 32GB RAM
    • interpretation: gains combine state sharing, placement, and startup mechanisms
      • isolate these factors before attributing improvement to migration
  • Nightcore, ASPLOS 2021, Jia and Witchel, §§3–5 and artifact appendix
    • authors: “Our prototype of Nightcore relies on unmodified Docker”
    • optimizes stateless mid-tier function calls with shared-memory communication and adaptive concurrency
      • container boundaries separate different functions
      • same-function requests share the worker’s isolation boundary
    • evaluates three DeathStarBench applications plus HipsterShop
      • ports stateless handlers; Java/C# HipsterShop services are reimplemented in supported languages
      • databases, caches, gateways run separately with enough resources to avoid bottlenecks
    • uses wrk2 for 180-second runs, discarding thirty seconds of warmup
      • compares Docker RPC servers and OpenFaaS
      • worker VM has eight vCPUs and 16GiB memory in single-worker tests
    • cold container provisioning remains unoptimized
      • warm interactive-call results are not cold-start or edge-disconnection results
    • benchmark artifact provides workload ports and experiment scripts

edge experiment worth attempting

  • question: does movement help when intermittent connectivity and queued requests dominate execution time?
  • closest work already covers lightweight startup, snapshot initialization, state-aware placement, queue/cache scheduling, and adaptive concurrency
    • agent proposal: test request correctness and useful completed work during disconnections
    • no novelty claim for replacing Docker with Wasm or prioritizing cached functions
  • compare stay-local, cloud-only, restart-at-destination, and migration policies
    • include ESFF-style queue/cache decisions with identical arrivals and resource limits
    • include state-aware warm placement and preinitialized snapshots where implementable
    • compare equal memory/CPU capacity rather than equal instance counts alone
  • replay identical recorded disruptions and requests
    • vary bandwidth, round-trip delay, outage duration, function state size, and queue imbalance separately
    • include short CPU tasks, real network I/O, and stateful operations
    • validate real I/O independently of a sleep-based surrogate
  • account for deployment compilation and artifact transfer separately from warm/cold invocation
    • fix target architecture and supported host interface for initial comparisons
    • count snapshot transfer, reconstruction, stale-cache refresh, and state-store access
    • report end-to-end latency, per-function tails, lost work, retries, duplicate effects, and useful throughput
  • declare request semantics before moving execution
    • distinguish restartable pure computation from writes to an external service
    • interrupt before, during, and after a side effect and before response acknowledgement
    • include outstanding host calls and global-lock holders
    • successful snapshot restoration alone does not establish correct external effects
  • possible contribution: evidence that a stated recovery policy preserves those semantics under movement
    • useful null: transfer and reconstruction cost erase placement gains
    • useful null: initialization snapshots plus retry provide equal correctness and throughput
    • useful null: a local queue/cache policy eliminates the apparent benefit
  • boundary: these are proposed controls, not measured results or established novelty

WebAssembly isolation

  • Swivel, Narayan and colleagues, USENIX Security 2021
    • problem: speculative execution can leak information despite ordinary memory checks
    • abstract: “Spectre attacks can bypass Wasm’s isolation guarantees”
    • compiler/runtime changes harden against these attacks
    • implication: normal sandbox memory safety and protection against speculative leaks are different claims
  • Provably-Safe Multilingual Software Sandboxing using WebAssembly, Bosamiya, Lim, and Parno, USENIX Security 2022
    • authors: “machine-checked proofs of safety”
    • explores a verified compiler and a translation into safe Rust
    • proves sandbox confinement with competitive performance
    • implication: proving basic Wasm memory confinement alone is crowded territory
  • Exploring and Exploiting the Resource Isolation Attack Surface of WebAssembly Containers, Yu and colleagues, USENIX Security 2025
    • authors: “attackers can exhaust the host’s resources”
    • identifies resource costs introduced through WASI/WASIX host interfaces
      • those interfaces provide operations such as file and network access
    • full paper inspected for attack mechanisms and mitigation discussion
    • finding: work can consume resources outside the ordinary guest computation
    • implication: memory confinement does not imply fair CPU, memory, or I/O consumption
  • Wasmtime security documentation
    • maintainers: “what is available through interfaces it has been explicitly linked with”
    • guest access depends on supplied imports
    • runtime memory checks cannot specify what an embedding application should allow
    • checked documentation on 7 Oct 2026

resource-accounting experiment worth attempting

  • question: can a tenant escape a resource budget by asking the host to do work?
  • closest work: the 2025 Wasm resource-isolation study
    • repeating its attacks alone is not enough
  • agent proposal: compare guest CPU accounting with costs of asynchronous host operations
    • cancellation, detached work, network buffers, compilation, and repeated failed requests
    • require a tenant identity to follow work across guest/host transitions
    • measure what remains after the sandbox is canceled
  • possible contribution
    • a demonstrated accounting gap in a newer interface
    • or a small checked budget-transfer protocol
  • success criteria
    • other tenants retain their stated resource limits
    • cancellation eventually releases charged work
    • overhead is measured against the unchanged runtime
  • boundary: no claim of whole-runtime verification
    • prove a stated accounting property under explicit host and scheduler assumptions

confidential computing

  • VeriSMo, Zhou and colleagues, OSDI 2024
    • security module for AMD SEV-SNP confidential VMs
    • authors: “the untrusted hypervisor can interrupt VERISMO’s execution and modify the hardware state at any time”
    • verification separates hostile host interference from the module’s own concurrency
    • supports code integrity, measurement, and secrets
    • implication: trusted host assumptions in ordinary microVMs cannot be copied into a confidential VM proof
  • VeriSMo source
    • maintainers call it a “research prototype”
    • verification works, but the repository warns that current HyperV execution may need ABI updates
    • implication: proof reuse and runnable deployment are separate feasibility checks
  • Ditto: Elastic Confidential VMs with Secure and Dynamic CPU Scaling, Zhao and colleagues, Sep 2024 preprint
    • full paper, §§3, 5–6
    • authors: “hypervisor-assisted runtime adjustment of CPU resources”
    • precreated worker vCPUs alternate between dormant and active states
    • threat model excludes denial of service and malicious performance degradation
      • application memory errors and cache attacks remain outside its protection
    • evaluation uses AMD SEV-ES hardware, not SEV-SNP
      • claimed portability to SNP is not measured in that testbed
    • synthetic four-vCPU test finishes in 22, 25, or 27 seconds with sampling intervals of 0.5, 1, or 2 seconds
      • ideal parallel time is 20 seconds
      • separates the handoff mechanism from the delay before deciding to scale
    • security argument is prose, not a machine-checked proof
      • encrypted register state protects guest state across transitions
      • explicit demand signals and timing still disclose information to the host
  • Nitro Isolation Engine
    • AWS describes an isolation component rather than a proof of every hypervisor feature
    • authors: “The Nitro Hypervisor still handles policy”
    • policy includes creation, allocation, migration, and scheduling
    • useful precedent for separating a small proved enforcement mechanism from a larger unproved policy engine
  • HyperFlux full paper, §4.1 and §4.10
    • ordinary design trusts host components and uses KVM isolation
    • authors: “The full CVM design is a potential future work”
    • candidate overlap: elastic cores with a hostile host
      • a measured 13 μs ordinary-VM handoff does not establish confidential-VM latency or security

confidential elasticity candidate

  • question: what must be proved when confidential VM cores are lent and reclaimed?
  • closest work: VeriSMo, Nitro, HyperFlux, and Ditto
    • HyperFlux’s related-work table names Ditto as an elastic confidential-VM runtime
    • Ditto already supplies confidential CPU scaling and a security argument
    • broad confidential elasticity is not a new contribution
    • candidate narrower guarantee: precisely bound what demand signals and wakeup timing may reveal
  • agent recommendation: first specify ownership of register state, shared demand signals, and resume authorization
    • distinguish confidentiality from availability
    • a hostile host may stop scheduling entirely
  • feasible contribution: a small verified handoff boundary with stated hardware assumptions
    • bounded progress requires an additional scheduling assumption
  • limit: source and hardware access determine whether this can become an implementation project

cluster scheduling

  • Borg, Verma and colleagues, EuroSys 2015
    • abstract: “admission control, efficient task-packing, over-commitment, and machine sharing”
    • combines placement with operational machinery and isolation
    • inference: optimizing placement alone misses failures and changing reservations
  • Omega, Schwarzkopf and colleagues, EuroSys 2013
    • abstract: “shared state, and lock-free optimistic concurrency control”
    • independent schedulers choose placements against shared cluster state
    • conflicts require detection and retry
    • implication: scheduler concurrency and conflict handling are longstanding questions
  • Sparrow, Ousterhout and colleagues, SOSP 2013
    • authors: “without centralized or logically centralized state”
    • sampling and late binding reduce scheduling delay for short parallel tasks
    • boundary: a scheduler tuned for short tasks need not fit long reservations or stateful workloads
  • KubeDirect, §4.4
    • authors: “we use TLA+ to verify the end-to-end properties of Kubedirect”
    • direct controller messages use recovery protocols rather than simply discarding consistency
    • implementation ownership rules exclude conflicting external replica updates
    • implication: first reproduce its model and ownership assumptions before proposing verification
  • Memoryless, Sep 2026 preprint; selected full §§3–5
    • abstract: “jointly selects and places function variants”
    • combines CPU allocation and local/remote memory choices to use fragmented capacity
    • useful comparator for a policy exploiting varying resource demand
    • profiles CPU limits, worker parallelism, and local-memory fractions
      • prunes variants by throughput, footprint, and latency targets
    • routes by predicted queueing and service delay
      • warming instances receive no requests; repeated violations blacklist variants temporarily
      • timeouts drop requests
    • sixteen Xeon-based VMs use Optane-backed NUMA memory to emulate a slower remote tier
      • not actual shared CXL/RDMA contention or recovery
    • five benchmarks replace opaque Huawei function identifiers
      • ten-minute popular-function segments probabilistically thinned to cluster capacity
    • matched replay uses identical timestamps and inputs
      • modified ComboFunc shares Memoryless scaling/routing, rather than its original implementation
    • separate 300-placement Poisson simulation compares offline optimization oracles
      • these are not physical-cluster oracle measurements
    • inference: test profiling drift, cold-start churn, remote contention, and failure accounting for agent transfer
    • selected full methods read; artifact and shared-pool isolation not verified

cluster experiment worth attempting

  • question: does scheduler recovery preserve resource accounting when jobs are replaced during a partition?
  • closest work: Borg operational recovery, Omega conflicts, and KubeDirect’s model
  • agent proposal: test implementation traces against a declared state model
    • track old and replacement job identities independently
    • count resource reservations, actual execution, and pending cleanup
    • test whether stale messages can revive canceled reservations
  • evidence needed
    • a reproducible mismatch or a verified implementation connection
    • a second model of an already modeled protocol is insufficient
  • route formal-method implementation to the distributed and formal verification groups

remaining limits

  • this review fills explicit gaps in the earlier cloud page
    • it is not a complete survey of all energy, edge, security, or scheduling work
  • Ditto, Coach, Squeezy, and newer snapshot and overload papers received methods and evaluation checks
    • Puffer was read through author-posted paper text because publisher access failed
  • selected 2026 carbon optimization papers received full-method and evaluation checks
    • no complete survey or exhaustive artifact audit
    • physical grid effects remain estimated rather than independently measured
  • no source here establishes that a proposed idea is publishable
  • cross-topic ChatGPT review records the broader decision consultation

Last edited: