Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

data that spans regions (authored by agents unless marked 🧑)

what this file is

  • a pass on data kept in several distant data centers: production databases, clocks, the latency price of strong guarantees, residency laws, edge data, losing a whole region, and published outage reports
  • it adds to three sibling files and does not repeat them
  • status: cut short on 7 Oct 2026 UTC by the Claude usage limit
    • web search stopped working after 9 queries; dblp, arXiv, and OpenAlex search all returned HTTP 429
    • so the 2025 to 2026 research-paper coverage is thin; the vendor documents and outage reports are solid
    • “what is not covered” at the end lists the holes
  • how to read the quotes
    • every quoted phrase was string-matched by me against the page or PDF I downloaded that day, unless marked “leftover, not rechecked”
    • “leftover” means an earlier agent that was cut off recorded it and I did not re-open the source
    • “my reading” and “my inference” mark what is mine
  • ChatGPT Extra High, original worker: no opinion was obtained; the shared browser was not signed in and I ran out of budget before writing the prompt file

takeaway, my opinion

  • production systems have converged on one shape for strong multi-region data
    • two regions hold full copies, a third only votes (a “witness”), a write waits for two of three
    • Spanner dual-region, DynamoDB multi-region strong consistency, and Aurora DSQL all document this shape
    • what differs is what each one gives up: DynamoDB drops transactions, DSQL gives snapshot isolation, Cosmos DB refuses strong consistency with several write regions
  • the published outage reports say replicas were rarely what failed
    • what failed was something every region shared: a DNS record, a global policy table, a config file, a deletion tool, a third-party store
    • so “survive the loss of a region” is mostly a question about hidden shared dependencies, and I found little research that measures them
  • residency is promised per row but leaks per key
    • CockroachDB and Spanner both document that index keys and split points leave the pinned region
    • I found no paper that measures this from the outside; that is a measurement paper we could write with tools we already know
  • what I would do first, in order
    1. a leak audit of pinned multi-region databases (research idea 1)
    2. a Verus model of the two-copies-plus-witness commit and what survives bad clocks (idea 2)
    3. a catalogue of outage reports tagged by which shared dependency crossed regions (idea 3)
  • these are hypotheses; I ran no experiment, and “I found none” is limited by the broken search

the problem in plain words

  • light takes time; a round trip between US coasts is tens of milliseconds, across an ocean more
  • to never lose a confirmed write when a region dies, some other region must have the write before you confirm it
    • so a safe write costs at least one round trip to the nearest other region
  • to let every region read the newest data locally, either writers wait longer or readers must know how fresh their copy is
    • synchronized clocks are the usual way to know
  • laws and contracts add a second constraint: some bytes may not leave a country
  • every system below is a different way to pay these costs

what production systems promise, in their own words

  • Azure Cosmos DB, consistency levels doc
    • the price of strong, stated as a formula: “the write latency is equal to two times round-trip time (RTT) between any of the two farthest regions, plus 10 milliseconds at the 99th percentile”
    • a distance cap: “Strong consistency for accounts with regions spanning more than 5,000 miles (8,000 kilometers) is blocked by default because of high write latency.”
    • a stated impossibility: accounts with multiple write regions “can’t use strong consistency because a distributed system can’t provide a recovery point objective (RPO) of zero and a recovery time objective (RTO) of zero”
      • RPO is how much confirmed data you may lose; RTO is how long you may be down
  • DynamoDB global tables, how it works
    • two modes: eventual (MREC) and strong (MRSC, launched June 2025)
    • eventual mode resolves conflicts by “last writer wins”, and with transactions “you might observe partially completed transactions” in another region while changes replicate
    • strong mode: changes are “synchronously replicated to at least one other Region before the write operation returns a successful response”
    • strong mode’s limits: “A MRSC global table must be deployed in exactly three Regions.” and such tables “do not support transaction operations”
    • a concurrent write to the same item in another region fails with ReplicatedWriteConflictException
    • my reading: DynamoDB bought zero data loss by rejecting conflicts and dropping transactions
  • Aurora DSQL, Brooker et al., arXiv July 2026; commit protocol is in distributed transactions
    • layout tested: two full regions “with a witness in us-west-2”
    • cost model: for read-write work the commit takes “approximately 2x the round-trip time between the nearest pair of regions”
    • reads stay local: latencies “significantly lower than one network round trip” between the two main regions
    • what bad clocks cost: “In other words, DSQL becomes merely snapshot isolated, rather than strong snapshot isolated.”
    • Brooker’s blog, Dec 2024: “DSQL offers active-active multi-writer capabilities in multiple availability zones (AZs) in a single region, or across multiple regions.”
  • Spanner
    • replication doc lists read-write, “read-only replicas, and witness replicas”
    • leftover, not rechecked, same doc: witness replicas “don’t maintain a full copy of data” and “don’t serve reads.”
    • leftover, not rechecked, OSDI 2012 paper: “If the uncertainty is large, Spanner slows down to wait out that uncertainty.”
  • CockroachDB
    • multi-region demo, VLDB 2022: a table is REGIONAL BY TABLE, REGIONAL BY ROW, or GLOBAL, plus a survival goal (zone or region)
    • global tables blog: writes “operate at future MVCC timestamps”
      • my reading: writers wait so that every region can read locally and still see the newest data; good only for data read far more than written
    • leftover, not rechecked, SIGMOD 2022 paper blog: the protocol “reduces tail latency by over 10x compared to prior approaches”
  • Aurora Global Database, doc
    • one writer region: “Only the primary cluster performs write operations.”
    • copies lag: replication “with latency typically under a second”
    • my inference: failover can lose the last moments of writes; this is the GitHub 2018 failure shape below
  • my inference from putting these side by side
    • the vendors state commit cost in different units (2x farthest pair plus 10 ms; 2x nearest pair; a future-timestamp wait), so nobody can compare them from documents
    • Gaia (below) starts that comparison for open systems; the managed services are still unmeasured as far as I found

clocks: what synchronized time buys

  • the idea: if every machine knows the time within a known error, a reader can tell whether its local copy is fresh enough without asking anyone
  • how small the error is now
    • Sundial, Li et al., OSDI 2020: “Sundial can achieve ~100ns time-uncertainty bound under different types of failures, which is more than two orders of magnitude lower than the state-of-the-art solutions.”
      • inside one data center, on a testbed of more than 500 machines
    • Graham, Najafi and Wei, NSDI 2022: “Graham reduces the clock drift of a commodity server by up to 2000×, reducing the maximum assumed drift in most situations from 200ppm to 100ppb.”
      • my reading: a machine that loses its time server stays trustworthy for longer
    • AWS, Nov 2023 announcement: “The Amazon Time Sync Service now gives you a way to synchronize time within microseconds of UTC on Amazon EC2 instances.”
    • AWS ClockBound, open source: “The window of uncertainty (the Clock Error Bound) is defined by two timestamps (earliest, latest) within which true time exists.”
      • my reading: this is Spanner’s interval clock offered to any tenant; ten years ago only Google had it
  • what happens when the bound is wrong
    • CockroachDB transaction layer doc: a node that sees itself too far off “crashes immediately”
    • same doc: skew past the bound can cause “violations of single-key linearizability between causally dependent transactions”
    • DSQL: drops to plain snapshot isolation (quote above)
    • my inference: each system states a different loss in prose; none of these statements is machine-checked as far as I found
  • leftovers on clocks, not rechecked: Huygens (NSDI 2018), Firefly (SIGCOMM 2025), Nezha (VLDB 2023), YugabyteDB’s hybrid clocks
  • K2 and Tiga, which use clocks to order transactions, are in the sibling files

measuring the latency price

  • Gaia, Mraz et al., arXiv May 2026, abstract
    • the complaint: “Existing benchmarks assume stable network conditions; they lack explicit settings for data and client locality, and they largely ignore data transfer costs across regions.”
    • the findings: “i) most systems are sensitive to network instabilities, ii) network costs dominate cloud deployment expenses iii) multi-region fault-tolerance mechanisms incur measurable critical-path overhead that is often overlooked in prior evaluations.”
    • I read only the abstract
  • my reading: finding ii is the under-studied one; most papers in the sibling files report latency and throughput, not the bill for bytes crossing regions

where to put data and copies

  • row-level movement between replication groups (PolyBase) and geography-aware locking (Bonspiel) are in transactions and regions
  • LEGOStore, Zare et al., VLDB 2022, arXiv, abstract
    • picks, per object, between full copies and erasure coding (split the object into pieces so any few rebuild it) and picks the data centers, “to minimize overall costs”
    • “We observe cost savings ranging from moderate (5-20%) to significant (60%) over baselines representing the state of the art while meeting tail latency SLOs.”
    • moving a key’s placement takes “3 to 4 inter-DC RTTs”
  • GeoLayer, arXiv Sept 2025: placement of graph data across data centers; found by search, abstract not read closely
  • not reached: Akkio, Shard Manager, SkyStore, Macaron, Skyplane, and other cost-driven placement work

laws that keep data in a country

  • what the law says
    • GDPR does not say “keep data in the EU”; Article 44 says a transfer to a third country “shall take place only if 
 the conditions laid down in this Chapter are complied with by the controller and processor”
      • my reading: keeping data inside is the way to need no transfer basis
    • US CLOUD Act, 18 U.S.C. 2713: a provider must disclose data in its “possession, custody, or control, regardless of whether such communication, record, or other information is located within or outside of the United States.”
      • my inference: this is why “EU region of a US company” does not settle the matter for EU buyers
    • leftovers, not rechecked: Schrems II press release, China PIPL thresholds, India DPDP, Russia localization (secondary sources only)
  • what vendors sell
    • AWS European Sovereign Cloud, launch post: “The first AWS Region in the AWS European Sovereign Cloud is located in the state of Brandenburg, Germany, and is generally available today.” and “This Region operates independently from existing AWS Regions.”
    • Microsoft EU Data Boundary: “These commitments are subject to limited circumstances where Customer Data, personal data, and Professional Services Data will continue to be transferred outside the EU Data Boundary.”
  • what databases admit leaks
    • CockroachDB data domiciling doc: indexed column data may land in system ranges, and “This synchronization does not respect any multi-region settings”
      • also on logs: “there is some cross-region leakage”
    • Spanner data residency doc: some key values “are used as split boundaries, which might be stored in the default placement”
    • my reading: both pin rows well and both let key values travel; an email address used as a primary key is exactly such a value
  • research on compliance
    • K9db, Albab et al., OSDI 2023: “K9db is a new, MySQL-compatible database that complies with privacy laws by construction.”
      • it handles who owns data and deletion, not where data sits (my inference from the abstract)
    • GDPRuler, arXiv June 2026: GDPR enforcement for key-value stores; abstract says existing approaches “overlook the integrity of compliance mechanisms themselves”
    • leftovers, not rechecked: GDPRbench (VLDB 2020), compliant geo-distributed query processing (SIGMOD 2021)
  • measuring where data really goes
    • Gamero-Garrido et al., PoPETs 2025, arXiv: “BGP, DNS and other Internet protocols were not designed to enforce jurisdictional constraints”
      • “on average, 2.3% of servers serving users in each EU country are located in non-adequate destination countries”
      • my inference: this measures where serving machines are, not where stored copies, backups, and logs end up
    • GeoFINDR, arXiv 2025: locate a cloud machine by network delay even with “a Cloud Service Provider (CSP) lying about the VM’s location”; accuracy “can be as high as 22.1km”

data at the edge and on devices

  • CRDTs and local-first software are in consistency guarantees
  • Cloudflare Durable Objects with SQLite, Varda, blog Sept 2024
    • the design: “Your application code runs exactly where the data is stored.”
    • one object has one writer, so no cross-region conflict; the cost is that a far user reaches the object over the network (my reading)
    • the write trick is called “Output Gates”: the code continues before the write is confirmed, and replies to the outside are held until it is
  • the same storage shares a hidden dependency, see the June 2025 outage below
  • Antipode (SOSP 2023): keeps “post before notification” order across several stores in several regions; known to me from the search summary only, PDF not opened
  • not reached: D1 read replicas, Turso, LiteFS, EdgeKV, and academic edge stores from 2022 on

losing a whole region

  • Maelstrom, Veeraraghavan et al., OSDI 2018: “Maelstrom provides a traffic management framework with modular, reusable primitives that can be composed to safely and efficiently drain the traffic of interdependent services from one or more failing datacenters to the healthy ones.”
  • Meta, The Evolution of Disaster Recovery at Meta, 2023
    • goal: “handle the loss of a single region without impacting site availability”
    • spare capacity kept for this is the “DR buffer”
    • practice: “DR Storms are our Disaster Readiness Exercises, during which we isolate a production region to validate the end-to-end (E2E) readiness of the DR buffer and service placement.”
    • a harder drill: “A Power Storm is a mammoth DR exercise where a typical production region is brought to a complete stop”
    • an admission about cutting a region off versus powering it down: “we hadn’t exercised our readiness relating to this particular mode of failure, so our recovery here was unknown.”
  • Metastable Failures in the Wild, Huang et al., OSDI 2022: “at least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failures”
    • a metastable failure keeps a system stuck in overload after the trigger is gone
    • my inference: failing over a region shifts load at once, which is a classic trigger

published outage reports that involved storage or databases

  • a shared name record took a region’s database away
    • AWS, DynamoDB disruption in us-east-1, 19 to 20 Oct 2025: “a latent race condition in the DynamoDB DNS management system that resulted in an incorrect empty DNS record for the service’s regional endpoint”
    • the report also describes EC2’s internal manager entering “congestive collapse” because it depends on DynamoDB
  • a globally copied row crashed every region
    • Google Cloud, 12 June 2025: a policy change went into “regional Spanner tables” and was “replicated globally”; then “the null pointer caused the binary to crash”
    • the code “did not have appropriate error handling nor was it feature flag protected”
    • my reading: replication worked perfectly and delivered the bad row everywhere within seconds
  • that outage took down someone else’s edge store
    • Cloudflare, 12 June 2025: “The cause of this outage was due to a failure in the underlying storage infrastructure used by our Workers KV service”
    • “Part of this infrastructure is backed by a third-party cloud provider, which experienced an outage today”
    • reach: “D1 databases share the same underlying storage infrastructure as Workers KV and Durable Objects.”
  • a database permission change broke a global file
    • Cloudflare, 18 Nov 2025: “a change to one of our database systems’ permissions which caused the database to output multiple entries into a” feature file used by bot management
  • single-site stores behind a multi-site service
    • Cloudflare, 2 Nov 2023: Kafka and ClickHouse “were only available in PDX-04 but had services that depended on them that were running in the high availability cluster”
  • failover with lagging copies split the data
    • GitHub, 21 Oct 2018: “Connectivity between these locations was restored in 43 seconds, but this brief outage triggered a chain of events that led to 24 hours and 11 minutes of service degradation.”
    • both coasts held writes the other lacked, so “we were unable to fail the primary back over to the US East Coast data center safely”
  • a deletion that replicas could not stop
    • Google Cloud and UniSuper, May 2024: “After the end of the system-assigned 1 year period, the customer’s GCVE Private Cloud was deleted.”
    • separate backups “were instrumental in aiding the rapid restoration”
  • an operator command and a cold restart
    • AWS S3, 28 Feb 2017: “one of the inputs to the command was entered incorrectly and a larger set of servers was removed than intended.”
    • “we have not completely restarted the index subsystem or the placement subsystem in our larger regions for many years.”
  • physical loss
    • Google Cloud Paris, 25 Apr 2023: “a cooling system water pipe leak occurred in one of the data centers in the europe-west9 region.” It “led to a fire.”
  • 2026, network inside or around one region
    • Google Cloud us-west1, 20 Aug 2026: “The disruption originated during scheduled fiber optic maintenance, which unexpectedly compromised network capacity between data centers within the us-west1 region.”
      • Cloud Storage, Cloud SQL, Bigtable, and AlloyDB are on the affected list
    • Azure West US, 23 July 2026, status history: “routes were withdrawn from multiple devices simultaneously, disrupting connectivity between the datacenter and the WAN”
  • leftovers, not rechecked: AWS Kinesis Nov 2020, AWS Dec 2021, South Korea government data center fire Sept 2025 (news snippet only)
  • my reading of the set
    • 6 of the 11 reports I checked are about something shared across sites, not about a lost copy
    • 2 are about operations nobody had practiced (S3 restart, Meta’s power-down admission)
    • only GitHub 2018 is a pure replication story, and its lesson is that failover with lagging copies makes two diverging databases

research we could do

  • idea 1: leak audit of pinned multi-region databases
    • question: with every table pinned to one region, which bytes still cross the border?
    • why open: the vendors’ documents admit key and log leaks (quotes above); I found no outside measurement, though my search was cut short
    • method: run CockroachDB and YugabyteDB with strict pinning in emulated regions, plant marker values in keys, indexes, and rows, capture all traffic between regions and scan disks, logs, and backups for the markers
    • fits: web measurement skills; result is a table of leak channels per system and version
    • stop if: a 2022 to 2026 paper already did this; check PoPETs, USENIX Security, and VLDB first
  • idea 2: a Verus model of two copies plus a witness
    • question: prove no confirmed write is lost with one region down, and state exactly which guarantee survives when the clock bound is broken
    • why open: three vendors ship this shape and each describes the bad-clock case in prose; DSQL says it used TLA+ and P model checking (transactions), which is not a proof of an implementation
    • first step: an executable Rust core with a ClockBound-style interval clock as a trusted interface, then weaken the clock assumption and see which proof breaks
    • stop if: the consensus slice already has a verified witness or flexible-quorum protocol that covers it
  • idea 3: outage reports as data
    • question: what share of large storage outages crossed regions through a shared dependency, and which kind?
    • method: collect vendor reports 2017 to 2026, tag trigger, shared component, whether data was lost, whether failover helped
    • 11 reports are already tagged above; the metastable-failures paper is the model for this kind of study
    • an LLM can do first-pass tagging, with a human check on a sample
  • idea 4: measure managed multi-region commit cost on the same region triples
    • Gaia does this for open systems; DSQL, DynamoDB strong mode, Spanner, and Cosmos DB publish formulas in different units
    • cost is cloud spend, not people; weakest on novelty since Gaia may extend to it
  • idea 5: test failover with lagging copies for split data
    • reproduce the GitHub 2018 shape on Aurora Global, MySQL, and Postgres setups and check histories with the checkers in consistency guarantees
    • lower priority; Jepsen-style work is close

what is not covered

  • research papers from 2025 and 2026 beyond the few named: I could not search SOSP, OSDI, NSDI, EuroSys, ATC, SIGCOMM, SIGMOD, VLDB, or CIDR programs before the limit
  • YugabyteDB, TiDB, OceanBase, PolarDB-X, FoundationDB, MongoDB zones, and Meta’s stores (TAO, ZippyDB, MySQL Raft)
  • cost-driven placement across regions and clouds (SkyStore, Macaron, Skyplane, Akkio)
  • edge and device stores beyond Durable Objects
  • wide-area clock sync and Google’s current TrueTime numbers
  • law beyond GDPR Article 44 and the CLOUD Act; the leftover notes on Schrems II, China, India, and Russia rest on secondary sources
  • Azure post-incident reviews beyond one line; OVH 2021, Atlassian 2022, Roblox 2021, Facebook 2021
  • Jepsen reports on multi-region behavior
  • no ChatGPT opinion; no second reviewer read this file
  • sources: about 40 pages or PDFs downloaded and string-checked by me; about 15 more appear only as marked leftovers
  • leftover notes for a later pass are in the scratch folder mr/ (sub1_production.md, sub2_clocks.md, sub4_residency.md, sub6_outages.md)

regional placement and copy-cost follow-up

closest prior work now read beyond abstracts

  • SkyStore, Liu et al., PVLDB 2025 paper, pp 2085–2088, 2094–2095
    • mechanism: write locally, copy on a remote read, choose eviction times from a byte-weighted histogram of gaps between reads
    • exact quote, §3.2.2: “collected per region per workload”
    • limit on the cost proof, §3.1: “For simplicity, we are ignoring the associated operation costs”
      • the two-times-optimal bound concerns the simplified two-region storage-and-egress model
      • it does not establish a bound for every multi-cloud deployment
    • §3.2.3 explicitly proposes grouping objects with similar access patterns
      • a proposal to learn groups for cheaper placement overlaps stated future work
    • §6.7.3 measures metadata overhead and discusses weaker metadata consistency as future work
      • my inference: copy-cost work must include metadata traffic and failure recovery, rather than counting only object bytes
    • read depth: placement and eviction mechanism, proof assumptions, metadata-overhead evaluation
      • not a full correctness or implementation audit
  • Macaron, Park et al., SOSP 2024 author-hosted paper, §§4–6
    • mechanism: a small DRAM cache above a cheap object-store cache
      • sampled trace simulations estimate missed bytes, miss rate, and average latency
      • the controller chooses object-store capacity by total dollar cost and DRAM capacity by latency
    • exact quote, §5: “set to 15 minutes by default”
      • this quote concerns the controller reconfiguration interval
      • default Macaron changes capacity using LRU eviction
      • section 5.1 describes Macaron-TTL as a variant that changes expiry times
    • exact quote, §4.3: “assumes data is immutable”
      • mutable applications must manage invalidation themselves
    • §5.1 excludes write-through transfer costs from the capacity formula because they do not change with cache capacity
      • my inference: that exclusion is reasonable for this decision, but unsuitable for comparing replication protocols with different write traffic
    • read depth: architecture, consistency assumptions, capacity objective, sampling, object packing and cache priming
      • not all evaluation plots or appendices
  • Skyplane, Jain et al., 2022 primary preprint, §§3–5
    • mechanism: choose relay regions, split traffic across paths, and choose VM counts from measured throughput and provider prices
    • exact quote, §4: “an application-specified throughput constraint”
      • cost minimization already supports a required transfer speed
    • exact quote, §4.1: “egress bandwidth is charged for each hop”
      • a fast detour can increase the bill
    • the planner uses a mixed-integer linear program
      • its optimum is relative to measured capacities and the supplied price model
    • my inference: a proposal for cheaper bulk copies must compare with transfer-path optimization as well as object placement
    • read depth: measured throughput grid, planner variables and constraints, routing and parallel connections
      • preprint read; final publication differences not checked
  • Akkio, Annamalai et al., OSDI 2018 paper, §§3–4.6
    • mechanism: move a small group of related data among existing replica groups
      • recent access counts choose location; resource use breaks ties
      • a limit on movement frequency prevents repeated moves back and forth
    • exact quote, §4.6.2: “once every 6 hours”
      • the example already limits migration frequency
    • exact quote, listing 2: “Writes are blocked during the migration”
      • the ZippyDB path uses permissions and transactions
      • the Cassandra path uses double writes, timestamps, and waits for location-cache expiry
    • §3 supports constraints on replica location and consistency
      • a placement proposal with allowed regions is not novel merely because it has geographic constraints
    • read depth: placement policy, metadata lookup, replication configurations, both migration paths and their partial-write race
      • no full recovery-code audit

what these papers change in our proposals

  • my inference: minimizing copy cost, learning placement from past accesses, adjusting eviction times, and limiting repeated movement are established mechanisms
    • none alone is a credible novelty claim
  • a narrower question worth checking: the total cost of changing placement while preserving a stated consistency guarantee
    • include copies, double writes, metadata, cache expiry waits, retries, and blocked-write time
    • compare Akkio’s migration paths, SkyStore’s replica management, Macaron’s immutable-data cache, and Skyplane’s transfer planner
    • these systems make different assumptions
      • any comparison must state the workload and guarantee before comparing the bill
  • my inference: residency restrictions could constrain the feasible copies and relay paths
    • Akkio already constrains replica locations
    • Skyplane introduces intermediate regions that an endpoint-only residency check could overlook
    • this reading does not establish a new legal or technical gap
  • stop condition: do not develop this into an experiment until checking SPANStore, SkyPIE, LEGOStore, and newer joint placement-and-transfer work

limits of this sweep

  • direct primary-source access worked on 8 Oct 2026
    • OpenAlex resolved titles; publisher or author sites supplied all four PDFs
    • ACM returned HTTP 403 for Macaron; the CMU author-hosted PDF worked
  • this closes the four named unread-paper holes, not the entire regional citation sweep
  • no experiment, independent novelty confirmation, or ChatGPT Extra High consultation was completed here

LEGOStore mechanism check

  • Zare et al., versioned primary text, sections 3.2–3.3
    • evidence: “assuming perfect knowledge of workload and system properties”
      • context: the placement optimizer’s inputs
    • evidence: “reconfigurations are applied sequentially by the reconfiguration controller”
      • context: the migration protocol
    • the objective already includes read/write network, storage, and VM costs
      • network metadata is included
      • stored metadata is excluded as negligible
    • migration copies a tagged value to the new configuration
      • some concurrent operations pause or restart
    • inference: a new migration-cost study must compare with this explicit protocol and cost objective
      • imperfect predictions and controller behavior are concrete dimensions to inspect
      • they are not established novel gaps
    • read depth: optimizer inputs/objective and migration mechanism
      • appendix proof and recovery implementation not audited
  • SPANStore and SkyPIE primary bodies remain unread in this follow-up
    • the attempted SPANStore author URL returned HTTP 404
    • title search did not resolve a SkyPIE primary copy
    • bibliographic discovery alone supplies no mechanism evidence

regional placement candidate remains deferred

recommendation

  • retire regional placement from the active research shortlist pending direct mechanism checks of SPANStore and SkyPIE
    • retain the four-paper survey as literature notes
    • the narrower total-cost question remains an unconfirmed possibility
    • do not present it as a research gap

why the remaining comparison matters

  • SkyStore explicitly compares with SPANStore
    • SkyStore primary paper, introduction: “SPANStore does not account for replication costs”
    • this is SkyStore’s characterization of another paper
      • I have not verified SPANStore’s own objective, cost exclusions, consistency constraints, or adaptation mechanism
    • the total-cost question could already be addressed by SPANStore or by the difference SkyStore addresses
  • SkyPIE has an accessible author-maintained implementation
    • primary repository README: “how to use the oracle API or the ILP baseline”
    • the README lists a real-trace experiment and precomputed oracles for different deployments and service objectives
    • this establishes available comparison machinery
      • it does not establish the paper’s migration assumptions or which transition costs its objective includes
  • my inference: these are material closest comparisons
    • without them, enough uncertainty remains to defer the candidate
    • no claim that the topic is exhausted or uninteresting follows

access attempts on 8 Oct 2026

  • SPANStore, primary DOI
    • OpenAlex metadata resolved title and authors
    • ACM PDF with download parameter returned HTTP 403
    • UCR author-paper paths returned HTTP 404
    • Michigan author-site alternatives failed to connect
    • Columbia author-site alternatives returned HTTP 404
    • arXiv title search returned no matching paper
  • SkyPIE, primary DOI
    • OpenAlex metadata resolved title and authors
    • ACM PDF with download parameter returned HTTP 403
    • Berkeley author-paper alternatives returned HTTP 404
    • Cornell author-site alternative returned HTTP 404
    • guessed project sites failed or returned HTTP 404
    • arXiv search returned unrelated astronomy results
    • GitHub primary repositories were accessible
      • recursively listed trees contained no PDF
  • shared web search returned HTTP 404
    • response: “Cannot POST /alpha/search”
  • conclusion: paper copies were not recovered through these alternatives
    • the blocker is access to these two primary papers in this session
    • it does not imply that all literature access is blocked

read depth and limits

  • SPANStore: metadata and SkyStore’s related-work characterization only
  • SkyPIE: metadata and primary implementation README only
  • objectives, constraints, adaptation and migration assumptions remain unverified against both papers
  • this retrieval follow-up ran no experiments
    • the coordinator’s completed Extra High consultation is recorded in the storage index

Last edited: