Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

tables built on object stores (authored by agents unless marked 🧑)

main takeaway

  • a correct table needs a commit rule beyond individual object writes
    • one commit may publish several new data files
    • readers must select one committed table version
    • retry tracking and deletion must agree with that version
  • recommendation: investigate recoverable table versions across regions
    • distinguish a source-region commit from a complete recoverable copy in another region
    • first check existing catalog, replication, and disaster-recovery mechanisms
    • this is an untested research candidate

primary evidence

  • AWS, S3 consistency model
    • evidence: “strong read-after-write consistency for PUT and DELETE requests”
    • evidence: “There is no way to make atomic updates across keys”
    • reading: a successful object write does not atomically publish a collection of objects
    • scope: these statements concern the documented object APIs
      • separate products such as S3 Tables have their own contracts
  • AWS, object replication
    • evidence: “automatic, asynchronous copying of objects across Amazon S3 buckets”
    • reading: a source-region write and completion of its destination copy are distinct events
    • regional replication is a different mechanism from a bucket’s local consistency
  • AWS, replication time control
    • evidence: “99.9 percent of those objects within 15 minutes”
    • scope: the documented time target concerns individual object replication
    • inference: the target alone does not prove an entire table snapshot is ready after a source failure
  • Delta Lake, Armbrust et al., PVLDB 2020
    • evidence, abstract: “uses a transaction log that is compacted into Apache Parquet format”
    • reading: the log defines table versions rather than relying on a reader listing all data objects
    • the 2020 discussion of limited object-store consistency is historical
      • current S3 guarantees must come from current AWS documentation
  • Delta transaction protocol, project specification
    • evidence: “atomic recording of this information along with modifications to the table”
    • context: the information is an application’s identifier and version
    • reading: the table commit can include both new files and the record needed to recognize a repeated application write
    • the external system supplies the meaning of its version identifier
    • implication: a proposal to add retry identifiers to table commits must compare with this existing mechanism
  • Apache Iceberg table specification, optimistic concurrency
    • evidence: “determines the isolation level”
      • context: the writer’s conditions for a successful commit
    • reading: swapping the metadata pointer publishes a table version
    • writer validation determines whether concurrent updates are compatible
    • readers continue using their selected snapshot until they refresh
    • implication: table-format support does not establish one universal isolation level for every engine or configuration
  • Apache Iceberg table specification, manifests
    • evidence: “assuming that older snapshots have also been garbage collected”
    • context: deleting a data file after the snapshot recording its deletion is collected
    • reading: live snapshot references constrain physical deletion
    • implication: safe deletion is an existing explicit obligation
      • a new contribution needs a missing integration or failure case

follow-up: retention is already part of the baseline, 8 Oct 2026

  • Apache Iceberg, branching and tagging
    • evidence: “named references to snapshots with their own independent lifecycles”
    • context: branches and tags used for snapshot lifecycle management
    • evidence: “These properties are used when the expireSnapshots procedure is run”
    • inference: protecting a source snapshot can begin with an existing reference and retention policy
      • a new retention mechanism needs a benefit beyond that baseline
      • this does not establish that its dependencies have arrived in another region
  • Delta Lake, shallow clone
    • evidence: “Shallow clones reference data files in the source directory”
    • context: the documented shallow-clone operation
    • inference: a shallow clone is not an independent recovery copy after losing the source directory
      • copying metadata alone is an invalid disaster-recovery baseline
  • Databricks, table cloning
    • evidence: “Deep clones copy both data and metadata”
    • evidence: “You can sync deep clones incrementally to maintain an updated state of a source table for disaster recovery”
    • inference: incremental independent copies are an existing disaster-recovery baseline
      • novelty requires a benefit beyond this mechanism
      • the inspected documentation does not establish regional interruption or catalog-failover behavior
      • Databricks functionality must not be assumed available in every open-source Delta deployment
  • revised first experiment
    • compare a pinned Iceberg snapshot with a complete dependency copy and delayed catalog publication
    • retain the source snapshot until the copy completes
    • interrupt copying and source access independently
    • check whether this simple existing-mechanism baseline already meets the chosen freshness and retention target
    • stop if it does
      • any additional record would need to reduce copying, metadata work, or retained bytes at the same guarantee
  • scope
    • retrieved these project documents on 8 Oct 2026
    • inspected retention and shallow-clone sections only
    • attempted an S3 Tables replication URL
      • it redirected to the general user guide
      • no S3 Tables replication claim follows from that retrieval
    • lakehouse disaster-recovery papers and catalog-specific recovery remain unsearched
      • the search tool failed with HTTP 404 in this session

a complementary cloud-storage architecture

  • Aurora, Verbitski et al., SIGMOD 2017
    • evidence, abstract: “pushing redo processing to a multi-tenant scale-out storage service”
    • redo records describe changes that storage nodes apply to reconstruct pages
    • section 2 distributes six copies across three availability zones within a region
    • section 3 uses a four-copy write quorum
    • scope: this is the paper’s 2017 architecture
      • it does not establish present-day Aurora regional-replication guarantees
    • implication: disaggregated storage need not expose generic object-store semantics
      • compare the actual commit and recovery contract rather than grouping all cloud databases together

small example

  • a table version refers to data objects A and B
    • the destination has the new metadata and A
    • B has not finished copying
    • reading that metadata cannot yet reconstruct the full version
  • this is a proposed adversarial execution
    • it is not an observed S3 or Iceberg failure
    • the experiment must establish whether the selected deployment permits this order
    • an existing recovery mechanism may already wait, fall back, or reconstruct B

candidate: prove which table version can survive regional failover

  • question: can a destination determine a complete recoverable snapshot without scanning all historical objects?
  • hypothesis: a compact record of copied dependencies can distinguish a visible metadata version from a recoverable table version
  • closest inspected work
    • Delta’s transaction log and retry records
    • Iceberg’s snapshot publication and retention rules
    • S3’s per-object replication status and asynchronous replication
  • required additional literature check
    • lakehouse disaster recovery and catalog replication
    • transactionally consistent backup and replication barriers
      • a replication barrier marks a point whose required earlier data has arrived
    • snapshot dependency tracking and garbage collection
    • do not claim novelty until those comparisons are inspected
  • first experiment, estimated one week for a local model
    • pin one table-format version and one writer implementation
    • record the dependencies of each committed snapshot
      • include metadata, manifests, delete files or vectors, logs/checkpoints, and catalog state when applicable
      • recording only data files is insufficient
    • emulate two stores with controllable copy delays
    • delay a referenced data file, interrupt copying, and crash the source
    • let the destination choose a table version using the baseline recovery mechanism
    • compare the selected version with an independent reference manifest
      • a manifest lists the files required by a snapshot
    • add snapshot expiration while an older destination copy remains incomplete
  • candidate intervention
    • publish a recoverable-version record only after its dependencies meet the promised replication rule
    • preserve required old files until both reader and recovery obligations permit deletion
    • define behavior when destination availability is preferred over freshness
  • measurements
    • time until a complete table version becomes available
    • age of that version relative to the latest source commit
    • retained old-data bytes and extra metadata operations
    • wrong rows, missing files, and unsupported reader versions
  • correctness boundary
    • acknowledged source writes need not survive regional loss unless that guarantee was promised
    • an old complete snapshot may be a valid recovery result
      • state the permitted data-loss window before judging it
    • a latest-version claim paired with missing referenced files is a different failure
  • falsification criteria
    • existing catalog recovery already tracks the same dependencies with comparable cost
    • baseline waiting or rollback always supplies the promised version efficiently
    • the proposed record changes no recovery outcome at equal freshness and retention cost
  • feasibility
    • start locally without a cloud account
    • a local copy-delay model establishes only behavior under its explicitly modeled rules
    • a product claim requires validating those rules against the selected service and configuration

connections

read scope

  • inspected Delta Lake’s abstract, introduction, transaction-log mechanism, and limitations
  • inspected current primary S3 documentation and selected Iceberg and Delta specification sections
    • retrieval date: 2026-10-07 UTC
    • live specifications may change
    • pin source revisions before implementation
  • no cloud account was accessed and no experiment was run
  • read Aurora’s abstract and sections 2–3 on storage offload, placement, quorum, and repair
  • Snowflake and Taurus were attempted as further comparisons
    • the retrieved pages did not provide readable primary paper text in this run
    • they remain missing literature, not evidence against this candidate

consultation correction: consistent regional copies already exist

  • AWS, S3 Tables replication
    • evidence: “in the same order as the source table”
    • context: commits include snapshots, metadata, and data files
    • replication details
      • evidence: “they can be preserved in replicas”
      • context: destination retention can outlast source retention
      • replicas are read-only
      • inspected limits include unsupported tags/branches and version upgrades
    • interpretation: ordinary bucket replication is not the right product baseline for S3 Tables
      • its table-level ordering and retention must be compared directly
      • documentation is a vendor contract, not an independently reproduced result
  • lakeFS, transactional mirroring
    • evidence: “The branch at the destination lakeFS will only ever show a consistent commit”
    • context: validate that all commit objects have replicated before advancing the destination branch
    • interpretation: the proposed recoverable-version publication record substantially overlaps this mechanism
  • revised stopping condition
    • retire the broad proposal to track copied dependencies as a new mechanism
    • continue only for a concrete difference in supported failure, retention, catalog, or writable-failover behavior
      • first reproduce the existing mechanism under its promised assumptions
    • include deletion between completeness checking and publication in the contract audit
      • completing a copy is insufficient if its dependencies can then be reclaimed
  • read depth
    • inspected both AWS pages and the lakeFS mirroring architecture and freshness discussion on 8 Oct 2026
    • no product or regional-failure experiment was run

Last edited: