A visual story about cloud storage

How one S3 upload earns eleven nines.

You PUT one file. About a second later, S3 answers 200 OK — having accepted a design target of 99.999999999% yearly durability for it. Here is what happened to your bytes in that second.

SEP 2026~9 MIN6 SCENESILLUSTRATIVE SHARDS · REAL MECHANISMS

One upload, one second, eleven nines

You PUT one 42 MB video — an illustrative object — and S3 answers 200 OK about a second later. In that second, AWS accepted a design target of eleven nines of yearly durability for it: 99.999999999%.

In AWS's own framing: store ten million objects and you would expect to lose one every ten thousand years. That is not a disk you can buy. It is a machine, and this article follows your bytes through it.

ILLUSTRATIVE OBJECT AND 6+3 SHARDS — S3'S REAL PARAMETERS ARE NOT PUBLIC · DURABILITY: AWS DESIGN TARGET, MODELED NOT MEASURED.

Act 01 · The modelTwo systems answer your request: a card and a stream

01 · flat keysA bucket holds objects named by flat string keys. videos/aurora.mp4 is one key — not a path through real folders.

02 · the splitIt is two systems. The card keeps the record — key, size, version, where the shards live. The stream is the encoded bytes, on disks.

03 · the geographyThe card must be always right. The bytes must be never lost. Below the stream wait at least three Availability Zones — separate buildings, power, and networks.

What to watch The object splits down the middle: a card that must be always right, and bytes that must be never lost.

Act 02 · The mechanismPUT: the reply waits for the shards

01 · the front doorFirst stop: stateless front-end servers. You are authenticated and authorized before anything touches data.

02 · the cutOn the way in, the stream is cut into nine 7 MB shards — six data, three parity. Illustrative 6+3.

03 · the dropThe shards drop into separate failure domains — different disks and racks, across at least three AZs.

04 · the releaseAll three checks tick. The card writes last. Only then does 200 OK leave.

What to watch The gate stays shut until every AZ check ticks; the card writes last, and only then 200 OK.

Act 03 · The critical detailWhy shards beat copies

01 · two rolesSix data shards carry the bytes. Three parity shards carry recovery math. Any six of nine rebuild the object.

02 · the costThis costs 1.5× the original size. Three full copies would cost 3.0× — and die on the third loss.

03 · the loopShards are not parked: checksums verify, scrubbing patrols, and losses are rebuilt from survivors. Redundancy is a loop.

What to watch The tiles differentiate into data and parity, then the rule lights: any six of nine rebuild. Parity is a checksum strong enough to redraw a missing piece.

Act 04 · The stress testLose a disk. Lose a datacenter. Still readable.

SHARD PANEL 6+3 · ILLUS.
AZ 1
AZ 2
AZ 3

shards alive 9/9 · needed 6 · margin +3 · status: readable

WORST CASE, SHOWN STATICAZ 1 LOST
AZ 1 D1 D4 P1
AZ 2 D2 D5 P2
AZ 3 D3 D6 P3

6 of 9 alive — still readable, zero margin. Enable JavaScript for the interactive version.

01 · break thingsTap shards to break them. The object stays readable while six of nine survive.

02 · lose a buildingKill a whole Availability Zone — three shards gone, six left. Readable, with zero margin.

03 · repairRepair rebuilds the missing shards from survivors, into healthy failure domains. Eleven nines is this loop, winning forever.

What to watch Each failure eats margin, not the object. Below six alive, the object is gone.

Act 05 · The way outGET: the card answers first, then the shards race

01 · the card firstAuth again — then the card says where this version's shards live, before any bytes move.

02 · the raceFetches race in parallel. The fastest six answer, the object is rebuilt, and it streams out — in milliseconds.

03 · separate promisesAvailability is a separate promise: 99.99% a year — tens of minutes. Express One Zone trades the other AZs for single-digit milliseconds.

What to watch Six tiles lift and answer first, numbered in arrival order; the stragglers never gate the read.

STANDARD VS EXPRESS ONE ZONE — WHAT A LOST BUILDING CHANGES
STANDARD At least three AZs. Milliseconds to tens of milliseconds. Designed to survive the loss of an Availability Zone.
EXPRESS ONE ZONE One AZ. Single-digit milliseconds. Still designed for eleven nines of durability — but the loss of that one AZ is not survived.

Both classes quote the same durability design target; the trade is latency and cost against building-level resilience.

Act 06 · The freshness ruleWrite v2, read v2 — always

01 · the old worldBefore December 2020, a read right after a write could briefly answer with the old value — caches lagged.

02 · the orderingNow every write takes a sequence number — v2 lands as 102 — and a witness tracks the latest committed version.

03 · fresh or bypassedA cache may answer only if it can prove it is not behind the witness. Otherwise, straight to the authoritative store. You always read v2.

What to watch The witness jumps to 102; the stale cache is checked and rerouted — never served.

Act 07 · The synthesis

Eleven nines is a loop you can audit

The whole machine in one picture: a front door that checks identity, a strongly consistent card, an erasure-coded fleet across at least three AZs, and a repair loop holding it all in place. AWS has published research on the parts — ShardStore, a formally verified Rust storage node, and Physalia, metadata built from millions of tiny consensus cells.

  1. Front doorIdentity and permissions checked before anything touches data or metadata.
  2. Metadata planeThe always-right index card — strongly consistent, with the witness as read barrier.
  3. Data fleetThe erasure-coded bytes, spread across at least three Availability Zones.
  4. Repair loopChecksums, scrubbing, and reconstruction holding redundancy in place forever.
  5. Ack gate200 OK only after redundancy is real — and DELETE honored as a valid command.

What this does not promise: your own DELETE. Durability protects against infrastructure failure — not against you.

THE HONEST LIMIT

Durability, availability, backup — three different axes. Eleven nines will not save you from an authorized delete; versioning will. Cross-region replication covers the regional disaster.

Keep four questions for any storage system — S3's answers are one polished instance. The questions transfer; the answers do not.

Q1

Where does the metadata live?S3: a strongly consistent index, replicated, with a witness read barrier.

Q2

How is the data spread?S3: erasure-coded shards across at least three Availability Zones.

Q3

When does it say done?S3: only after the redundancy target is met — then metadata commits, then 200 OK.

Q4

What repairs it?S3: checksums, background scrubbing, and reconstruction from surviving shards.

HISTORY — PER-PREFIX CEILINGS (ABOUT 3,500 WRITES/S, 5,500 READS/S) SHAPED KEY-NAMING HABITS. AWS HAS SINCE AUTOMATED PARTITION SPLITTING; TREAT THEM AS HISTORY, NOT LIMITS.

The whole idea

Durability is a race between failure and repair, won by design.

Eleven nines is not a property of any disk, rack, or building. It is a loop: split every object into shards, spread them across at least three Availability Zones, acknowledge only when redundancy is real, and repair losses faster than the world causes them.

The same loop tells you what S3 is not: not a local disk, and not a backup against your own DELETE. Meet any storage system and ask: where is the metadata, how is the data spread, when does it say done, and what repairs it?