The reason quantum computers don't work yet, the reason they eventually will, and the exact door AI walked in through.
You'll be able to explain how a surface code finds an error without reading the data, and why a threshold matters.
One qubit is too fragile to trust, so you weave one logical qubit out of many physical ones. You never look at the data. You measure checks between neighbors, and a decoder infers what went wrong. Done well, errors get exponentially rarer as you add qubits; done badly, adding qubits makes things worse. The inferring is a pattern-recognition problem, which is why an AI now out-guesses thirty years of hand-written algorithms on accuracy, with speed as the remaining fight.
A qubit is a singer who slowly drifts off key. Everything nudges it: a warm wire, a stray bit of radiation, the qubit next door. On good hardware, roughly one operation in a thousand goes wrong. That sounds tolerable until you hear that useful algorithms need billions of error-free steps (trillions of physical operations underneath) to land correctly, back to back.
A classical computer solves this with copies: store the bit three times, take the majority vote. Quantum mechanics slams both doors on that. You can't copy an unknown qubit: the no-cloning theorem isn't an engineering limit, it's the law. And you can't peek at one to see if it's still right, because measuring a qubit destroys the very superposition you were trying to protect.
So here's the trick, and it's one of the best ideas physics produced in the last forty years: don't listen to any singer alone; listen to the harmony between neighbors. Ask pairs, "are you two still in tune with each other?" That question is safe: it reveals nothing about the song, only whether someone drifted, and roughly where. A conductor, the decoder, collects those harmony reports and works out who went off key, without ever hearing a solo.
For thirty years, humans wrote that conductor's rulebook by hand, assuming the choir fails in tidy, textbook ways. Real choirs don't. In 2024, Google DeepMind replaced the rulebook with an AI, AlphaQubit, that learned a real machine's actual bad habits from its data, and it made fewer wrong calls than every hand-written rulebook it was tested against. (It's still too slow to run live; making learned decoders fast is the current race.) The thing AI is worst at (being exactly right) is fixing the thing quantum computing can't exist without. That loop is this site's whole subject.
Real terms for the same picture: Working: keep going ↓
WorkingEncode one logical qubit into a grid, a surface code, where data qubits are interleaved with measurement qubits. The measurement qubits repeatedly run parity checks (stabilizers) on their data-qubit neighbors. A check that fires tells you an odd number of its neighbors flipped; the full pattern of fired checks is the syndrome. It's a diagnosis from symptoms: you never see the disease, only which alarms it tripped.
| Piece | Job | Failure mode |
|---|---|---|
| Data qubits | Hold the encoded information | Flip and drift, constantly |
| Parity checks | Report "someone here flipped" without reading data | The measurements themselves are noisy too |
| Decoder | Infer the most likely error from the syndrome | Guesses wrong when noise breaks its assumptions |
| Logical qubit | The one perfect qubit the whole grid pretends to be | Fails only when errors slip past all of the above |
The grid's size is its code distance d: a d=3 code corrects any single flip and detects any pair, but a chain of 3 errors crossing the grid is invisible, changing the logical qubit while every check stays quiet. Bigger d, longer chains needed, safer qubit.
And the number that rules the whole field: the threshold. If the physical error rate is below roughly 1% per operation (model-dependent), enlarging the grid makes the logical qubit exponentially better. Above it, enlarging the grid makes it worse: more qubits, more noise, no rescue. Every headline about qubit quality is really a headline about which side of this line a machine is on. That's also why decoders matter so much: a smarter decoder effectively moves you further below threshold with zero new hardware.
Drag the physical error rate as a fraction of the threshold. Below 1, a bigger code drives the logical error rate down exponentially, and the slider shows how many physical qubits that costs. At 1 or above, the same code that was helping now hurts — the curve turns red and no distance reaches the target.
This is the scaling law the formal section below states in full, pL ≈ A (p/pth)⌊(d+1)/2⌋, drawn with A = 1. The marker is the smallest odd distance whose pL reaches the target, solved in log space (rE underflows a raw comparison by a rounding hair). It reproduces both worked figures below — r = 1/10, 10⁻¹² → d = 23, ≈1,057 physical qubits; r = 1/3 → d = 51, 5,201 — checked in a verification script before shipping. Real overhead is higher: routing, magic-state factories and correlated noise all sit on top.
This is the real thing: the distance-3 surface code ("surface-17": 9 data + 8 check qubits), the first layout serious superconducting labs build. Click data qubits to flip them. Amber patches are parity checks firing; the dashed teal rings are the decoder's proposed repair. See if you can get an error past it.
⬤ red = your bit-flip · ◼ amber = a check firing · ◌ teal dashed = the decoder's repair · faded dashed patches = phase-flip checks (idle in this demo)
Honest toy model: the layout and check structure are exact (the standard d=3 rotated surface code, Tomita & Svore 2014), and the decoder is true minimum-weight inference over all 512 bit-flip patterns. Simplified two ways: only bit-flips are injected (phase-flips behave identically under the faded X checks), and checks here are perfect one-shot measurements, whereas real hardware re-measures them ~a million times per second and the checks themselves misfire.
The flat picture squeezes two kinds of check into one plane, as a checkerboard. Here they are pulled apart: the nine data qubits sit in the middle, the teal layer above catches bit flips and the violet layer below catches phase flips. Slide the layers together and you are back at the checkerboard above.
A qubit can suffer a bit flip (X), a phase flip (Z), or both at once (Y). Each layer only ever sees its own kind, so in this toy the decoder solves two separate matching problems. That is a property of this family of codes, not a trick of the drawing. (A real decoder can do a little better by using the fact that a Y is one event, not two.) Turn it, tap a qubit, and try the presets.
Honest toy model: the layout and the checks are the same as the flat explorer above (the standard distance-3 rotated surface code, Tomita & Svore 2014), and the decoder is exact minimum-weight decoding over all 512 patterns, run once for each kind of error. The unfolding is a picture, not hardware: on a chip all 17 qubits sit in one plane. It models one round of perfect checks. Real machines repeat the checks over time, and the checks themselves can be wrong, which adds a third dimension of a different kind: time.
The same code with its time axis, plus three hardware pictures in 3D: the decoding graph, the 15-to-1 magic-state recipe, ion shuttling and the Rydberg blockade.
The widget above let you see the error. A real decoder never does. It sees which alarms fired, and nothing else, and it has to bet. So here is the actual job.
Ten rounds on a chip with a defect: somewhere in this lattice two qubits are coupled, and when one flips its partner tends to flip with it. Cross-talk is a real failure mode, and one of the exact things that breaks textbook decoders. Your opponent is minimum-weight matching, the hand-written decoder the field has run on since 2001. It is very good, and it is running the textbook noise model it was handed: every qubit fails on its own, every failure equally likely. Nobody told it about this chip. You can hand-fit a correlation back into a matching decoder (that is what “correlated matching” is), but somebody has to know what to fit.
You can. After every round you see which qubits really flipped, the same kind of device data AlphaQubit was fine-tuned on. Find the defect and you will beat it. That is the whole thesis of this site, played out on a nine-qubit grid.
◆ This is operations research, running inside a quantum computer. "Minimum-weight matching" is Edmonds' blossom algorithm from 1965, a graph-optimisation result that predates the field it now decodes for, from the paper that gave computer science its definition of an efficient algorithm. Topic 07: Matching and assignment ▸Rotated surface code with distance d: d² data qubits, d²−1 syndrome qubits; the stabilizer group is generated by weight-4 X- and Z-plaquette operators in the bulk and weight-2 operators on the boundaries. Logical operators are X- and Z-strings spanning the lattice between their respective boundaries; the code is [[d², 1, d]]. Decoding is inference over the syndrome volume: given syndrome history σ, return the most probable logical equivalence class of errors. The exact physical error is unrecoverable and irrelevant; any correction in the right class (differing from the truth by a stabilizer) succeeds. The demo above shows this: flip one qubit of a weight-2 check and the decoder may "repair" its partner instead, and that's a win.
The threshold theorem is what makes "below threshold" more than a hopeful phrase: it is the proof that arbitrarily long fault-tolerant quantum computation is possible at all, provided the physical error rate sits below a fixed constant pth and the code is scaled up to compensate, exactly the claim the scaling law below quantifies. It was proven independently, by three different routes, in the late 1990s: by concatenating small codes inside themselves recursively (Aharonov & Ben-Or, arXiv quant-ph/9906129 (SIAM J. Comput. 38, 1207 (2008), expanding their 1997 STOC result); Knill, Laflamme & Zurek, arXiv quant-ph/9702058 (Proc. R. Soc. Lond. A 454, 365 (1998))), and by a topological route closer to the surface code used here (Kitaev, Russian Mathematical Surveys 52, 1191 (1997)). The proof holds under a specific noise model, though: errors independent across qubits and gates, local rather than striking many qubits at once, and already below threshold before scaling starts. Real hardware breaks the first two by default (crosstalk correlates neighboring failures, leakage escapes the two-level qubit the proof assumes), and that gap, far from being a footnote to the theorem, is the reason correlated matching and AlphaQubit exist below.
Below threshold, logical error per round scales as
so each two-step increase in distance (d → d+2) multiplies suppression by Λ ≈ pth/p.
Google's Willow processor was the first hardware demonstration on the right side of the exponent: across a d = 3, 5, 7 series, increasing the distance by two suppressed the logical error rate by Λ = 2.14 ± 0.02 (obtained by fitting the log of the logical error rate against distance across the whole series, not from any single step; the lone d = 5 → 7 ratio is ≈ 2.1), with the d=7 logical qubit's lifetime exceeding its best physical qubit's (arXiv 2408.13687, Nature 638, 920 (2025)). The overhead is the price: 2d²−1 physical qubits per logical, so at p ≈ 10⁻³ and a 10⁻¹² logical target, d ≈ 23 → roughly a thousand physical qubits per logical before routing and magic-state factories. The same overhead logic, run at the gentler error budget a real attack circuit needs (~9×10⁷ Toffoli gates → a ~10⁻⁹–10⁻¹⁰ target, not 10⁻¹²), is how "≈1,200 logical qubits to break a Bitcoin key" (2026 estimate) becomes half a million physical ones: about 400 physical per logical, magic-state factories included.
Magic-state factories, named above but never explained: Clifford-group gates alone (H, S, CNOT) are not universal. The Gottesman–Knill theorem shows any circuit built only from Clifford gates, stabilizer-state inputs, and Pauli measurements is efficiently simulable on an ordinary classical computer (Gottesman's 1997 Caltech thesis, arXiv quant-ph/9705052, already gives a polynomial-time simulation algorithm; sped up and extended to mixed states by Aaronson & Gottesman, arXiv quant-ph/0406196, Phys. Rev. A 70, 052328 (2004)). So a fault-tolerant computer that only ever ran Clifford gates would be classically simulable too, i.e. pointless. Universality needs one more ingredient, typically the T-gate, and most codes can't apply it directly and stay fault-tolerant, so instead it's injected via a prepared magic state. Injected magic states start noisy, and they don't inherit the surrounding code's protection for free, so they are purified by distillation: consume several noisy copies, output fewer higher-fidelity ones, repeat until the error is low enough for the algorithm ahead (Bravyi & Kitaev, arXiv quant-ph/0403025, Phys. Rev. A 71, 022316 (2005)). That consumption is not a rounding error on top of the surface-code overhead above: early resource estimates found distillation dominating the physical-qubit budget so consistently that a later paper's entire contribution was proving it didn't have to (Litinski, arXiv 1905.06903, Quantum 3, 205 (2019), titled, pointedly, "Magic State Distillation: Not as Costly as You Think").
Distillation, done for real. Until 2025 that whole picture was circuit-level accounting on paper. In July 2025 a QuEra/Harvard/MIT team ran distillation on logical qubits, not raw physical ones, on a neutral-atom processor: encoding magic states in d=3 and d=5 color codes and distilling them, they raised output fidelity to 99.4% from a 95.1% input at d=3 (an 8× cut in infidelity) and to 98.6% from 92.5% at d=5 (a 6× cut) (arXiv 2412.15165, Nature 645, 620 (2025)). The output beating every input is the result: distillation stacked on top of error correction, the architecture every serious resource estimate assumes, run end-to-end on hardware for the first time, on a different platform from Willow's superconducting transmons and a different layer of the stack from AlphaQubit's decoding.
Transversal gates. A transversal gate applies the same physical operation to every physical qubit of a logical block independently (one qubit at a time, never a two-qubit gate between two physical qubits inside the same block), so a single faulty gate can corrupt at most one physical qubit's worth of information and has nowhere to spread within the block; that's what makes the gate fault-tolerant by construction rather than by extra circuitry bolted on afterward. It has a hard ceiling, though: the Eastin–Knill theorem proves that any code able to detect an arbitrary error on a single physical qubit cannot have a gate set that is both transversal and universal (Eastin & Knill, arXiv 0811.4262 (Phys. Rev. Lett. 102, 110502 (2009))). That forces universality to come from outside the transversal gate set entirely, which is exactly why the magic-state machinery two paragraphs up isn't a separate topic from this one: distillation is how that outside ingredient gets purified to usable fidelity, precisely because transversality alone has been proven unable to supply it. On the surface code specifically: logical CNOT between two patches is transversal, but only if the patches are arranged with each physical qubit adjacent to its partner, which a flat 2D layout can't generally give two side-by-side patches; logical Hadamard is transversal on a single patch too, but a bare transversal-H layer swaps the patch's X- and Z-type boundaries, so recovering the original orientation costs a geometric rotation of the patch on top of the gate; S is not transversal on the surface code at all and has to be reached some other way, such as gate teleportation through an ancilla state.
Lattice surgery. Instead of executing a two-qubit logical gate transversally (which, as above, needs qubit i of one patch physically adjacent to qubit i of the other), lattice surgery gets there by physically merging adjacent surface-code patches: briefly, a new set of joint stabilizers is measured across the boundary where two patches touch, amounting to a joint measurement of a logical operator, X̄X̄ or Z̄Z̄, after which the patches are split back apart. A single merge-and-split measures a joint parity; it is not by itself a deterministic CNOT. The actual logical-CNOT protocol (Horsman, Fowler, Devitt & Van Meter, arXiv 1111.4022 (New Journal of Physics 14, 123011 (2012))) needs a third patch, an ancilla prepared in a fixed logical state, merged with the control patch and split, then merged with the target patch and split: two merge-split cycles across three patches, sequenced right, rather than one operation between two. The saving is still real, just not "no ancilla at all": every patch in that sequence is an ordinary surface-code patch measured in the same 2D plane, never a separate teleportation channel or a third dimension to line qubits up in. The paper's own stated aim is the previous paragraph made concrete: coupling planar codes without transversal operations, so the array stays nearest-neighbor, touching only along shared edges. That's the mechanism, not a side detail, behind moving logical information around the kind of array the overhead figures above describe: the [[d²,1,d]] structure and its 2d²−1-per-logical-qubit cost describe a static patch, and a real circuit still has to bring logical qubits together to interact, and merge-and-split is how that happens without breaking the flat, nearest-neighbor layout those qubit counts assume in the first place.
Where AI enters. Minimum-weight perfect matching decodes the graph with edge weights from an assumed independent-Pauli noise model: provably good under that model, degraded by what real chips do (correlated errors, leakage, cross-talk, drift). AlphaQubit (arXiv 2310.05900, Nature 2024) is a recurrent transformer over per-round syndrome inputs, crucially consuming analog "soft" readout rather than binarized measurements, pretrained on simulated samples and fine-tuned on device data. Reported: ~6% fewer logical errors than tensor-network decoders (near-optimal but far too slow for real time) and ~30% fewer than correlated matching, maintained to d=11 in simulation.
The real-time constraint. A decoder slower than the syndrome cycle accumulates backlog and the machine stalls: accuracy without latency is a benchmark, not a decoder. The frontier moved in 2026: an FPGA neural decoder reached 550 ns system latency (124 ns inference) inside a 1.25 μs cycle on superconducting hardware (arXiv 2605.04892). Open problems: scaling learned decoders to large d in real time, and doing it for the multi-logical-qubit patches an actual algorithm needs.
• Surface codes, canonical intro: Fowler, Mariantoni, Martinis & Cleland, arXiv 1208.0928 (PRA 2012)
• The d=3, 17-qubit layout ("Surface-17", the name came later): Tomita & Svore, arXiv 1404.3747
• Threshold theorem: Aharonov & Ben-Or, arXiv quant-ph/9906129 (SIAM J. Comput. 38, 1207 (2008)); Knill, Laflamme & Zurek, arXiv quant-ph/9702058 (Proc. R. Soc. Lond. A 454, 365 (1998)); Kitaev, Russian Mathematical Surveys 52, 1191 (1997)
• Below-threshold scaling (Willow): Google Quantum AI, arXiv 2408.13687 (Nature 638, 920 (2025); online Dec 2024)
• Gottesman–Knill theorem: Gottesman, arXiv quant-ph/9705052; Aaronson & Gottesman, arXiv quant-ph/0406196 (Phys. Rev. A 70, 052328 (2004))
• Magic-state distillation: Bravyi & Kitaev, arXiv quant-ph/0403025 (Phys. Rev. A 71, 022316 (2005)); Litinski, arXiv 1905.06903 (Quantum 3, 205 (2019))
• Logical-qubit magic-state distillation (QuEra/Harvard/MIT, neutral atoms): arXiv 2412.15165 (Nature 645, 620 (2025))
• Transversal gates, the fundamental limit: Eastin & Knill, arXiv 0811.4262 (Phys. Rev. Lett. 102, 110502 (2009))
• Lattice surgery: Horsman, Fowler, Devitt & Van Meter, arXiv 1111.4022 (New Journal of Physics 14, 123011 (2012))
• AlphaQubit: arXiv 2310.05900 (Nature 2024) · Google announcement
• Real-time FPGA neural decoder: arXiv 2605.04892
Check yourself
A parity check measures whether two neighbouring qubits agree. Why does that not destroy the logical state?
A stabilizer measurement returns a parity, and parity is exactly the information that is orthogonal to the encoded value. The check learns that something flipped without learning what the logical qubit is, which is why a surface code survives being measured a million times a second. This is ordinary stabilizer physics, and it is the single idea the whole field rests on.
Error correction is the multiplier in every "quantum breaks Bitcoin" estimate. The numbers, with caveats.
Can quantum break Bitcoin? →Who's furthest below threshold. The table is still empty, and the page says why.
Scoreboard →