SilverSight/docs/weird_machine_conservation_law.md
openresearch 723992c567 docs: add π tape LUT coda — cleanest conservation law proof
Measured on real π (1M digits): offset digits ≈ data digits,
slope exactly 1. The pointer-into-π is the same size as the data.

BBP formula makes the tape free to read (random access without
storage), but the address carries all the bits. Free shelf, call
number as long as the book.

π-normality only conjectured → losslessness not guaranteed.

This is the cleanest single proof of the base-conversion conservation
law in the entire arc: real π, slope-1, half a second to run.
Substrate-independent: the law holds whether the tape is stored,
computed, or given by physics.

Implication for dense computation: even with a free tape, look-up
= base conversion = no gain. Target genuinely sparse structure
(low-rank, k-sparse, RIP-compliant), not look-up from big tables.
2026-07-03 20:57:15 +00:00

9.2 KiB
Raw Blame History

Weird Machine Conservation Law: Proven with Real Bytes

The Claim That Was Tested

"A Turing-complete weird machine can beat unpredictability by finding generating programs instead of predicting."

The Conservation Law (now measured)

compressed_size = program_size + residual_size ≥ entropy_floor × data_size

The weird machine moves bits between the program column and the residual column. It never reduces the sum below the entropy floor.

Measured Results (Claude Code demo, lossless round-trip PASS)

order k tape B model B TOTAL B amortized B
0 101,812 440 102,252 101,812
1 75,806 9,728 85,534 75,806
3 55,777 501,392 557,169 55,777
xz -9 35,492 ~60KB 35,492 35,492

As k increases:

  • Tape SHRINKS (better prediction, smaller residual)
  • Model EXPLODES (every new context = bytes to ship)
  • TOTAL bottoms out at k=1, then BLOWS UP at k=3

Why xz Wins

xz's decoder is ~60KB, amortized across all files by the standard. It never ships a fat per-file model. The model column is effectively zero per file. That's why total = tape = 35,492.

The One Real Win (not Hutter)

Frozen model + arithmetic coder: k=3 amortized = 55,777 bytes, sub-xz on tape alone. A real frozen LLM would drive this lower. But the model must be shared out-of-band (not scored). The instant you ship the model (Hutter Prize), the model column dominates and you lose.

What This Permanently Gates

  • "Turing-complete weird machine beats unpredictability" = FALSE
  • Conservation forbids it. The machine is never free; it's on the invoice.
  • Generation = prediction. The generating program = the model. The residual = what can't be predicted/generated. Sum is conserved.
  • The Braille/T9/hachimoji substrate is a different decomposition, not a different bound. It changes where bits go, not whether they exist.

The GW SNR Sweep (same law, different data)

SNR program residual total ratio
clean 9 coeff 0 tiny 583x (zero-noise artifact)
60 dB 9 coeff small small 2.3x
30 dB 9 coeff noise ~floor 1.5x (ties LPC)
20 dB 9 coeff more noise ~floor 1.5x (LPC wins)

Same conservation: bits move from program to residual as noise increases. Total converges to entropy floor. Nobody beats it.

The Honest Map

Approach Text (enwik8) Signals (GW) Verdict
Order-2 PPM 3.088 b/B Honest baseline
Braille/T9 4.167 b/B Dead (worse than PPM)
16D braid 1.5x (ties LPC) Dead (adds nothing)
Polynomial Receipt Receipt Receipt, not compressor
xz 1.989 b/B The floor
cmix ~1.2 b/B SOTA (461 models)
LPC ~1.5x The signal floor
Frozen LLM + AC sub-xz (amortized) Real, but model not scored

Semantic Mass Number: Base Conversion Proof (Final Branch)

The Test

"Encode data as a semantic mass number A(H)" = represent the message as one big number (the nuclide address / 10-adic residue reading).

Measured Results (lossless round-trip PASS)

Method bits/char bytes ratio lossless
M1 base-256 mass number 8.000 100,000 1.00 PASS
M2 mixed-radix (155 symbols) 7.276 91,107 1.10 PASS
M3 freq-weighted (=arithmetic) 5.401 67,823 1.47 needs model
xz -9 (order-N + matching) 2.551 31,892 3.14 PASS

Why It Fails

Base conversion is a bijection. A bijection moves information around, never destroys it — so it cannot compress below its radix. M1 IS the data (1.00x). M2 only beats 1.00 because the data uses 155 of 256 byte values (dropping unused-symbol slack, not compression). M3 = arithmetic coding wearing a nuclide costume — and the model must ship = conservation wall.

The Doctrine Already Knew

The de-anthropocentric revision explicitly flags "English-facing semantic compression" as the OLD ERROR and redefines: "MassNumber is the admissibility / recoverability RECEIPT projected from SemanticMass."

The measurement just put numbers behind the flag.

Final Sealed Map

Idea As compressor Honest home
char-poly receipt, adds overhead GCCL integrity receipt
Braille/T9 4.167 b/B, lose to xz dead
16D / 583x GW zero-noise artifact, LPC win LPC in costume
weird-machine conservation, k=3 worst total bits relocate, never shrink
semantic mass number base conversion, 1.001.10x recoverability receipt

One rule: you can move bits between columns, never beat K(data). Everything that "compresses" is either base conversion (bijection, no gain) or arithmetic coding (needs model, ship cost). The clever geometry buys nothing over boring xz/LPC.

LLM Recoverable Drop: Same Law, Different Substrate

How LLMs Actually "Drop Data Recoverably"

  1. Residual stream = accumulate, don't overwrite. No bytes saved — the data was never dropped. The stream is a workspace, not a compressor.

  2. Superposition = pack N features into d < N dimensions via near-orthogonal directions (Anthropic's Toy Models of Superposition). Recovery is exact only when k-sparse: d ≥ k·log(N/k) (RIP bound). Past that → interference → lossy.

  3. Attention = soft retrieval, keeps everything. KV-cache eviction (StreamingLLM/beacons) is explicitly lossy — evicted tokens gone, reconstructed only approximately.

  4. Quantization = drop precision bits, recover approximately. Lossy, bounded by weight tolerance.

The Law Under All of It

Recoverable ⟺ sparse/redundant. Every mechanism recovers the structured part and loses the noise part. Feed it dense/random data and recovery fails — same wall as every compression branch measured.

Connection to Pipeline Measurements

  • GW ringdown at 30dB: signal is k=5 sparse in N=10,000 samples. RIP bound: d ≥ 5·log(2000) ≈ 37.5. Our d=9 polynomial is BELOW the bound → lossy recovery, residual = noise → 1.5x (ties LPC).

  • Order-2 PPM on enwik8: 256×256 co-occurrence packs N=65,536 transitions into d=256 contexts. Sparse text = good recovery. Dense Wikipedia = interference = 3.088 b/B residual.

  • QR decomposition (O-AMMR) IS compressed sensing. Eigenvalue spectrum IS the sparse feature set. Golden spiral IS RIP-compliant recovery. But it's LOSSY — the residual is lost.

The Honest Split

  • Lossless: dead. K(data) is conserved across every substrate.
  • Lossy recoverability: real, useful, must be judged vs JPEG/Opus at equal distortion. The QR projection recovers the dominant structure (sparse part), loses the residual (noise part). The mass number IS the honesty tag for what was kept vs lost.

Final Consistency Check

The doctrine says: "MassNumber is the admissibility / recoverability RECEIPT projected from SemanticMass." The measurement confirms:

  • As lossless compressor: dead (base conversion, 1.00x)
  • As lossy projection: real (recovers sparse structure, loses noise)
  • As receipt: exactly what the spec says (records what was kept)
  • LLM superposition: same mechanism, same RIP bound, same lossy floor

Coda: π as Tape LUT (Cleanest Proof)

The Test

"Encode data as an offset into π — every string appears in π's digits, so store just the offset."

Measured on Real π (1,000,000 digits)

k (data digits) avg 1st-occurrence pos log10(pos) = offset digits
1 9 0.96
2 118 2.07
3 983 2.99
4 10,007 4.00
5 109,256 5.04
6 403,165* 5.61*

*only 198/300 found in 10^6 digits (survivorship-biased small; extend to 10^7 and avg → ~10^6, offset → 6.00)

The Result

Slope = exactly 1. Offset digits ≈ data digits. The pointer-into-π is the same size as the data it points to. Base conversion — the offset is the data's rank in another base. Zero gain.

The BBP Nuance (Real but Doesn't Beat Conservation)

π has the BBP formula: compute the n-th hex digit without previous ones, O(n log n) time, tiny space. So π is a random-access tape you never store. The "LUT is astronomically big" objection doesn't apply.

But: random access is free; the address is not. You still must store a ~k-digit offset to name k digits of data. BBP makes the shelf free; the call number is still as long as the book.

Also: π-normality is only conjectured. "Every string appears" isn't proven. Losslessness not guaranteed for all inputs.

Verdict

π-as-tape-LUT = base conversion + a free-to-read shelf. The address carries all the bits. Same wall as semantic mass number, now wearing π. Cleanest single proof of the base-conversion law: real π, slope-1, measured in half a second.

Implication for Dense Computation Targeting

Even with a free-to-read tape (π via BBP), the address IS the information. The conservation law is substrate-independent — it doesn't matter whether the tape is stored, computed, or given by physics. The sparse structure you target must be genuinely sparse (low-rank, k-sparse, RIP-compliant), not just "looked up from a big table." Look-up = base conversion = no gain.