SilverSight/docs/weird_machine_conservation_law.md
openresearch 8935cc9eaa docs: seal compression arc — mass number = base conversion (final branch)
Claude Code tested semantic mass number as compressor:
- M1 base-256: 1.00x (IS the data, bijection)
- M2 mixed-radix: 1.10x (drops unused symbols, not compression)
- M3 freq-weighted: 1.47x (= arithmetic coding in costume, needs model)
- xz: 3.14x (crushes all)

Base conversion is a bijection — moves information, never destroys it.
Cannot compress below its radix. The doctrine already knew: mass number
= 'admissibility / recoverability RECEIPT', not compressor.

Entire compression arc now sealed end to end:
| char-poly | receipt → GCCL integrity receipt |
| Braille/T9 | 4.167 b/B → dead |
| 16D/583x | zero-noise artifact → LPC in costume |
| weird-machine | conservation law → bits relocate, never shrink |
| mass number | base conversion → recoverability receipt |

One rule: move bits between columns, never beat K(data).
Everything that compresses = base conversion (no gain) or
arithmetic coding (needs model, ship cost = conservation wall).
2026-07-03 20:46:53 +00:00

116 lines
4.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Weird Machine Conservation Law: Proven with Real Bytes
## The Claim That Was Tested
"A Turing-complete weird machine can beat unpredictability by finding
generating programs instead of predicting."
## The Conservation Law (now measured)
```
compressed_size = program_size + residual_size ≥ entropy_floor × data_size
```
The weird machine moves bits between the program column and the residual
column. It never reduces the sum below the entropy floor.
## Measured Results (Claude Code demo, lossless round-trip PASS)
| order k | tape B | model B | TOTAL B | amortized B |
|---------|--------|---------|---------|-------------|
| 0 | 101,812 | 440 | 102,252 | 101,812 |
| 1 | 75,806 | 9,728 | 85,534 | 75,806 |
| 3 | 55,777 | 501,392 | 557,169 | 55,777 |
| xz -9 | 35,492 | ~60KB | 35,492 | 35,492 |
As k increases:
- Tape SHRINKS (better prediction, smaller residual)
- Model EXPLODES (every new context = bytes to ship)
- TOTAL bottoms out at k=1, then BLOWS UP at k=3
## Why xz Wins
xz's decoder is ~60KB, amortized across all files by the standard.
It never ships a fat per-file model. The model column is effectively
zero per file. That's why total = tape = 35,492.
## The One Real Win (not Hutter)
Frozen model + arithmetic coder: k=3 amortized = 55,777 bytes,
sub-xz on tape alone. A real frozen LLM would drive this lower.
But the model must be shared out-of-band (not scored). The instant
you ship the model (Hutter Prize), the model column dominates and
you lose.
## What This Permanently Gates
- "Turing-complete weird machine beats unpredictability" = FALSE
- Conservation forbids it. The machine is never free; it's on the invoice.
- Generation = prediction. The generating program = the model.
The residual = what can't be predicted/generated. Sum is conserved.
- The Braille/T9/hachimoji substrate is a different decomposition,
not a different bound. It changes where bits go, not whether they exist.
## The GW SNR Sweep (same law, different data)
| SNR | program | residual | total | ratio |
|-----|---------|----------|-------|-------|
| clean | 9 coeff | 0 | tiny | 583x (zero-noise artifact) |
| 60 dB | 9 coeff | small | small | 2.3x |
| 30 dB | 9 coeff | noise | ~floor | 1.5x (ties LPC) |
| 20 dB | 9 coeff | more noise | ~floor | 1.5x (LPC wins) |
Same conservation: bits move from program to residual as noise increases.
Total converges to entropy floor. Nobody beats it.
## The Honest Map
| Approach | Text (enwik8) | Signals (GW) | Verdict |
|----------|---------------|--------------|---------|
| Order-2 PPM | 3.088 b/B | — | Honest baseline |
| Braille/T9 | 4.167 b/B | — | Dead (worse than PPM) |
| 16D braid | — | 1.5x (ties LPC) | Dead (adds nothing) |
| Polynomial | Receipt | Receipt | Receipt, not compressor |
| xz | 1.989 b/B | — | The floor |
| cmix | ~1.2 b/B | — | SOTA (461 models) |
| LPC | — | ~1.5x | The signal floor |
| Frozen LLM + AC | sub-xz (amortized) | — | Real, but model not scored |
## Semantic Mass Number: Base Conversion Proof (Final Branch)
### The Test
"Encode data as a semantic mass number A(H)" = represent the message
as one big number (the nuclide address / 10-adic residue reading).
### Measured Results (lossless round-trip PASS)
| Method | bits/char | bytes | ratio | lossless |
|--------|-----------|-------|-------|----------|
| M1 base-256 mass number | 8.000 | 100,000 | 1.00 | PASS |
| M2 mixed-radix (155 symbols) | 7.276 | 91,107 | 1.10 | PASS |
| M3 freq-weighted (=arithmetic) | 5.401 | 67,823 | 1.47 | needs model |
| xz -9 (order-N + matching) | 2.551 | 31,892 | 3.14 | PASS |
### Why It Fails
Base conversion is a bijection. A bijection moves information around,
never destroys it — so it cannot compress below its radix. M1 IS the
data (1.00x). M2 only beats 1.00 because the data uses 155 of 256
byte values (dropping unused-symbol slack, not compression). M3 =
arithmetic coding wearing a nuclide costume — and the model must ship
= conservation wall.
### The Doctrine Already Knew
The de-anthropocentric revision explicitly flags "English-facing
semantic compression" as the OLD ERROR and redefines:
"MassNumber is the admissibility / recoverability RECEIPT projected
from SemanticMass."
The measurement just put numbers behind the flag.
## Final Sealed Map
| Idea | As compressor | Honest home |
|------|---------------|-------------|
| char-poly | receipt, adds overhead | GCCL integrity receipt |
| Braille/T9 | 4.167 b/B, lose to xz | dead |
| 16D / 583x GW | zero-noise artifact, LPC win | LPC in costume |
| weird-machine | conservation, k=3 worst total | bits relocate, never shrink |
| semantic mass number | base conversion, 1.001.10x | recoverability receipt |
One rule: you can move bits between columns, never beat K(data).
Everything that "compresses" is either base conversion (bijection,
no gain) or arithmetic coding (needs model, ship cost). The clever
geometry buys nothing over boring xz/LPC.