Records the full compression investigation as a repo doc plus the scripts that back every number (real coders, byte-exact lossless round-trips, no straw baselines). One law: recoverable <=> sparse/structured; no method beats K(data), schemes only relocate bits between model and residual columns. Findings (all measured): char-poly = integrity receipt not compressor; Braille/T9 = 4.167 b/B, loses to xz; "16D/583x" GW ringdown = zero-noise self-fit artifact, ~1.5x tying/losing to LPC on noisy strain; frozen-model conservation law (k=3 smallest tape, worst total); Semantic Mass Number = base conversion (1.00-1.10x, bijection); capstone superposition/compressed-sensing cliff (recoverable iff k <= ~d/log N). Honest home for all: receipts/addresses/ recoverability gates (GCCL/RRC), never the ratio column. docs/research/COMPRESSION_HONEST_FINDINGS.md + scripts/compression/ (7 scripts + README). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
7.4 KiB
Compression: Honest Findings (Receipt, Not Ratio)
Date: 2026-07-03 Status: measured, reproducible Doctrine: OTOM honest-measurement / anti-smuggle
This document records a series of compression ideas that were tested head-on with
real coders and lossless round-trips, and the single law they all obey. Every
"novel compressor" explored here turns out to be either base conversion (a
bijection — no gain) or arithmetic coding in a costume (needs a shipped model —
conservation), and every one loses to stock xz/LPC. The recurring honest
conclusion: these constructions are receipts / addresses / recoverability gates,
never compression ratios.
Reproducible scripts live in scripts/compression/.
All numbers below are from real runs (arithmetic coders verified lossless by
byte-exact round-trip; entropies are the rate a coder provably reaches).
The one law
Recoverable ⟺ sparse/structured. No method beats
K(data)(Kolmogorov complexity). A scheme can only move bits between two columns — the model/program and the residual/tape — never reduce their sum below the data's information content. "Compression" that appears otherwise is measuring against a straw baseline, or is lossy recoverability (rate–distortion) mislabeled as lossless.
Baselines (enwik8, 100 MB, all lossless)
| method | bits/byte | note |
|---|---|---|
| order-0 entropy | 5.08 | symbol frequencies only |
| order-2 PPM (our real arithmetic coder) | 3.088 | lossless round-trip verified; our honest floor |
| gzip -9 | 2.92 | LZ77 + Huffman |
| zstd -19 | 2.16 | |
| xz -9 (LZMA) | 1.989 | long-range dictionary matching — the number to beat |
| cmix (SOTA) | ~1.2 | 461-model context mixing |
Wall 1 — a low-rank/spectral summary of the byte co-occurrence matrix can never
beat the raw counts (the empirical counts are the MLE; any rank-k reconstruction has
cross-entropy ≥ the full model). Wall 2 — even the full order-2 counts (3.088)
lose to stock xz (1.989), because LZMA sees kilobytes of context, not 2 bytes.
Findings
1. Characteristic polynomial / chiral LUT — a receipt, not a compressor
Text reconstructs from vocab + indices; the Faddeev–LeVerrier char-poly is computed
downstream of the indices and is never read during decompression. It adds bytes
(poly +3 B, chiral LUT +60 B on a toy). Its honest role is a GCCL integrity receipt
(tamper-evidence), never a ratio. Script: spectral_vs_empirical.py, ac_roundtrip.py.
2. Braille → T9 → hachimoji (three-layer) on text — dead
Real lossless three-layer coder: 4.167 bits/byte on enwik8 — worse than order-2 PPM (3.088), far behind xz (1.989). A 64-cell alphabet cannot represent 256 byte values without lossy mapping; the T9 disambiguation cost exceeds the 6→3-bit saving.
3. "16D geometric / golden-spiral braid" GW ringdown — 583× is a zero-noise artifact
Claim: a 16D parametric fit compresses a gravitational-wave ringdown 583×. It was
measuring the generator round-tripping itself (a signal synthesized from ~9 coefficients,
"compressed" back to those coefficients). Honest test (gw_honest_test.py): synthetic
2-QNM ringdown + realistic detector noise, int16-quantized, lossless, with the
parametric coder handed the TRUE mode parameters (steelman), raced vs LPC-8/xz on the
same stream:
| SNR regime | PARAM* b/s | LPC-8 b/s | best ratio |
|---|---|---|---|
| clean (σ=0) | 0.14 | 0.35 | 111× ← the big ratio lives ONLY here |
| 60 dB | 7.02 | 7.57 | 2.3× |
| 40 dB | 9.81 | 9.87 | tie |
| 30 dB (realistic loud) | 10.58 | 10.61 | 1.5× (tie) |
| 20 dB (weak) | 10.94 | 10.92 | LPC wins |
PARAM handed the true params — upper bound on any 16D/braid fit. Lossless ⇒ store the residual = signal − fit = the noise, which is incompressible. Honest ceiling on noisy strain ≈ 1.5×, tying/losing to 50-year-old LPC. "16D braid" = LPC in a costume.
4. "Turing-complete weird machine" — the conservation law
A decompressor is already a Turing machine; total self-contained size = |program (frozen
model)| + |tape (coded residual)|. Frozen order-k byte model + real arithmetic coder,
150 KB held-out test, all rows lossless (frozen_model_compress.py):
| order k | test b/char | tape B | model B | TOTAL self | amortized |
|---|---|---|---|---|---|
| 0 | 5.430 | 101,812 | 440 | 102,252 | 101,812 |
| 1 | 4.043 | 75,806 | 9,728 | 85,534 | 75,806 |
| 2 | 3.253 | 61,000 | 104,116 | 165,116 | 61,000 |
| 3 | 2.975 | 55,777 | 501,392 | 557,169 | 55,777 |
As k↑, tape shrinks but model explodes; k=3 has the smallest tape and the worst total.
You only relocate bits. xz-9 self-contained (35,492 B) beats every total. A frozen
LLM + arithmetic coder is a genuine strong compressor in the amortized column only
(model shared out-of-band, sub-1 b/byte — cf. Bellard ts_zip, DeepMind "Language Modeling
Is Compression"); the instant the model must ship (Hutter), the model column dominates and
it loses. Long-context LLM stacks (StreamingLLM, Infini-attention, Activation Beacon, CEPE,
E2LLM) are lossy — right for RAG/QA, wrong for lossless reconstruction.
5. Semantic Mass Number as a lossless compressor — base conversion
Encoding a message as one "mass number" (OTOM Semantic Nuclide Address / "10-adic residue")
IS base conversion — a bijection, hence cannot compress below its radix
(mass_number_compress.py, 100 KB, round-trips PASS):
| method | bits/char | ratio | note |
|---|---|---|---|
| M1 base-256 mass number | 8.000 | 1.00× | literally the data |
| M2 mixed-radix over Σ ( | Σ | =155) | 7.276 |
| M3 freq-weighted (=arithmetic) | 5.401 | 1.47× | radices ARE the model → must ship |
| xz -9 | 2.551 | 3.14× | crushes all |
OTOM's own de-anthropocentric revision already flags "English-facing semantic compression" as the OLD ERROR and defines MassNumber as the admissibility/recoverability RECEIPT projected from SemanticMass. The measurement confirms the doctrine.
6. Capstone — "how LLMs drop data recoverably" = superposition = compressed sensing
Pack N=1024 features into d=128 dims via a random projection (8× squeeze); recover a
k-sparse vector (superposition_recovery.py, OMP + matched-filter, lossless-in-the-limit):
| k active | OMP rel-err | support | verdict |
|---|---|---|---|
| ≤16 | ~3e-16 | 100% | lossless-ish (superposition works) |
| 24 | 9.5e-5 | 100% | degrading |
| 32 | 6.7e-2 | 93% | degrading |
| 48 | 0.80 | 34% | LOST |
| ≥64 | ≥1.0 | ~15% | LOST |
| 1024 (dense/random) | 1.5 | 13% = chance | unrecoverable |
Compressed-sensing bound k* ≈ d/(2 ln N) = 9. An LLM "drops data recoverably" only in
the sparse regime — recoverability is bought entirely by sparsity, not free compression.
Past the bound → polysemantic interference → data lost. Dense/random data goes to chance.
Bottom line
Structure (sparsity/redundancy) is the only thing any method — polynomial, mass-number, LLM superposition — can recover; the incompressible part is always lost or stored raw. Keep these constructions where they honestly belong: receipts, addresses, and recoverability gates (GCCL/RRC), never in the compression-ratio column. To actually reduce bytes, the levers are the boring, established ones: context depth + long-range matching (LZMA, context mixing) for text, linear prediction (LPC/FLAC) for smooth signals.