SilverSight/docs/research/COMPRESSION_HONEST_FINDINGS.md
allaun 99c943dcd0 docs(compression): honest-findings writeup + reproducible scripts
Records the full compression investigation as a repo doc plus the scripts that
back every number (real coders, byte-exact lossless round-trips, no straw
baselines). One law: recoverable <=> sparse/structured; no method beats K(data),
schemes only relocate bits between model and residual columns.

Findings (all measured): char-poly = integrity receipt not compressor; Braille/T9
= 4.167 b/B, loses to xz; "16D/583x" GW ringdown = zero-noise self-fit artifact,
~1.5x tying/losing to LPC on noisy strain; frozen-model conservation law
(k=3 smallest tape, worst total); Semantic Mass Number = base conversion
(1.00-1.10x, bijection); capstone superposition/compressed-sensing cliff
(recoverable iff k <= ~d/log N). Honest home for all: receipts/addresses/
recoverability gates (GCCL/RRC), never the ratio column.

docs/research/COMPRESSION_HONEST_FINDINGS.md + scripts/compression/ (7 scripts + README).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 15:52:27 -05:00

7.4 KiB
Raw Blame History

Compression: Honest Findings (Receipt, Not Ratio)

Date: 2026-07-03 Status: measured, reproducible Doctrine: OTOM honest-measurement / anti-smuggle

This document records a series of compression ideas that were tested head-on with real coders and lossless round-trips, and the single law they all obey. Every "novel compressor" explored here turns out to be either base conversion (a bijection — no gain) or arithmetic coding in a costume (needs a shipped model — conservation), and every one loses to stock xz/LPC. The recurring honest conclusion: these constructions are receipts / addresses / recoverability gates, never compression ratios.

Reproducible scripts live in scripts/compression/. All numbers below are from real runs (arithmetic coders verified lossless by byte-exact round-trip; entropies are the rate a coder provably reaches).

The one law

Recoverable ⟺ sparse/structured. No method beats K(data) (Kolmogorov complexity). A scheme can only move bits between two columns — the model/program and the residual/tape — never reduce their sum below the data's information content. "Compression" that appears otherwise is measuring against a straw baseline, or is lossy recoverability (ratedistortion) mislabeled as lossless.

Baselines (enwik8, 100 MB, all lossless)

method bits/byte note
order-0 entropy 5.08 symbol frequencies only
order-2 PPM (our real arithmetic coder) 3.088 lossless round-trip verified; our honest floor
gzip -9 2.92 LZ77 + Huffman
zstd -19 2.16
xz -9 (LZMA) 1.989 long-range dictionary matching — the number to beat
cmix (SOTA) ~1.2 461-model context mixing

Wall 1 — a low-rank/spectral summary of the byte co-occurrence matrix can never beat the raw counts (the empirical counts are the MLE; any rank-k reconstruction has cross-entropy ≥ the full model). Wall 2 — even the full order-2 counts (3.088) lose to stock xz (1.989), because LZMA sees kilobytes of context, not 2 bytes.

Findings

1. Characteristic polynomial / chiral LUT — a receipt, not a compressor

Text reconstructs from vocab + indices; the FaddeevLeVerrier char-poly is computed downstream of the indices and is never read during decompression. It adds bytes (poly +3 B, chiral LUT +60 B on a toy). Its honest role is a GCCL integrity receipt (tamper-evidence), never a ratio. Script: spectral_vs_empirical.py, ac_roundtrip.py.

2. Braille → T9 → hachimoji (three-layer) on text — dead

Real lossless three-layer coder: 4.167 bits/byte on enwik8 — worse than order-2 PPM (3.088), far behind xz (1.989). A 64-cell alphabet cannot represent 256 byte values without lossy mapping; the T9 disambiguation cost exceeds the 6→3-bit saving.

3. "16D geometric / golden-spiral braid" GW ringdown — 583× is a zero-noise artifact

Claim: a 16D parametric fit compresses a gravitational-wave ringdown 583×. It was measuring the generator round-tripping itself (a signal synthesized from ~9 coefficients, "compressed" back to those coefficients). Honest test (gw_honest_test.py): synthetic 2-QNM ringdown + realistic detector noise, int16-quantized, lossless, with the parametric coder handed the TRUE mode parameters (steelman), raced vs LPC-8/xz on the same stream:

SNR regime PARAM* b/s LPC-8 b/s best ratio
clean (σ=0) 0.14 0.35 111× ← the big ratio lives ONLY here
60 dB 7.02 7.57 2.3×
40 dB 9.81 9.87 tie
30 dB (realistic loud) 10.58 10.61 1.5× (tie)
20 dB (weak) 10.94 10.92 LPC wins

PARAM handed the true params — upper bound on any 16D/braid fit. Lossless ⇒ store the residual = signal fit = the noise, which is incompressible. Honest ceiling on noisy strain ≈ 1.5×, tying/losing to 50-year-old LPC. "16D braid" = LPC in a costume.

4. "Turing-complete weird machine" — the conservation law

A decompressor is already a Turing machine; total self-contained size = |program (frozen model)| + |tape (coded residual)|. Frozen order-k byte model + real arithmetic coder, 150 KB held-out test, all rows lossless (frozen_model_compress.py):

order k test b/char tape B model B TOTAL self amortized
0 5.430 101,812 440 102,252 101,812
1 4.043 75,806 9,728 85,534 75,806
2 3.253 61,000 104,116 165,116 61,000
3 2.975 55,777 501,392 557,169 55,777

As k↑, tape shrinks but model explodes; k=3 has the smallest tape and the worst total. You only relocate bits. xz-9 self-contained (35,492 B) beats every total. A frozen LLM + arithmetic coder is a genuine strong compressor in the amortized column only (model shared out-of-band, sub-1 b/byte — cf. Bellard ts_zip, DeepMind "Language Modeling Is Compression"); the instant the model must ship (Hutter), the model column dominates and it loses. Long-context LLM stacks (StreamingLLM, Infini-attention, Activation Beacon, CEPE, E2LLM) are lossy — right for RAG/QA, wrong for lossless reconstruction.

5. Semantic Mass Number as a lossless compressor — base conversion

Encoding a message as one "mass number" (OTOM Semantic Nuclide Address / "10-adic residue") IS base conversion — a bijection, hence cannot compress below its radix (mass_number_compress.py, 100 KB, round-trips PASS):

method bits/char ratio note
M1 base-256 mass number 8.000 1.00× literally the data
M2 mixed-radix over Σ ( Σ =155) 7.276
M3 freq-weighted (=arithmetic) 5.401 1.47× radices ARE the model → must ship
xz -9 2.551 3.14× crushes all

OTOM's own de-anthropocentric revision already flags "English-facing semantic compression" as the OLD ERROR and defines MassNumber as the admissibility/recoverability RECEIPT projected from SemanticMass. The measurement confirms the doctrine.

6. Capstone — "how LLMs drop data recoverably" = superposition = compressed sensing

Pack N=1024 features into d=128 dims via a random projection (8× squeeze); recover a k-sparse vector (superposition_recovery.py, OMP + matched-filter, lossless-in-the-limit):

k active OMP rel-err support verdict
≤16 ~3e-16 100% lossless-ish (superposition works)
24 9.5e-5 100% degrading
32 6.7e-2 93% degrading
48 0.80 34% LOST
≥64 ≥1.0 ~15% LOST
1024 (dense/random) 1.5 13% = chance unrecoverable

Compressed-sensing bound k* ≈ d/(2 ln N) = 9. An LLM "drops data recoverably" only in the sparse regime — recoverability is bought entirely by sparsity, not free compression. Past the bound → polysemantic interference → data lost. Dense/random data goes to chance.

Bottom line

Structure (sparsity/redundancy) is the only thing any method — polynomial, mass-number, LLM superposition — can recover; the incompressible part is always lost or stored raw. Keep these constructions where they honestly belong: receipts, addresses, and recoverability gates (GCCL/RRC), never in the compression-ratio column. To actually reduce bytes, the levers are the boring, established ones: context depth + long-range matching (LZMA, context mixing) for text, linear prediction (LPC/FLAC) for smooth signals.