mirror of
https://github.com/allaunthefox/SilverSight.git
synced 2026-08-12 13:10:35 +00:00
Records the full compression investigation as a repo doc plus the scripts that back every number (real coders, byte-exact lossless round-trips, no straw baselines). One law: recoverable <=> sparse/structured; no method beats K(data), schemes only relocate bits between model and residual columns. Findings (all measured): char-poly = integrity receipt not compressor; Braille/T9 = 4.167 b/B, loses to xz; "16D/583x" GW ringdown = zero-noise self-fit artifact, ~1.5x tying/losing to LPC on noisy strain; frozen-model conservation law (k=3 smallest tape, worst total); Semantic Mass Number = base conversion (1.00-1.10x, bijection); capstone superposition/compressed-sensing cliff (recoverable iff k <= ~d/log N). Honest home for all: receipts/addresses/ recoverability gates (GCCL/RRC), never the ratio column. docs/research/COMPRESSION_HONEST_FINDINGS.md + scripts/compression/ (7 scripts + README). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
134 lines
7.4 KiB
Markdown
134 lines
7.4 KiB
Markdown
# Compression: Honest Findings (Receipt, Not Ratio)
|
||
|
||
**Date:** 2026-07-03 **Status:** measured, reproducible **Doctrine:** OTOM honest-measurement / anti-smuggle
|
||
|
||
This document records a series of compression ideas that were tested **head-on with
|
||
real coders and lossless round-trips**, and the single law they all obey. Every
|
||
"novel compressor" explored here turns out to be either **base conversion** (a
|
||
bijection — no gain) or **arithmetic coding in a costume** (needs a shipped model —
|
||
conservation), and every one loses to stock `xz`/`LPC`. The recurring honest
|
||
conclusion: these constructions are **receipts / addresses / recoverability gates**,
|
||
never compression ratios.
|
||
|
||
Reproducible scripts live in [`scripts/compression/`](../../scripts/compression/).
|
||
All numbers below are from real runs (arithmetic coders verified lossless by
|
||
byte-exact round-trip; entropies are the rate a coder provably reaches).
|
||
|
||
## The one law
|
||
|
||
> **Recoverable ⟺ sparse/structured.** No method beats `K(data)` (Kolmogorov complexity).
|
||
> A scheme can only **move bits between two columns** — the *model/program* and the
|
||
> *residual/tape* — never reduce their **sum** below the data's information content.
|
||
> "Compression" that appears otherwise is measuring against a straw baseline, or is
|
||
> lossy recoverability (rate–distortion) mislabeled as lossless.
|
||
|
||
## Baselines (enwik8, 100 MB, all lossless)
|
||
|
||
| method | bits/byte | note |
|
||
|---|---|---|
|
||
| order-0 entropy | 5.08 | symbol frequencies only |
|
||
| **order-2 PPM (our real arithmetic coder)** | **3.088** | lossless round-trip verified; our honest floor |
|
||
| gzip -9 | 2.92 | LZ77 + Huffman |
|
||
| zstd -19 | 2.16 | |
|
||
| **xz -9 (LZMA)** | **1.989** | long-range dictionary matching — the number to beat |
|
||
| cmix (SOTA) | ~1.2 | 461-model context mixing |
|
||
|
||
**Wall 1** — a low-rank/spectral summary of the byte co-occurrence matrix can never
|
||
beat the raw counts (the empirical counts are the MLE; any rank-k reconstruction has
|
||
cross-entropy ≥ the full model). **Wall 2** — even the *full* order-2 counts (3.088)
|
||
lose to stock `xz` (1.989), because LZMA sees kilobytes of context, not 2 bytes.
|
||
|
||
## Findings
|
||
|
||
### 1. Characteristic polynomial / chiral LUT — a *receipt*, not a compressor
|
||
Text reconstructs from **vocab + indices**; the Faddeev–LeVerrier char-poly is computed
|
||
*downstream* of the indices and is never read during decompression. It adds bytes
|
||
(poly +3 B, chiral LUT +60 B on a toy). Its honest role is a **GCCL integrity receipt**
|
||
(tamper-evidence), never a ratio. Script: `spectral_vs_empirical.py`, `ac_roundtrip.py`.
|
||
|
||
### 2. Braille → T9 → hachimoji (three-layer) on text — dead
|
||
Real lossless three-layer coder: **4.167 bits/byte on enwik8** — worse than order-2 PPM
|
||
(3.088), far behind xz (1.989). A 64-cell alphabet cannot represent 256 byte values
|
||
without lossy mapping; the T9 disambiguation cost exceeds the 6→3-bit saving.
|
||
|
||
### 3. "16D geometric / golden-spiral braid" GW ringdown — 583× is a zero-noise artifact
|
||
Claim: a 16D parametric fit compresses a gravitational-wave ringdown 583×. **It was
|
||
measuring the generator round-tripping itself** (a signal synthesized from ~9 coefficients,
|
||
"compressed" back to those coefficients). Honest test (`gw_honest_test.py`): synthetic
|
||
2-QNM ringdown **+ realistic detector noise**, int16-quantized, lossless, with the
|
||
parametric coder **handed the TRUE mode parameters** (steelman), raced vs LPC-8/xz on the
|
||
same stream:
|
||
|
||
| SNR regime | PARAM* b/s | LPC-8 b/s | best ratio |
|
||
|---|---|---|---|
|
||
| clean (σ=0) | 0.14 | 0.35 | 111× ← the big ratio lives ONLY here |
|
||
| 60 dB | 7.02 | 7.57 | 2.3× |
|
||
| 40 dB | 9.81 | 9.87 | tie |
|
||
| 30 dB (realistic loud) | 10.58 | 10.61 | 1.5× (tie) |
|
||
| 20 dB (weak) | 10.94 | 10.92 | **LPC wins** |
|
||
|
||
*PARAM handed the true params — upper bound on any 16D/braid fit.* Lossless ⇒ store the
|
||
residual = signal − fit = **the noise**, which is incompressible. Honest ceiling on noisy
|
||
strain ≈ **1.5×, tying/losing to 50-year-old LPC**. "16D braid" = LPC in a costume.
|
||
|
||
### 4. "Turing-complete weird machine" — the conservation law
|
||
A decompressor is already a Turing machine; total self-contained size = |program (frozen
|
||
model)| + |tape (coded residual)|. Frozen order-k byte model + real arithmetic coder,
|
||
150 KB held-out test, all rows lossless (`frozen_model_compress.py`):
|
||
|
||
| order k | test b/char | tape B | model B | TOTAL self | amortized |
|
||
|---|---|---|---|---|---|
|
||
| 0 | 5.430 | 101,812 | 440 | 102,252 | 101,812 |
|
||
| 1 | 4.043 | 75,806 | 9,728 | 85,534 | 75,806 |
|
||
| 2 | 3.253 | 61,000 | 104,116 | 165,116 | 61,000 |
|
||
| 3 | 2.975 | 55,777 | 501,392 | **557,169** | 55,777 |
|
||
|
||
As k↑, tape shrinks but model explodes; k=3 has the **smallest tape and the worst total**.
|
||
You only **relocate** bits. xz-9 self-contained (35,492 B) beats every total. A frozen
|
||
**LLM** + arithmetic coder is a genuine strong compressor **in the amortized column only**
|
||
(model shared out-of-band, sub-1 b/byte — cf. Bellard `ts_zip`, DeepMind "Language Modeling
|
||
Is Compression"); the instant the model must ship (Hutter), the model column dominates and
|
||
it loses. Long-context LLM stacks (StreamingLLM, Infini-attention, Activation Beacon, CEPE,
|
||
E2LLM) are **lossy** — right for RAG/QA, wrong for lossless reconstruction.
|
||
|
||
### 5. Semantic Mass Number as a lossless compressor — base conversion
|
||
Encoding a message as one "mass number" (OTOM Semantic Nuclide Address / "10-adic residue")
|
||
**IS base conversion** — a bijection, hence cannot compress below its radix
|
||
(`mass_number_compress.py`, 100 KB, round-trips PASS):
|
||
|
||
| method | bits/char | ratio | note |
|
||
|---|---|---|---|
|
||
| M1 base-256 mass number | 8.000 | 1.00× | literally the data |
|
||
| M2 mixed-radix over Σ (|Σ|=155) | 7.276 | 1.10× | drops unused-symbol slack only |
|
||
| M3 freq-weighted (=arithmetic) | 5.401 | 1.47× | radices ARE the model → must ship |
|
||
| xz -9 | 2.551 | 3.14× | crushes all |
|
||
|
||
OTOM's own **de-anthropocentric revision** already flags "English-facing semantic
|
||
compression" as the OLD ERROR and defines MassNumber as the **admissibility/recoverability
|
||
RECEIPT** projected from SemanticMass. The measurement confirms the doctrine.
|
||
|
||
### 6. Capstone — "how LLMs drop data recoverably" = superposition = compressed sensing
|
||
Pack N=1024 features into d=128 dims via a random projection (8× squeeze); recover a
|
||
k-sparse vector (`superposition_recovery.py`, OMP + matched-filter, lossless-in-the-limit):
|
||
|
||
| k active | OMP rel-err | support | verdict |
|
||
|---|---|---|---|
|
||
| ≤16 | ~3e-16 | 100% | **lossless-ish** (superposition works) |
|
||
| 24 | 9.5e-5 | 100% | degrading |
|
||
| 32 | 6.7e-2 | 93% | degrading |
|
||
| 48 | 0.80 | 34% | **LOST** |
|
||
| ≥64 | ≥1.0 | ~15% | LOST |
|
||
| 1024 (dense/random) | 1.5 | 13% = chance | unrecoverable |
|
||
|
||
Compressed-sensing bound `k* ≈ d/(2 ln N) = 9`. **An LLM "drops data recoverably" only in
|
||
the sparse regime** — recoverability is bought entirely by sparsity, not free compression.
|
||
Past the bound → polysemantic interference → data lost. Dense/random data goes to chance.
|
||
|
||
## Bottom line
|
||
|
||
Structure (sparsity/redundancy) is the only thing **any** method — polynomial, mass-number,
|
||
LLM superposition — can recover; the incompressible part is always lost or stored raw. Keep
|
||
these constructions where they honestly belong: **receipts, addresses, and recoverability
|
||
gates** (GCCL/RRC), never in the compression-ratio column. To actually reduce bytes, the
|
||
levers are the boring, established ones: **context depth + long-range matching** (LZMA,
|
||
context mixing) for text, **linear prediction** (LPC/FLAC) for smooth signals.
|