Adds python/spectral_codebook_db.py: sync codebook rows into
ene.rrc_predictions on the neon-64gb Postgres (NEON_PG convention from
scripts/auto/auto_pipeline.py, default research_stack DB).
- Row shape (flat, SQL-typed, Spark-JDBC readable): equation_id,
proxy_pred = cluster codeword C0..C8, exact_pred = shape from exact
lambda under CURRENT ClassifyN.lean thresholds (1.5/4.0 Q16.16,
integer semantics mirrored), matrix_hash =
'charpoly=<c1..c8>;pos10=<base-10 positional hash>' (similarity +
injective identity keys), confidence = 1.0 unique fingerprint / 1/k
in k-way charpoly collision class; deterministic uuid5 ids so reruns
upsert idempotently.
- SAFE BY DEFAULT: dry run prints summary + sample SQL and writes
nothing; --apply required to insert (psycopg2, with --emit-sql
data/spectral_codebook_sync.sql fallback when the driver is absent).
--apply has NOT been run; live DB untouched. --verify-schema does a
read-only column check; schema verified offline against
scripts/auto/ene_schema.sql in tests (live check left to the user
per the ask-before-DB-work rule).
- spectral_codebook.py gains --sync-db (always dry-run from that entry
point). Dry-run counts: 250 rows; C0=35 C1=20 C2=13 C3=79 C4=29
C5=15 C6=22 C7=19 C8=18; Logogram=69 Signal=131 CognitiveLoad=50;
183 rows at confidence 1.0.
- docs: 'Neon data layer' section — ENE table map, stale
ene.rrc_classifications finding (120 rows with artifact spectral
radii 0.3-0.85 predating the exact-eigenvalue fix; recommend
re-classification via this codebook), empty landing tables, Spark
JDBC snippet, arxiv-pg (podman-exec only) citation layer note.
- tests: 8 new dry-run/no-network tests (23 total).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Implements the Next Steps of docs/SPECTRAL_CODEBOOK_ANALYSIS.md as
python/spectral_codebook.py (stdlib-only; NumPy optional fast path):
- Parses the 250 8x8 braid adjacency matrices from PIST/Matrices250.lean.
- Primary fingerprint: exact integer characteristic-polynomial
coefficients via Faddeev-LeVerrier in Fraction arithmetic (196 unique
over 238 distinct matrices, vs 179 unique lambda at 4dp; 6 cospectral
non-identical groups; 11 exact-duplicate matrix groups / 23 ids).
- Corrects the analysis doc: the 9 'mid band' lambda in (0.5,1) are
power-iteration non-convergence artifacts - exact rho = 1.0 for all 9
(peripheral spectra). spectral_radius is now the exact max root
modulus (numpy eigvals or Durand-Kerner on the exact char poly);
power-iteration lambda kept only for traceability.
- Gap-aware quantization: dedupe to 238 distinct matrices, boundaries at
gaps > 3x median gap, min-support guard (>=10 distinct per cluster),
sparse-tail outlier flagging above lambda ~= 7.66. Result: 9 clusters.
- Round trip encode(matrix) -> (codeword, index) -> decode -> equation_id
verified bijective over all 250 in tests/test_spectral_codebook.py.
- 278-row RRC/Q16_16Manifold corpus: 28 extra rows are repeated ids;
boundaries reproduce exactly, no new gaps or clusters.
- Emits data/spectral_codebook.json (schema spectral_codebook_v2) with
explicit collision classes; docs/SPECTRAL_CODEBOOK_GENERATOR.md notes
the hashMatrix base-5 injectivity gap and the ClassifyN threshold
(1.5/4.0 Q16.16) vs analysis-doc (0.5/1.0) mismatch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>