mirror of
https://github.com/allaunthefox/SilverSight.git
synced 2026-07-31 01:25:21 +00:00
Lake build results (run 019f2f3f, 24min, 8 vCPUs): - HCMR.lean: BUILT ✅ (agent fix worked — removed excess omega, downgraded false theorem from > to ≥) - CacheSieve.lean: FAILED ❌ (agent fix incomplete — still has errors) - Blitter6502OISC.lean: FAILED ❌ (type class synthesis L62, rewrite failures L135/L142) - YangMillsPerformance.lean: FAILED ❌ (10 errors — omega/linarith can't handle Nat.div, 1 sorry) - WorkloadTestbench.lean: NOT REACHED (depends on failed CacheSieve) - CRTSidonN.lean: FAILED ❌ (errors after agent fix) Root causes: - YangMills: proofs use omega/linarith for Nat.div goals (wrong tactic) - Blitter: type class instance missing, rw patterns don't match - CacheSieve: agent fix didn't fully resolve all issues - CRTSidonN: compilation errors remain For future builds: add 'lake exe cache get' before 'lake build' to download precompiled Mathlib oleans (saves ~20min).
154 lines
9.1 KiB
Markdown
154 lines
9.1 KiB
Markdown
# Complete Test Matrix: Every Assumption, Every Question
|
||
|
||
**Status:** DEFINITIVE — test every claim, answer every question
|
||
**Date:** 2026-07-04
|
||
**Principle:** Good science tests all assumptions and answers all questions that can be answered.
|
||
|
||
---
|
||
|
||
## The Questions
|
||
|
||
For each component of the framework, we ask:
|
||
|
||
1. **Does it work?** (receipt: measurement)
|
||
2. **When is it needed?** (receipt: regime where removing it causes failure)
|
||
3. **When is it redundant?** (receipt: regime where removing it changes nothing)
|
||
4. **What are its limits?** (receipt: input where it breaks)
|
||
5. **What does it cost?** (receipt: timing/complexity measurement)
|
||
|
||
---
|
||
|
||
## Test Matrix
|
||
|
||
### CRT Encode Engine
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T01 | Does CRT preserve Sidon? | Sidon labels, coprime moduli | 0 collisions after embed | CRTSidon.lean, 0 sorries | DONE (proven) |
|
||
| T02 | Does CRT CREATE Sidon from non-Sidon? | [0,1,2,3,4], wrapping moduli | collisions drop to 0 | sidon_preservation_creation.md §6.3 | TESTED (small instances) |
|
||
| T03 | Is CRT redundant for already-Sidon input? | [1,2,4,8,16,32,64,128] | Same collision count with/without CRT | C3 run 019f2f07 | DONE: YES, redundant |
|
||
| T04 | Does CRT wrapping work at scale? | Non-Sidon, n=21 labels | collisions drop to 0 | C11 (new) | PENDING |
|
||
| T05 | Is CRT position-invariant for Sidon? | Permute labels across positions | Same collision count | C3 run 019f2f07 | DONE: YES, invariant |
|
||
| T06 | What does CRT cost vs direct check? | n=21 labels | CRT slower by CRT overhead factor | Timing comparison | PENDING |
|
||
|
||
### Chiral System (ChiralLabel, Rossby drift)
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T07 | Does Rossby drift vary across configs? | 256 positional swaps | drift varies | C3 run 019f2f07 | DONE: NO, constant (-65536) |
|
||
| T08 | Does Kelvin regime (drift=0) exist in 65K? | 4^8 cross-enriched | Some configs have drift=0 | C5 | PENDING |
|
||
| T09 | Does COUCH Kelvin check filter anything? | 65K configs | Some configs rejected (drift=0) | C5 | PENDING |
|
||
| T10 | Is Rossby drift position-invariant? | Permute chiral labels | Same drift (symmetric sum) | C3 + math | DONE: YES, symmetric |
|
||
| T11 | Does chiral permutation affect quaternion products? | 1,i,j,k basis, permuted | Different products for different perms | C4 | PENDING |
|
||
|
||
### Sidon Filter
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T12 | Is Sidon check sufficient for uniqueness? | Sidon labels | All pairwise sums distinct | CRTSidon.lean | DONE (proven) |
|
||
| T13 | Does quaternion Sidon differ from CRT Sidon? | Same labels, both filters | Different collision counts | C4 | PENDING |
|
||
| T14 | Can Sidon filter discriminate chiral configs? | 256 positional swaps | Variance in collision count | C3 | DONE: NO, invariant |
|
||
| T15 | What inputs break the Sidon filter? | Adversarial labels | Find input with false positive/negative | Fuzz test | PENDING |
|
||
|
||
### COUCH Gate
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T16 | Does COUCH reject Kelvin configs? | drift=0 configs | Rejected | C5 | PENDING |
|
||
| T17 | Does COUCH reject high-scarred configs? | All-scarred config | Rejected (self_loop ≥ threshold) | C3 run | DONE: YES (all passed, but no all-scarred in 64) |
|
||
| T18 | Is COUCH necessary (what fails without it)? | Remove COUCH, run pipeline | More configs reach Sidon (wasted work) | A/B test | PENDING |
|
||
| T19 | Does COUCH correlate with QUBO tractability? | QUBO instances with known difficulty | Kelvin = hard, Rossby = easy | C10 | PENDING |
|
||
|
||
### HCMR (Hardware Contention)
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T20 | Are self-loop values measured or assumed? | EPYC KVM benchmark | Log with 0.823, 0.885, 0.0 | C7 | PENDING (search Research Stack) |
|
||
| T21 | Does self-loop = Sidon collision rate? | Correlate self-loop with collision count | Positive correlation | C10 variant | PENDING |
|
||
| T22 | Does ring dispatch = 0 collisions? | All-Sidon labels, ring regime | self_loop=0, 100% Sidon | C3 run | DONE: consistent (0 collisions, but self_loop not measured independently) |
|
||
|
||
### Hoffman Bound / Spectral Database
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T23 | Is Hoffman gap=1 universal for unit-distance? | de Grey 1581, other 5-chromatic | gap=1 | C9 | PENDING (needs --full run) |
|
||
| T24 | Is Hoffman tight for all regular graphs? | Path, cycle, complete, more | tight | hn_spectral_database.json | DONE: YES (3/3) |
|
||
| T25 | Does Welch-Wynn improve on Hoffman? | Same graphs | WW ≥ Hoffman | hn_spectral_database.json | DONE: NO (WW=None, implementation incomplete) |
|
||
|
||
### q-Profile
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T26 | Is q>1 preference robust? | 10+ random Sidon label sets | q>1 always 100% Sidon | C8 | PENDING |
|
||
| T27 | Does q=1 always degenerate? | q=1 configs | 0% or low Sidon | crt_qprofile_sweep.json | DONE: q=1 has some Sidon (not always degenerate) |
|
||
| T28 | Is q-profile about modulus selection or encoding? | Vary L0 directly (not via q) | Same result | Controlled experiment | PENDING |
|
||
|
||
### Conservation Law
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T29 | Is compression always bounded by K(data)? | 8 different compressors | All fail to beat K(data) | weird_machine_conservation_law.md | DONE: YES (8×) |
|
||
| T30 | Does filtering avoid the conservation bound? | Filter configs, measure output size | Output < K(data) is NOT required (filtering ≠ compression) | By construction | DONE: YES (filtering is selection, not compression) |
|
||
|
||
### Pipeline
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T31 | Does the pipeline reduce search space? | 256 configs | Output < input | C3 run 019f2f07 | DONE: 256→64 (packer cap) |
|
||
| T32 | Is each stage necessary? | Remove each stage, re-run | Some stages are redundant for some inputs | A/B per stage | PENDING |
|
||
| T33 | Does cross-enrichment (4^8) run on GPU? | chiral_cross_enrich.wgsl | 65K results | C6 | PENDING |
|
||
| T34 | Is the pipeline faster on GPU than CPU? | Same input, GPU vs CPU | GPU faster for n≥8 | Timing comparison | PENDING |
|
||
|
||
### Lean Formalization
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T35 | Does CRTSidonN compile? | lake build | exit 0 | C1 | DONE: FAIL ❌ (errors after agent fix) |
|
||
| T36 | Does HCMR suite compile? | lake build | exit 0, 2 sorries | C2 | DONE: PARTIAL — HCMR ✅, CacheSieve ❌, Blitter ❌, YangMills ❌, WorkloadTestbench not reached |
|
||
| T37 | Are the proven theorems non-tautological? | Review each proof | At least sidon_preserved_mod is non-trivial | Adversarial review | DONE: sidon_preserved (componentwise) is tautological, sidon_preserved_mod is real |
|
||
| T38 | Does native_decide use violate AGENTS.md? | helical_coverage_74 | Uses native_decide | HopfFibration.lean | DONE: YES, uses native_decide (documented exception for finite decidability) |
|
||
|
||
### QUBO/QAOA Bridge
|
||
|
||
| Test | Question | Input | Prediction | Receipt | Status |
|
||
|------|----------|-------|------------|---------|--------|
|
||
| T39 | Does COUCH predict QAOA tractability? | QUBO instances | Kelvin=hard, Rossby=easy | C10 | PENDING |
|
||
| T40 | Does golden angle select good QAOA architecture? | 28 architectures | Some are better than random | Experiment | PENDING |
|
||
| T41 | Does 65K→100 pre-filtering help QAOA? | Filtered vs unfiltered QAOA | Filtered converges faster | Experiment | PENDING |
|
||
| T42 | Is the quaternion gate mapping (1=I,i=X,j=Y,k=Z) used? | Check if pipeline applies gates | Gates are NOT applied (filtering only) | Code review | DONE: NO, pipeline filters, doesn't apply gates |
|
||
|
||
---
|
||
|
||
## Execution Priority
|
||
|
||
### Blocking (must answer first)
|
||
- T35/T36: Does the Lean compile? (lake build running, ~89%)
|
||
- T04: Does CRT wrapping work at scale? (C11)
|
||
- T11: Does chiral permutation affect quaternion products? (C4)
|
||
|
||
### High value (changes the theory)
|
||
- T08/T09: Does Kelvin regime exist and get filtered? (C5)
|
||
- T13: Does quaternion Sidon differ from CRT Sidon? (C4)
|
||
- T19: Does COUCH correlate with QUBO? (C10)
|
||
- T23: Is Hoffman gap=1 universal? (C9)
|
||
|
||
### Robustness (confirms existing results)
|
||
- T26: Is q>1 robust across label sets? (C8)
|
||
- T06: What does CRT cost? (timing)
|
||
- T32: Is each pipeline stage necessary? (A/B)
|
||
|
||
### Quantum (speculative, needs Phase 1-2 first)
|
||
- T39-T41: QUBO/QAOA experiments
|
||
|
||
---
|
||
|
||
## Rule
|
||
|
||
Every test has:
|
||
1. A QUESTION (what we want to know)
|
||
2. An INPUT (what to feed it)
|
||
3. A PREDICTION (what we expect)
|
||
4. A RECEIPT (what would prove it)
|
||
5. A REJECTION (what would disprove it — implicit: opposite of prediction)
|
||
|
||
No test is skipped. No result is assumed. Every assumption gets a receipt or gets killed.
|