This squashes all local history (768 commits) onto the scrubbed PR #90
baseline. Individual commits were lost during filter-repo corruption;
the working tree content is preserved intact.
Build: N/A (working tree state only)
Squash the four overlapping feature branches into a single change set against
main, eliminating cross-PR merge conflicts and the duplicated CI-fix scripts.
What this brings in (merge order #79 -> #80 -> #81 -> #89):
- #79 refactor(infra): shared utilities (4-Infrastructure/lib/*: q16, hashing,
jsonl, fraction_utils) + the scripts/math-first/* validators that the
math-check CI requires.
- #80 feat(lean): Semantics.E8Sidon (1025 lines) -- Eisenstein coefficient
identity E4^2 = E8 and the Sidon framework. E4_sq_eq_E8_coeff is fully proved
(all Fourier-coefficient extraction machine-checked); the single residual gap
is pinned to E4_sq_eq_E8_qExpansion (Mathlib lacks the valence formula /
dim M8 = 1). 4 sorries + 1 axiom (e8_additive_completeness), all TODO(lean-port).
- #81 refactor(lean): Float-free FixedPoint core (integer-only sqrt/log2/expNeg).
E8Sidon.lean kept at #80's final 1025-line version (the #81 intermediate
438-line copy was overridden by merge order).
- #89 feat(lean): Semantics.RRC.PolyFactorIdentity -- short-sleeve polynomial
detection at the zerocopy limb boundary; now imports Semantics.E8Sidon for
sigma3/sigma7/convolutionLHS (single source of truth) instead of inlining them.
Conflict resolution:
- flake.nix -> canonical rs-surface removal (Garnix shutdown).
- scripts/math-first/* -> byte-identical across branches, clean.
- .cursorrules / AGENTS.md -> unified; baselines + sorry inventory refreshed.
Verification:
- lake build (default aggregator): 3573 jobs, 0 errors.
- lake build Semantics.RRC.PolyFactorIdentity (E8Sidon + FixedPoint + PolyFactor):
3655 jobs, 0 errors. Witnesses verified (sigma7 4 = 16513, convolutionLHS 6 = 2350).
- Python tests: 68/68 pass.
Note: the "Workers Builds: researchstack" check is a preexisting external
Cloudflare build unrelated to this change (no branch touches 4-Infrastructure/cloudflare/).
Build: 3573 jobs (default), 3655 jobs (narrow), 0 errors
Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com>
Follow-up to PR #10. Addresses comments left by Devin Review.
Primary fix (the BUG comment, .pre-commit-config.yaml:86-87):
The receipt-required-for-math-content hook used files: '<math-track
regex>' with pass_filenames: true. Pre-commit applies that regex to
the staged file list BEFORE invoking the hook, so evidence files
(receipts under shared-data/artifacts/deepseek_review/, claims.yaml)
were stripped from argv. require_math_evidence.py then saw only the
math-track files, found no evidence, and exited 1 -- even when proper
evidence was committed alongside. The only case that worked was
Lean-only commits, because Lean files are dual-classified as both
math-track and evidence.
Fix: drive the hook from the index instead of argv.
* require_math_evidence.py grows a --staged mode that runs
'git diff --cached --name-only' itself, plus a mutex check so
--staged, --from-git-diff, and explicit FILES cannot be combined.
* .pre-commit-config.yaml hook switches to always_run: true,
pass_filenames: false, and 'entry: ... --staged'. The script
exits 0 early when no math-track files are staged, so the cost
of always_run is negligible.
Polish:
* claims-registry.schema.json: add required: ["status"] inside each
'if' subschema. Without it, an entry missing 'status' would also
spuriously trip the 'then' clauses (review_receipts, lean) before
the top-level required catch. Pure error-message cleanup.
* validate_claims_registry.py: replace the catch-all
re.compile(r'^[A-Za-z]+:') with a closed list of well-known URI
schemes (http, https, arxiv, doi, isbn, mailto, urn).
Module-name-shaped strings like 'Module:Theorem' will no longer
silently bypass the on-disk path check.
* validate_claims_registry.py: thread a FormatChecker through the
Draft202012Validator so format-keyword behaviour matches
validate_deepseek_receipts.py. No-op for today's schema but
cheap insurance for the next contributor who adds 'format'.
Regression tests:
* New scripts/math-first/test_require_math_evidence.py covers ten
classification cases plus the actual --staged regression: it spins
up a temp git repo, stages a math-track file + a receipt, invokes
the script with --staged, and asserts exit 0. Without the fix this
case fails, demonstrating the bug end-to-end.
* math-check.yml runs the new self-tests in CI.
Docs: * docs/math-first-tooling.md: document the --staged contract, why
always_run + pass_filenames: false is necessary, and how to run
the new self-tests.
Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com>
pre-commit stashes unstaged changes, runs hooks, then pops the stash.
When the runner's working tree has LFS pointer files but git's LFS
smudge filter is configured (per .gitattributes), the stash/pop cycle
reports a phantom diff against the binary content git thinks it
should smudge, and the pop fails with
"the patch applies to ... which does not match the current contents".
All hooks themselves pass on this PR (validated locally and visible in
the previous CI run for #10). Clearing the LFS filters locally for the
pre-commit job removes the disagreement without mutating the repo or
any LFS-tracked files.
Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com>
Adds automated guardrails so mathematical rigor is enforced by tooling
instead of by convention. See docs/math-first-tooling.md for the full
contract.
Schemas + registry:
- shared-data/schemas/deepseek-review-receipt.schema.json
Draft 2020-12 schema for the existing ollama_deepseek_review_receipt_v1
and ollama_deepseek_review_continuation_receipt_v1 receipt formats. Pins
sha256:<hex> hashes, non-negative token counts, repo-relative POSIX
paths, and rejects additional fields.
- shared-data/schemas/claims-registry.schema.json
Schema for claims.yaml. Requires review_receipts when status is
verified-by-ai and a lean source when status is formally-proven.
- claims.yaml
Initial registry entry: prime-gap-entropy-collapse (verified-by-ai)
linked to the two existing receipts under
shared-data/artifacts/deepseek_review/.
Validators (scripts/math-first/):
- validate_deepseek_receipts.py: validates tracked or passed receipts
against the JSON Schema; shared by pre-commit and CI.
- test_validate_deepseek_receipts.py: positive + 7 negative fixtures
asserting exit-code behaviour.
- validate_claims_registry.py: schema check + unique id check + on-disk
existence check for every referenced repo-relative path.
- require_math_evidence.py: gate that requires a DeepSeek receipt, a
Lean change, or a claims.yaml update alongside edits to math-track
surfaces (Lean Semantics kernels, ArithmeticSpec docs, stack
solidification receipts).
Pre-commit (.pre-commit-config.yaml):
- check-json, check-yaml, end-of-file-fixer, trim trailing whitespace,
detect-private-key (scoped to math-first files only per AGENTS.md
Do Not Sweep).
- Local hooks wiring all three math-first validators above.
CI (.github/workflows/math-check.yml):
- validate-schemas: compiles every schema, runs both validators, runs
the validator self-tests, then re-invokes the canonical Ollama
emitter in --verify-only mode against every tracked receipt to
re-check answer_sha256 against the answer-file bytes on disk.
- require-evidence: enforces the math-track evidence rule at PR scope.
- pre-commit: runs all pre-commit hooks against the PR diff so the
contract holds even for contributors who skip installing hooks
locally.
MCP (.mcp.json):
- filesystem, sympy, wolfram-alpha, lean, deepseek-review entries
pointing at off-the-shelf upstream servers and at the canonical
ollama_deepseek_review_emitter.py. Secrets stay in the runtime env
(WOLFRAM_ALPHA_APPID, OLLAMA_API_KEY) and are never embedded.
Docs (docs/math-first-tooling.md):
- Philosophy, surfaces, schema reference, registry workflow, hook
catalogue, CI catalogue, MCP catalogue, end-to-end verify command.
shared-data/schemas/*.schema.json and claims.yaml live under paths the
top-level .gitignore would normally exclude; they are force-added via
git add -f the same way existing promoted receipts under
shared-data/artifacts/deepseek_review/ are tracked (per AGENTS.md).
Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com>