Follow-up to PR #10. Addresses comments left by Devin Review. Primary fix (the BUG comment, .pre-commit-config.yaml:86-87): The receipt-required-for-math-content hook used files: '<math-track regex>' with pass_filenames: true. Pre-commit applies that regex to the staged file list BEFORE invoking the hook, so evidence files (receipts under shared-data/artifacts/deepseek_review/, claims.yaml) were stripped from argv. require_math_evidence.py then saw only the math-track files, found no evidence, and exited 1 -- even when proper evidence was committed alongside. The only case that worked was Lean-only commits, because Lean files are dual-classified as both math-track and evidence. Fix: drive the hook from the index instead of argv. * require_math_evidence.py grows a --staged mode that runs 'git diff --cached --name-only' itself, plus a mutex check so --staged, --from-git-diff, and explicit FILES cannot be combined. * .pre-commit-config.yaml hook switches to always_run: true, pass_filenames: false, and 'entry: ... --staged'. The script exits 0 early when no math-track files are staged, so the cost of always_run is negligible. Polish: * claims-registry.schema.json: add required: ["status"] inside each 'if' subschema. Without it, an entry missing 'status' would also spuriously trip the 'then' clauses (review_receipts, lean) before the top-level required catch. Pure error-message cleanup. * validate_claims_registry.py: replace the catch-all re.compile(r'^[A-Za-z]+:') with a closed list of well-known URI schemes (http, https, arxiv, doi, isbn, mailto, urn). Module-name-shaped strings like 'Module:Theorem' will no longer silently bypass the on-disk path check. * validate_claims_registry.py: thread a FormatChecker through the Draft202012Validator so format-keyword behaviour matches validate_deepseek_receipts.py. No-op for today's schema but cheap insurance for the next contributor who adds 'format'. Regression tests: * New scripts/math-first/test_require_math_evidence.py covers ten classification cases plus the actual --staged regression: it spins up a temp git repo, stages a math-track file + a receipt, invokes the script with --staged, and asserts exit 0. Without the fix this case fails, demonstrating the bug end-to-end. * math-check.yml runs the new self-tests in CI. Docs: * docs/math-first-tooling.md: document the --staged contract, why always_run + pass_filenames: false is necessary, and how to run the new self-tests. Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com> |
||
|---|---|---|
| .agents/plugins | ||
| .consolidation-manifests | ||
| .github | ||
| .vscode | ||
| 0-Core-Formalism | ||
| 1-Distributed-Systems | ||
| 2-Search-Space | ||
| 3-Mathematical-Models | ||
| 4-Infrastructure | ||
| 5-Applications | ||
| 6-Documentation | ||
| 6-Kernel-Shim | ||
| ai-math-discovery-systems | ||
| docs | ||
| hardware/tangnano9k/rtl/generated | ||
| hooks | ||
| info | ||
| plugins/substack-connector | ||
| scripts | ||
| shared-data | ||
| tools/kotc | ||
| workspace-config | ||
| .codeiumignore | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| .mcp.json | ||
| .pre-commit-config.yaml | ||
| .python-version | ||
| Agda_Test.agda | ||
| Agda_Test.agdai | ||
| AGENTS.md | ||
| ARCHITECTURE.md | ||
| CHANGELOG.md | ||
| changes.zip | ||
| CITATION.cff | ||
| claims.yaml | ||
| CONCEPTS.md | ||
| config | ||
| description | ||
| GETTING_STARTED.md | ||
| HEAD | ||
| lean_hardware_checker_summary.json | ||
| LICENSE | ||
| NOTICE | ||
| package-lock.json | ||
| package.json | ||
| PROJECT_MAP.md | ||
| pyrightconfig.json | ||
| README.md | ||
| test.lean | ||
| THIRD_PARTY_NOTICES.md | ||
| TODO_MAP.md | ||
Research-Stack (OTOM)
Ultra-low power, zero-decimal data routing and compression.
If you just stumbled across this repository, you might see words like "Topological State Machine" and "Manifold Points" and assume this is dense, academic magic. It isn't.
This project is actually built on a very simple, grounded idea: Modern computing is incredibly wasteful.
Right now, running AI or compressing massive datasets requires giant, power-hungry GPUs because they rely on Floating-Point Math (heavy decimals like 3.14159...). We prove that you don't need decimals. You can map complex data (like the grammar of the English language) into structural shapes, and navigate them using only simple integers (whole numbers).
Because we only use addition, subtraction, multiplication, and modulo, this system can run on a $15 blank-slate microchip (an FPGA) instead of a massive server farm.
🛠️ How we know it works (Zero Guesswork)
We do not guess that our integer math works. We prove it. The core of this project is written in Lean 4, a strict mathematical theorem prover. If our logic has a flaw, the code physically will not compile. We currently have over 3,500 mathematically verified proofs securing this engine.
Python, Rust, and Verilog only exist in this repository to act as "dumb pipes" to feed data into our proven mathematical core.
📁 Repository Structure (By Goal)
Everything is numbered so you know exactly what depends on what.
| Folder | What it actually is (Plain English) |
|---|---|
0-Core-Formalism/ |
The Brain. Lean 4 code. The mathematically proven integer arithmetic. This is the source of truth. |
1-Distributed-Systems/ |
The Network. Code for making multiple computers talk to each other to share the workload. |
2-Search-Space/ |
The Navigator. Algorithms that search through our data shapes to find the best routes. |
3-Mathematical-Models/ |
The Library. Where we store our databases of equations and compressed English grammar shapes. |
4-Infrastructure/ |
The Drivers. Code that physically talks to the hardware, GPUs, and APIs. |
5-Applications/ |
The Executables. Python scripts that run the system end-to-end. (These are just shims connecting data to our Lean 4 Brain). |
6-Documentation/ |
The Manual. Where you'll find plain-English explanations and our theoretical papers. |
shared-data/ |
Raw data, cache, and exported files. |
📖 Where to start?
If you are new here, read these two files first:
- Explanation for Humans - A translation guide for our technical jargon.
- Calculator-Plain Math - Proof that every complex concept we use can be calculated on a high-school graphing calculator.
🚀 Quick Start
# 1. Compile the mathematically proven core (Takes ~1-2 minutes)
cd "0-Core-Formalism/lean/Semantics"
lake build
# 2. Run the English Manifold Builder (Compresses English via grammar shapes)
cd "../../.."
python3 5-Applications/scripts/redpajama_english_manifold.py
⚖️ The One Rule for Contributors
Lean is the source of truth.
If you add logic, it goes in 0-Core-Formalism/lean/Semantics/ and must be mathematically proven. Python scripts may not contain complex math, branching logic, or cost functions. Python is just the delivery boy for Lean.
Research Stack — All Rights Reserved
