Adds automated guardrails so mathematical rigor is enforced by tooling instead of by convention. See docs/math-first-tooling.md for the full contract. Schemas + registry: - shared-data/schemas/deepseek-review-receipt.schema.json Draft 2020-12 schema for the existing ollama_deepseek_review_receipt_v1 and ollama_deepseek_review_continuation_receipt_v1 receipt formats. Pins sha256:<hex> hashes, non-negative token counts, repo-relative POSIX paths, and rejects additional fields. - shared-data/schemas/claims-registry.schema.json Schema for claims.yaml. Requires review_receipts when status is verified-by-ai and a lean source when status is formally-proven. - claims.yaml Initial registry entry: prime-gap-entropy-collapse (verified-by-ai) linked to the two existing receipts under shared-data/artifacts/deepseek_review/. Validators (scripts/math-first/): - validate_deepseek_receipts.py: validates tracked or passed receipts against the JSON Schema; shared by pre-commit and CI. - test_validate_deepseek_receipts.py: positive + 7 negative fixtures asserting exit-code behaviour. - validate_claims_registry.py: schema check + unique id check + on-disk existence check for every referenced repo-relative path. - require_math_evidence.py: gate that requires a DeepSeek receipt, a Lean change, or a claims.yaml update alongside edits to math-track surfaces (Lean Semantics kernels, ArithmeticSpec docs, stack solidification receipts). Pre-commit (.pre-commit-config.yaml): - check-json, check-yaml, end-of-file-fixer, trim trailing whitespace, detect-private-key (scoped to math-first files only per AGENTS.md Do Not Sweep). - Local hooks wiring all three math-first validators above. CI (.github/workflows/math-check.yml): - validate-schemas: compiles every schema, runs both validators, runs the validator self-tests, then re-invokes the canonical Ollama emitter in --verify-only mode against every tracked receipt to re-check answer_sha256 against the answer-file bytes on disk. - require-evidence: enforces the math-track evidence rule at PR scope. - pre-commit: runs all pre-commit hooks against the PR diff so the contract holds even for contributors who skip installing hooks locally. MCP (.mcp.json): - filesystem, sympy, wolfram-alpha, lean, deepseek-review entries pointing at off-the-shelf upstream servers and at the canonical ollama_deepseek_review_emitter.py. Secrets stay in the runtime env (WOLFRAM_ALPHA_APPID, OLLAMA_API_KEY) and are never embedded. Docs (docs/math-first-tooling.md): - Philosophy, surfaces, schema reference, registry workflow, hook catalogue, CI catalogue, MCP catalogue, end-to-end verify command. shared-data/schemas/*.schema.json and claims.yaml live under paths the top-level .gitignore would normally exclude; they are force-added via git add -f the same way existing promoted receipts under shared-data/artifacts/deepseek_review/ are tracked (per AGENTS.md). Co-Authored-By: Allaun Silverfox <bigdataiscoming+9i37y6j2@protonmail.com> |
||
|---|---|---|
| .agents/plugins | ||
| .consolidation-manifests | ||
| .github | ||
| .vscode | ||
| 0-Core-Formalism | ||
| 1-Distributed-Systems | ||
| 2-Search-Space | ||
| 3-Mathematical-Models | ||
| 4-Infrastructure | ||
| 5-Applications | ||
| 6-Documentation | ||
| 6-Kernel-Shim | ||
| ai-math-discovery-systems | ||
| docs | ||
| hardware/tangnano9k/rtl/generated | ||
| hooks | ||
| info | ||
| plugins/substack-connector | ||
| scripts | ||
| shared-data | ||
| tools/kotc | ||
| workspace-config | ||
| .codeiumignore | ||
| .gitattributes | ||
| .gitignore | ||
| .gitmodules | ||
| .mcp.json | ||
| .pre-commit-config.yaml | ||
| .python-version | ||
| Agda_Test.agda | ||
| Agda_Test.agdai | ||
| AGENTS.md | ||
| ARCHITECTURE.md | ||
| CHANGELOG.md | ||
| changes.zip | ||
| CITATION.cff | ||
| claims.yaml | ||
| CONCEPTS.md | ||
| config | ||
| description | ||
| GETTING_STARTED.md | ||
| HEAD | ||
| lean_hardware_checker_summary.json | ||
| LICENSE | ||
| NOTICE | ||
| package-lock.json | ||
| package.json | ||
| PROJECT_MAP.md | ||
| pyrightconfig.json | ||
| README.md | ||
| test.lean | ||
| THIRD_PARTY_NOTICES.md | ||
| TODO_MAP.md | ||
Research-Stack (OTOM)
Ultra-low power, zero-decimal data routing and compression.
If you just stumbled across this repository, you might see words like "Topological State Machine" and "Manifold Points" and assume this is dense, academic magic. It isn't.
This project is actually built on a very simple, grounded idea: Modern computing is incredibly wasteful.
Right now, running AI or compressing massive datasets requires giant, power-hungry GPUs because they rely on Floating-Point Math (heavy decimals like 3.14159...). We prove that you don't need decimals. You can map complex data (like the grammar of the English language) into structural shapes, and navigate them using only simple integers (whole numbers).
Because we only use addition, subtraction, multiplication, and modulo, this system can run on a $15 blank-slate microchip (an FPGA) instead of a massive server farm.
🛠️ How we know it works (Zero Guesswork)
We do not guess that our integer math works. We prove it. The core of this project is written in Lean 4, a strict mathematical theorem prover. If our logic has a flaw, the code physically will not compile. We currently have over 3,500 mathematically verified proofs securing this engine.
Python, Rust, and Verilog only exist in this repository to act as "dumb pipes" to feed data into our proven mathematical core.
📁 Repository Structure (By Goal)
Everything is numbered so you know exactly what depends on what.
| Folder | What it actually is (Plain English) |
|---|---|
0-Core-Formalism/ |
The Brain. Lean 4 code. The mathematically proven integer arithmetic. This is the source of truth. |
1-Distributed-Systems/ |
The Network. Code for making multiple computers talk to each other to share the workload. |
2-Search-Space/ |
The Navigator. Algorithms that search through our data shapes to find the best routes. |
3-Mathematical-Models/ |
The Library. Where we store our databases of equations and compressed English grammar shapes. |
4-Infrastructure/ |
The Drivers. Code that physically talks to the hardware, GPUs, and APIs. |
5-Applications/ |
The Executables. Python scripts that run the system end-to-end. (These are just shims connecting data to our Lean 4 Brain). |
6-Documentation/ |
The Manual. Where you'll find plain-English explanations and our theoretical papers. |
shared-data/ |
Raw data, cache, and exported files. |
📖 Where to start?
If you are new here, read these two files first:
- Explanation for Humans - A translation guide for our technical jargon.
- Calculator-Plain Math - Proof that every complex concept we use can be calculated on a high-school graphing calculator.
🚀 Quick Start
# 1. Compile the mathematically proven core (Takes ~1-2 minutes)
cd "0-Core-Formalism/lean/Semantics"
lake build
# 2. Run the English Manifold Builder (Compresses English via grammar shapes)
cd "../../.."
python3 5-Applications/scripts/redpajama_english_manifold.py
⚖️ The One Rule for Contributors
Lean is the source of truth.
If you add logic, it goes in 0-Core-Formalism/lean/Semantics/ and must be mathematically proven. Python scripts may not contain complex math, branching logic, or cost functions. Python is just the delivery boy for Lean.
Research Stack — All Rights Reserved
