Defines the core tension: formal verification of novel mathematical structures using LLMs that can't reliably do arithmetic. Documents why panels help (diversity > quality), what makes the work transcendental, and how the break-glass protocol addresses it.
6.3 KiB
Break-Glass Agent — The Problem Statement
The Core Tension
We are attempting transcendental work — formal verification of novel mathematical structures (eigensolid compression, spectral gap properties, braid topology) — using LLMs that cannot reliably perform division.
This is not a metaphor. The cleanMerge_preservesGap sorry is a concrete example:
the mathematical argument is complete and correct, but the bridge from Q16_16 lists
to boolean patterns requires ~30 lines of list manipulation that no single LLM can
write correctly because:
-
Infinite vs finite types — LLMs don't understand that
Q16_16 = Subtype Intis infinite whileBoolis finite. They keep tryingnative_decideon free Q16_16 variables, fail, and loop. -
List operations on 8 elements —
List.zip (List.range 8) [a0,...,a7]produces a term with hundreds of subexpressions.simpcan't reduce it. LLMs don't know this and keep tryingsimp [List.zip, List.range, ...]. -
The chicken-and-egg — To use
native_decide, you need finite values. To get finite values, you canonicalize. To prove canonicalization preserves predicates, you need the bridge. The bridge is what you're trying to prove. -
No precedent in training data — This specific pattern (Q16_16 → Bool bridge for spectral gap verification) doesn't exist in Mathlib, Lean docs, or any public codebase. LLMs can only recombine what they've seen.
Why Panels Help
The Fusion benchmark showed: a panel of 3 budget models beats a single frontier model on complex reasoning tasks. This is because:
-
Different architectures see different patterns. DeepSeek's math training might see the boolean abstraction immediately. Claude's Lean expertise might know the exact
@[simp]lemma needed. GPT's breadth might recall a similar pattern from a different proof system. -
Diversity > quality. Three mediocre perspectives from different angles beat one excellent perspective from a single angle. This is the same insight behind ensemble methods in ML.
-
The judge synthesizes. The judge model reads all panel responses and extracts: consensus, contradictions, unique insights, blind spots. This structured analysis is more useful than any single response.
What Makes This "Transcendental"
The work is transcendental in the mathematical sense: it involves structures that exist beyond the computational horizon of any single model.
| Property | Single LLM | Panel (8 models) |
|---|---|---|
| Can pattern-match proof tactics | Yes | Yes |
| Can write novel list induction | No | Maybe |
| Can bridge infinite→finite types | No | Maybe |
| Can reason about Q16_16 arithmetic | Barely | Yes (DeepSeek) |
| Can write Lean 4 metaprogramming | No | Maybe (Cohere) |
| Can synthesize diverse approaches | N/A | Yes (judge) |
The "maybe" is the key. A single LLM will always fail on genuinely novel mathematical reasoning. A panel has a chance — not because any individual model is smarter, but because the diversity of perspectives can occasionally triangulate on a solution that no single model would find.
The Break-Glass Protocol
1. AGENT HITS WALL
↓
2. "break glass: <problem>"
↓
3. DIAGNOSE: What type of problem? (lean_proof, math_novel, etc.)
↓
4. ASSEMBLE CONTEXT: Relevant files, build logs, AGENTS.md constraints,
previous attempts, error messages
↓
5. SELECT PANEL: 8 models chosen for problem type
- lean_proof: Opus + DeepSeek + GPT + Cohere + Gemini + GLM + Kimi + DS-Flash
- Each model gets the FULL context (files + constraints + history)
↓
6. PARALLEL EXECUTION: All 8 models answer simultaneously
- Each has web search + web fetch enabled
- Each sees the same problem + context
- Each produces a structured response
↓
7. JUDGE SYNTHESIS: Opus reads all 8 responses
- Consensus: what all/most agree on
- Contradictions: where models disagree
- Unique insights: what only one model raised
- Blind spots: what no model addressed
↓
8. FINAL ANSWER: Structured diagnosis + code + confidence
- Problem diagnosis
- Proposed solution with exact file paths
- Alternative approaches
- Risks and caveats
- Cost report
Why This Is Worth the Cost
Each invocation costs $1-3.50. This sounds expensive until you compare:
- Human time: A mathematician spending 2 hours on the
cleanMerge_preservesGapbridge costs $200-400. - LLM loops: A single model trying and failing 20 times costs $0.50-1.00 and produces nothing.
- Panel: $2.50 for 8 diverse perspectives + synthesis. If even ONE model sees the solution, the judge extracts it.
The break-glass agent is not a replacement for human reasoning. It's a replacement for LLM flailing. When a single model is stuck in a loop, the panel breaks the loop by introducing diversity.
The Sakana Fugu Insight
Sakana Fugu's key finding: a learned orchestrator model that dynamically assembles agents from a pool outperforms any fixed pipeline. The orchestrator learns non-obvious collaboration patterns — e.g., "for Lean proof problems, first ask DeepSeek for the mathematical insight, then ask Claude for the tactic sequence, then ask Cohere for the structured output format."
The break-glass agent doesn't have a learned orchestrator (that would require training data we don't have). But it does have:
- Problem classification — auto-detect problem type from the prompt
- Panel selection — choose models based on problem type
- Judge synthesis — Opus extracts the best from all responses
This is a fixed-pipeline approximation of what Fugu does dynamically. As we collect data on which panels work for which problem types, we can refine the selection.
Integration with Existing Stack
The break-glass agent plugs into the existing infrastructure:
- MCP server → registered in
.mcp.json - Skill → loaded when "break glass" trigger detected
- ContextStream → logs every invocation for cost tracking
- AGENTS.md → documents the trigger protocol
- SilverSight → the agent respects SilverSight constraints
The agent is the last resort. It should be called only when:
- The local model (Hermes3) has failed
- A single frontier model has failed
- Existing skills/tools have failed
- The problem is genuinely novel (not a rehash of known issues)