12 KiB
DeepSeek Review Process
Source: Home ·
shared-data/artifacts/deepseek_review/
The Research Stack incorporates DeepSeek AI models for formal mathematical
review and validation of canonical specifications, Lean kernels, and
statistical-interpretation receipts. This page documents the wiki-level view
of the review pipeline, the receipt schema, and the canonical example
(prime-gap entropy collapse) shipped under
shared-data/artifacts/deepseek_review/.
Important boundary: the canonical artifacts documented on this page are
Ollama-compatible review receipts. The tracked
5-Applications/tools-scripts/llm/deepseek_review_adapter.py is a separate
Anthropic-compatible DeepSeek adapter with its own schema_version: "1.0"
receipt format. Do not use that adapter as the emitter for the
ollama_deepseek_review_receipt_v1 schema unless it has first been reconciled
to this schema.
AI-Assisted Mathematical Review
DeepSeek review receipts provide reproducibility and integrity tracking for
AI-assisted mathematical validation. Each review run emits a paired
.md answer file and a .receipt.json sidecar that records the exact model,
endpoint, token usage, and SHA-256 hashes needed to re-validate the review
without re-running the model.
Review Artifacts
| Path | Purpose |
|---|---|
shared-data/artifacts/deepseek_review/ |
Root for all review answers and receipt sidecars |
*_deepseek-v3.2_*.md |
Primary review answer (markdown, model-authored) |
*_deepseek-v3.2_*.receipt.json |
Receipt for the primary review (schema ollama_deepseek_review_receipt_v1) |
*_deepseek-v4-flash_continuation_*.md |
Continuation answer when the primary review was truncated |
*_deepseek-v4-flash_continuation_*.receipt.json |
Receipt for the continuation (schema ollama_deepseek_review_continuation_receipt_v1) |
Receipt filenames are aligned with their answer files by sharing the same
<topic>_<model>_<ISO-timestamp> stem.
Emitter Boundary
The existing prime-gap receipts were not emitted by
5-Applications/tools-scripts/llm/deepseek_review_adapter.py. That adapter
targets https://api.deepseek.com/anthropic, records schema_version,
completed_at, tokens.{input,output,cache_creation,cache_read},
response_sha256, and structured cached_context_files, and is suitable for
Anthropic-compatible DeepSeek reviews.
The artifacts in shared-data/artifacts/deepseek_review/ instead record the
Ollama-compatible schema documented below: schema, created_at, endpoint,
usage.{prompt_tokens,completion_tokens,total_tokens}, plain string
context_files, and answer_sha256. As of this page, no tracked canonical
Ollama emitter script has been identified in the repository. Treat the checked
receipts themselves as the schema authority until such an emitter is added or
the adapter is reconciled.
Receipt Schema
Both receipt schemas are versioned JSON records that track every field needed to re-derive the review independently of the model server.
| Field | Type | Notes |
|---|---|---|
schema |
string | ollama_deepseek_review_receipt_v1 (primary) or ollama_deepseek_review_continuation_receipt_v1 (continuation) |
created_at |
ISO-8601 timestamp | UTC time at which the review was produced |
model |
string | Model identifier (e.g. deepseek-v3.2, deepseek-v4-flash) |
endpoint |
URL | API endpoint that served the review (e.g. https://ollama.com/v1/chat/completions) |
prompt_sha256 |
sha256:<hex> |
SHA-256 of the exact serialized prompt body sent to the endpoint |
answer_sha256 |
sha256:<hex> |
SHA-256 of the answer payload written to answer_path |
usage.prompt_tokens |
int | Tokens consumed by the prompt as reported by the endpoint |
usage.completion_tokens |
int | Tokens produced by the model |
usage.total_tokens |
int | Sum of the two above (recorded for cross-checks) |
context_files |
string[] | Repo-relative paths to every file that participated in the review prompt context (primary receipt only; continuation consumers reconstruct this through previous_answer_path and the primary receipt) |
answer_path |
string | Repo-relative path to the answer markdown |
previous_answer_path |
string | Repo-relative path to the answer being continued (continuation receipt only) |
message_keys |
string[] | Field names returned alongside content by the continuation endpoint (e.g. role, content, reasoning) (continuation receipt only) |
Receipts are committed alongside the answers so that any future agent can:
- Verify integrity by recomputing the SHA-256 of
answer_pathand comparing toanswer_sha256. - Reconstruct the prompt by joining the listed
context_filesagainst the commit at which the receipt was authored. - Audit token usage and cost without re-querying the endpoint.
Reconstructing context for continuations
Continuation receipts intentionally omit context_files — a continuation
inherits the prompt context of the primary review it extends. Consumers should
not treat the missing field as lost context; they reconstruct it through the
previous answer and primary receipt. To reconstruct the full context for a
continuation answer:
- Read
previous_answer_pathfrom the continuation receipt. - Locate the sibling primary receipt by replacing the
.mdsuffix onprevious_answer_pathwith.receipt.json(the primary receipt and its answer share a stem). - Read
context_filesfrom that primary receipt and treat it as the continuation's effective context bundle. - The continuation prompt body itself is the primary answer at
previous_answer_pathplus any continuation directive recorded in the continuation answer's preamble;prompt_sha256on the continuation receipt covers that combined body.
Review Process
Reviews are executed as a two-stage pipeline:
- Primary analysis with
deepseek-v3.2against a curated prompt that bundles the canonical spec, statistical receipts, and Lean kernels for the topic under review. - Continuation with
deepseek-v4-flashwhen the primary review is truncated by the per-completion token cap. The continuation receipt references the primary answer viaprevious_answer_pathand inherits the review topic via filename stem.
Every review answer is structured to include:
- A YES/NO verdict table answering the specific review questions posed in the prompt.
- An arithmetic recheck that recomputes every numeric claim in the reviewed material against the canonical specification.
- A statistical interpretation section that classifies thresholds as
CANONICAL,HEURISTIC, orDETERMINISTIC WINDOW FEATURE, and flags null-model mismatches. - A failure mode analysis that enumerates the conditions under which the reviewed method would emit false positives or false negatives.
- A canonical spec patches block of explicit warnings to copy verbatim into the reviewed specification.
Example Review: Prime Gap Entropy Collapse
The canonical example shipped with the initial DeepSeek review tracking is the prime-gap entropy-collapse analysis:
| Artifact | Path |
|---|---|
| Primary answer | shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.md |
| Primary receipt | shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.receipt.json |
| Continuation answer | shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.md |
| Continuation receipt | shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.receipt.json |
Context files cited by the primary receipt:
6-Documentation/docs/distilled/ArithmeticSpec_Corrected_2026-05-11.mdshared-data/data/stack_solidification/prime_gap_k21_rerun_receipt_2026-05-11.md0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/EntropyCollapseDetector.lean
The review covers:
- Arithmetic verification of the canonical spec (crossing counts,
D₂,σ_q, Kendall SD, and exact tail probabilities atW=8). - Statistical interpretation of threshold selection — confirming
K=7as non-selective (94.57% FPR under random-permutation null) andK=21as a heuristic ~5% FPR calibration (strict>21: 3.05%, inclusive>=21: 5.43%). - Failure mode analysis for the random-permutation null model when applied to prime gaps (ties, non-uniform marginal distribution, local dependence make the FPR estimates unreliable).
- Canonical spec patches that add HEURISTIC warnings, clarify that
window-level
σ_qandD₂are deterministic features rather than estimators, document the null-model mismatch for prime gaps, and address multiple-testing and edge-effect considerations. - Continuation (deepseek-v4-flash) extending the canonical spec patches with predictive-versus-contemporaneous fusion semantics, ground-truth caveats, parameter-sensitivity reporting requirements, edge-effect handling, reproducibility requirements, and an interpretation caveat for the final verdict surface.
Adding a New Review
When emitting new review artifacts:
- Write the answer to
shared-data/artifacts/deepseek_review/<topic>_deepseek_<model>_<ISO-Z>.md. The literal_deepseek_segment is the provider tag and is held constant across all DeepSeek-family reviews;<model>is the specific model identifier (e.g.deepseek-v3.2,deepseek-v4-flash). The two segments are kept separate so future provider-level tooling (rate caps, budget accounting, fleet-wide audits) can filter on the provider tag without parsing the<model>slug. This is why the canonical example filenames contain_deepseek_deepseek-v3.2_— the doubleddeepseekis intentional and reflects<provider>_<model>. - Write the matching receipt alongside it with the same stem and
.receipt.jsonsuffix, usingollama_deepseek_review_receipt_v1for the primary review andollama_deepseek_review_continuation_receipt_v1for any continuation. Do not generate these receipts with5-Applications/tools-scripts/llm/deepseek_review_adapter.py; it emits a different Anthropic-compatible schema unless explicitly updated. - Populate
context_fileswith repo-relative paths to every file consumed by the prompt so future agents can reproduce the prompt body. Continuation receipts omitcontext_filesandmessage_keysrecords the alternate shape of the continuation response; consumers reconstruct continuation context via the primary receipt indexed byprevious_answer_path(see #Reconstructing context for continuations above). - Record
prompt_sha256andanswer_sha256for integrity verification. - Commit the answer and receipt together — the receipt is meaningless without the answer it indexes, and the answer is unverifiable without the receipt.
Related
- Build-System — pinned Python interpreter for review tooling and prompt-hash reproducibility across runs.
5-Applications/tools-scripts/llm/deepseek_review_adapter.py— Anthropic-compatible DeepSeek adapter with a different receipt schema; useful as related infrastructure, but not the emitter for the current Ollama-style artifacts.6-Documentation/docs/distilled/— canonical specs that reviews validate against (e.g.ArithmeticSpec_Corrected_2026-05-11.md).0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/— Lean kernels cross-referenced as review context.