Research-Stack/6-Documentation/wiki/DeepSeek-Review-Process.md
devin-ai-integration[bot] 1187b4ab44 docs(wiki): disambiguate continuation receipt instructions
Split the run-on sentence that read 'Continuation receipts omit
context_files and message_keys records the alternate shape...' into
two distinct clauses so agents following the 'Adding a New Review'
checklist don't misread it as 'omit [context_files and message_keys]'.
Explicitly state that continuation receipts MUST populate message_keys
to record the alternate response shape.
2026-05-17 12:20:57 -05:00

12 KiB
Raw Permalink Blame History

DeepSeek Review Process

Source: Home · shared-data/artifacts/deepseek_review/

The Research Stack incorporates DeepSeek AI models for formal mathematical review and validation of canonical specifications, Lean kernels, and statistical-interpretation receipts. This page documents the wiki-level view of the review pipeline, the receipt schema, and the canonical example (prime-gap entropy collapse) shipped under shared-data/artifacts/deepseek_review/.

Important boundary: the canonical artifacts documented on this page are Ollama-compatible review receipts. Use 5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py for this schema. The tracked 5-Applications/tools-scripts/llm/deepseek_review_adapter.py is a separate Anthropic-compatible DeepSeek adapter with its own schema_version: "1.0" receipt format.


AI-Assisted Mathematical Review

DeepSeek review receipts provide reproducibility and integrity tracking for AI-assisted mathematical validation. Each review run emits a paired .md answer file and a .receipt.json sidecar that records the exact model, endpoint, token usage, and SHA-256 hashes needed to re-validate the review without re-running the model.

Review Artifacts

Path Purpose
shared-data/artifacts/deepseek_review/ Root for all review answers and receipt sidecars
*_deepseek-v3.2_*.md Primary review answer (markdown, model-authored)
*_deepseek-v3.2_*.receipt.json Receipt for the primary review (schema ollama_deepseek_review_receipt_v1)
*_deepseek-v4-flash_continuation_*.md Continuation answer when the primary review was truncated
*_deepseek-v4-flash_continuation_*.receipt.json Receipt for the continuation (schema ollama_deepseek_review_continuation_receipt_v1)

Receipt filenames are aligned with their answer files by sharing the same <topic>_<model>_<ISO-timestamp> stem.

Emitter Boundary

The existing prime-gap receipts were not emitted by 5-Applications/tools-scripts/llm/deepseek_review_adapter.py. That adapter targets https://api.deepseek.com/anthropic, records schema_version, completed_at, tokens.{input,output,cache_creation,cache_read}, response_sha256, and structured cached_context_files, and is suitable for Anthropic-compatible DeepSeek reviews.

The artifacts in shared-data/artifacts/deepseek_review/ instead record the Ollama-compatible schema documented below: schema, created_at, endpoint, usage.{prompt_tokens,completion_tokens,total_tokens}, plain string context_files, and answer_sha256. The canonical emitter is 5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py; it writes the answer file first, computes answer_sha256 from bytes read back from disk, writes the receipt, and immediately verifies the receipt against the answer path before returning success. This write-time verification gate exists because the first manually landed Ollama receipts had stale answer_sha256 values.

Receipt Schema

Both receipt schemas are versioned JSON records that track every field needed to re-derive the review independently of the model server.

Field Type Notes
schema string ollama_deepseek_review_receipt_v1 (primary) or ollama_deepseek_review_continuation_receipt_v1 (continuation)
created_at ISO-8601 timestamp UTC time at which the review was produced
model string Model identifier (e.g. deepseek-v3.2, deepseek-v4-flash)
endpoint URL API endpoint that served the review (e.g. https://ollama.com/v1/chat/completions)
prompt_sha256 sha256:<hex> SHA-256 of the exact serialized prompt body sent to the endpoint
answer_sha256 sha256:<hex> SHA-256 of the answer payload written to answer_path
usage.prompt_tokens int Tokens consumed by the prompt as reported by the endpoint
usage.completion_tokens int Tokens produced by the model
usage.total_tokens int Sum of the two above (recorded for cross-checks)
context_files string[] Repo-relative paths to every file that participated in the review prompt context (primary receipt only; continuation consumers reconstruct this through previous_answer_path and the primary receipt)
answer_path string Repo-relative path to the answer markdown
previous_answer_path string Repo-relative path to the answer being continued (continuation receipt only)
message_keys string[] Field names returned alongside content by the continuation endpoint (e.g. role, content, reasoning) (continuation receipt only)

Receipts are committed alongside the answers so that any future agent can:

  1. Verify integrity by recomputing the SHA-256 of answer_path and comparing to answer_sha256.
  2. Reconstruct the prompt by joining the listed context_files against the commit at which the receipt was authored.
  3. Audit token usage and cost without re-querying the endpoint.

Reconstructing context for continuations

Continuation receipts intentionally omit context_files — a continuation inherits the prompt context of the primary review it extends. Consumers should not treat the missing field as lost context; they reconstruct it through the previous answer and primary receipt. To reconstruct the full context for a continuation answer:

  1. Read previous_answer_path from the continuation receipt.
  2. Locate the sibling primary receipt by replacing the .md suffix on previous_answer_path with .receipt.json (the primary receipt and its answer share a stem).
  3. Read context_files from that primary receipt and treat it as the continuation's effective context bundle.
  4. The continuation prompt body itself is the primary answer at previous_answer_path plus any continuation directive recorded in the continuation answer's preamble; prompt_sha256 on the continuation receipt covers that combined body.

Review Process

Reviews are executed as a two-stage pipeline:

  1. Primary analysis with deepseek-v3.2 against a curated prompt that bundles the canonical spec, statistical receipts, and Lean kernels for the topic under review.
  2. Continuation with deepseek-v4-flash when the primary review is truncated by the per-completion token cap. The continuation receipt references the primary answer via previous_answer_path and inherits the review topic via filename stem.

Every review answer is structured to include:

  • A YES/NO verdict table answering the specific review questions posed in the prompt.
  • An arithmetic recheck that recomputes every numeric claim in the reviewed material against the canonical specification.
  • A statistical interpretation section that classifies thresholds as CANONICAL, HEURISTIC, or DETERMINISTIC WINDOW FEATURE, and flags null-model mismatches.
  • A failure mode analysis that enumerates the conditions under which the reviewed method would emit false positives or false negatives.
  • A canonical spec patches block of explicit warnings to copy verbatim into the reviewed specification.

Example Review: Prime Gap Entropy Collapse

The canonical example shipped with the initial DeepSeek review tracking is the prime-gap entropy-collapse analysis:

Artifact Path
Primary answer shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.md
Primary receipt shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.receipt.json
Continuation answer shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.md
Continuation receipt shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.receipt.json

Context files cited by the primary receipt:

  • 6-Documentation/docs/distilled/ArithmeticSpec_Corrected_2026-05-11.md
  • shared-data/data/stack_solidification/prime_gap_k21_rerun_receipt_2026-05-11.md
  • 0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/EntropyCollapseDetector.lean

The review covers:

  • Arithmetic verification of the canonical spec (crossing counts, D₂, σ_q, Kendall SD, and exact tail probabilities at W=8).
  • Statistical interpretation of threshold selection — confirming K=7 as non-selective (94.57% FPR under random-permutation null) and K=21 as a heuristic ~5% FPR calibration (strict >21: 3.05%, inclusive >=21: 5.43%).
  • Failure mode analysis for the random-permutation null model when applied to prime gaps (ties, non-uniform marginal distribution, local dependence make the FPR estimates unreliable).
  • Canonical spec patches that add HEURISTIC warnings, clarify that window-level σ_q and D₂ are deterministic features rather than estimators, document the null-model mismatch for prime gaps, and address multiple-testing and edge-effect considerations.
  • Continuation (deepseek-v4-flash) extending the canonical spec patches with predictive-versus-contemporaneous fusion semantics, ground-truth caveats, parameter-sensitivity reporting requirements, edge-effect handling, reproducibility requirements, and an interpretation caveat for the final verdict surface.

Adding a New Review

When emitting new review artifacts:

  1. Write the answer to shared-data/artifacts/deepseek_review/<topic>_deepseek_<model>_<ISO-Z>.md. The literal _deepseek_ segment is the provider tag and is held constant across all DeepSeek-family reviews; <model> is the specific model identifier (e.g. deepseek-v3.2, deepseek-v4-flash). The two segments are kept separate so future provider-level tooling (rate caps, budget accounting, fleet-wide audits) can filter on the provider tag without parsing the <model> slug. This is why the canonical example filenames contain _deepseek_deepseek-v3.2_ — the doubled deepseek is intentional and reflects <provider>_<model>.
  2. Write the matching receipt alongside it with the same stem and .receipt.json suffix, using ollama_deepseek_review_receipt_v1 for the primary review and ollama_deepseek_review_continuation_receipt_v1 for any continuation. Use 5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py; do not generate these receipts with 5-Applications/tools-scripts/llm/deepseek_review_adapter.py, which emits a different Anthropic-compatible schema.
  3. Populate context_files with repo-relative paths to every file consumed by the prompt so future agents can reproduce the prompt body. Continuation receipts omit context_files. Continuation receipts MUST populate message_keys to record the alternate shape of the continuation response (e.g. ["role", "content", "reasoning"]). Consumers reconstruct continuation context via the primary receipt indexed by previous_answer_path (see #Reconstructing context for continuations above).
  4. Record prompt_sha256 and answer_sha256 for integrity verification.
  5. Commit the answer and receipt together — the receipt is meaningless without the answer it indexes, and the answer is unverifiable without the receipt.

  • Build-System — pinned Python interpreter for review tooling and prompt-hash reproducibility across runs.
  • 5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py — canonical Ollama-compatible emitter for this schema; verifies answer_sha256 after writing.
  • 5-Applications/tools-scripts/llm/deepseek_review_adapter.py — Anthropic-compatible DeepSeek adapter with a different receipt schema; useful as related infrastructure, but not the emitter for the current Ollama-style artifacts.
  • 6-Documentation/docs/distilled/ — canonical specs that reviews validate against (e.g. ArithmeticSpec_Corrected_2026-05-11.md).
  • 0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/ — Lean kernels cross-referenced as review context.