mirror of
https://github.com/allaunthefox/Research-Stack.git
synced 2026-07-31 03:05:21 +00:00
Split the run-on sentence that read 'Continuation receipts omit context_files and message_keys records the alternate shape...' into two distinct clauses so agents following the 'Adding a New Review' checklist don't misread it as 'omit [context_files and message_keys]'. Explicitly state that continuation receipts MUST populate message_keys to record the alternate response shape.
226 lines
12 KiB
Markdown
226 lines
12 KiB
Markdown
# DeepSeek Review Process
|
||
|
||
> **Source:** [[Home|Wiki Home]] · `shared-data/artifacts/deepseek_review/`
|
||
|
||
The Research Stack incorporates DeepSeek AI models for formal mathematical
|
||
review and validation of canonical specifications, Lean kernels, and
|
||
statistical-interpretation receipts. This page documents the wiki-level view
|
||
of the review pipeline, the receipt schema, and the canonical example
|
||
(prime-gap entropy collapse) shipped under
|
||
`shared-data/artifacts/deepseek_review/`.
|
||
|
||
Important boundary: the canonical artifacts documented on this page are
|
||
Ollama-compatible review receipts. Use
|
||
`5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py` for this
|
||
schema. The tracked `5-Applications/tools-scripts/llm/deepseek_review_adapter.py`
|
||
is a separate Anthropic-compatible DeepSeek adapter with its own
|
||
`schema_version: "1.0"` receipt format.
|
||
|
||
---
|
||
|
||
## AI-Assisted Mathematical Review
|
||
|
||
DeepSeek review receipts provide reproducibility and integrity tracking for
|
||
AI-assisted mathematical validation. Each review run emits a paired
|
||
`.md` answer file and a `.receipt.json` sidecar that records the exact model,
|
||
endpoint, token usage, and SHA-256 hashes needed to re-validate the review
|
||
without re-running the model.
|
||
|
||
### Review Artifacts
|
||
|
||
| Path | Purpose |
|
||
|---|---|
|
||
| `shared-data/artifacts/deepseek_review/` | Root for all review answers and receipt sidecars |
|
||
| `*_deepseek-v3.2_*.md` | Primary review answer (markdown, model-authored) |
|
||
| `*_deepseek-v3.2_*.receipt.json` | Receipt for the primary review (schema `ollama_deepseek_review_receipt_v1`) |
|
||
| `*_deepseek-v4-flash_continuation_*.md` | Continuation answer when the primary review was truncated |
|
||
| `*_deepseek-v4-flash_continuation_*.receipt.json` | Receipt for the continuation (schema `ollama_deepseek_review_continuation_receipt_v1`) |
|
||
|
||
Receipt filenames are aligned with their answer files by sharing the same
|
||
`<topic>_<model>_<ISO-timestamp>` stem.
|
||
|
||
### Emitter Boundary
|
||
|
||
The existing prime-gap receipts were not emitted by
|
||
`5-Applications/tools-scripts/llm/deepseek_review_adapter.py`. That adapter
|
||
targets `https://api.deepseek.com/anthropic`, records `schema_version`,
|
||
`completed_at`, `tokens.{input,output,cache_creation,cache_read}`,
|
||
`response_sha256`, and structured `cached_context_files`, and is suitable for
|
||
Anthropic-compatible DeepSeek reviews.
|
||
|
||
The artifacts in `shared-data/artifacts/deepseek_review/` instead record the
|
||
Ollama-compatible schema documented below: `schema`, `created_at`, `endpoint`,
|
||
`usage.{prompt_tokens,completion_tokens,total_tokens}`, plain string
|
||
`context_files`, and `answer_sha256`. The canonical emitter is
|
||
`5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py`; it writes
|
||
the answer file first, computes `answer_sha256` from bytes read back from disk,
|
||
writes the receipt, and immediately verifies the receipt against the answer
|
||
path before returning success. This write-time verification gate exists because
|
||
the first manually landed Ollama receipts had stale `answer_sha256` values.
|
||
|
||
### Receipt Schema
|
||
|
||
Both receipt schemas are versioned JSON records that track every field needed
|
||
to re-derive the review independently of the model server.
|
||
|
||
| Field | Type | Notes |
|
||
|---|---|---|
|
||
| `schema` | string | `ollama_deepseek_review_receipt_v1` (primary) or `ollama_deepseek_review_continuation_receipt_v1` (continuation) |
|
||
| `created_at` | ISO-8601 timestamp | UTC time at which the review was produced |
|
||
| `model` | string | Model identifier (e.g. `deepseek-v3.2`, `deepseek-v4-flash`) |
|
||
| `endpoint` | URL | API endpoint that served the review (e.g. `https://ollama.com/v1/chat/completions`) |
|
||
| `prompt_sha256` | `sha256:<hex>` | SHA-256 of the exact serialized prompt body sent to the endpoint |
|
||
| `answer_sha256` | `sha256:<hex>` | SHA-256 of the answer payload written to `answer_path` |
|
||
| `usage.prompt_tokens` | int | Tokens consumed by the prompt as reported by the endpoint |
|
||
| `usage.completion_tokens` | int | Tokens produced by the model |
|
||
| `usage.total_tokens` | int | Sum of the two above (recorded for cross-checks) |
|
||
| `context_files` | string[] | Repo-relative paths to every file that participated in the review prompt context (primary receipt only; continuation consumers reconstruct this through `previous_answer_path` and the primary receipt) |
|
||
| `answer_path` | string | Repo-relative path to the answer markdown |
|
||
| `previous_answer_path` | string | Repo-relative path to the answer being continued (continuation receipt only) |
|
||
| `message_keys` | string[] | Field names returned alongside `content` by the continuation endpoint (e.g. `role`, `content`, `reasoning`) (continuation receipt only) |
|
||
|
||
Receipts are committed alongside the answers so that any future agent can:
|
||
|
||
1. Verify integrity by recomputing the SHA-256 of `answer_path` and comparing
|
||
to `answer_sha256`.
|
||
2. Reconstruct the prompt by joining the listed `context_files` against the
|
||
commit at which the receipt was authored.
|
||
3. Audit token usage and cost without re-querying the endpoint.
|
||
|
||
### Reconstructing context for continuations
|
||
|
||
Continuation receipts intentionally omit `context_files` — a continuation
|
||
inherits the prompt context of the primary review it extends. Consumers should
|
||
not treat the missing field as lost context; they reconstruct it through the
|
||
previous answer and primary receipt. To reconstruct the full context for a
|
||
continuation answer:
|
||
|
||
1. Read `previous_answer_path` from the continuation receipt.
|
||
2. Locate the sibling primary receipt by replacing the `.md` suffix on
|
||
`previous_answer_path` with `.receipt.json` (the primary receipt and its
|
||
answer share a stem).
|
||
3. Read `context_files` from that primary receipt and treat it as the
|
||
continuation's effective context bundle.
|
||
4. The continuation prompt body itself is the primary answer at
|
||
`previous_answer_path` plus any continuation directive recorded in the
|
||
continuation answer's preamble; `prompt_sha256` on the continuation
|
||
receipt covers that combined body.
|
||
|
||
### Review Process
|
||
|
||
Reviews are executed as a two-stage pipeline:
|
||
|
||
1. **Primary analysis** with `deepseek-v3.2` against a curated prompt that
|
||
bundles the canonical spec, statistical receipts, and Lean kernels for the
|
||
topic under review.
|
||
2. **Continuation** with `deepseek-v4-flash` when the primary review is
|
||
truncated by the per-completion token cap. The continuation receipt
|
||
references the primary answer via `previous_answer_path` and inherits the
|
||
review topic via filename stem.
|
||
|
||
Every review answer is structured to include:
|
||
|
||
- A **YES/NO verdict table** answering the specific review questions posed
|
||
in the prompt.
|
||
- An **arithmetic recheck** that recomputes every numeric claim in the
|
||
reviewed material against the canonical specification.
|
||
- A **statistical interpretation** section that classifies thresholds as
|
||
`CANONICAL`, `HEURISTIC`, or `DETERMINISTIC WINDOW FEATURE`, and flags
|
||
null-model mismatches.
|
||
- A **failure mode analysis** that enumerates the conditions under which the
|
||
reviewed method would emit false positives or false negatives.
|
||
- A **canonical spec patches** block of explicit warnings to copy verbatim
|
||
into the reviewed specification.
|
||
|
||
### Example Review: Prime Gap Entropy Collapse
|
||
|
||
The canonical example shipped with the initial DeepSeek review tracking is
|
||
the prime-gap entropy-collapse analysis:
|
||
|
||
| Artifact | Path |
|
||
|---|---|
|
||
| Primary answer | `shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.md` |
|
||
| Primary receipt | `shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v3.2_20260512T033551Z.receipt.json` |
|
||
| Continuation answer | `shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.md` |
|
||
| Continuation receipt | `shared-data/artifacts/deepseek_review/prime_gap_entropy_collapse_deepseek_deepseek-v4-flash_continuation_20260512T033849Z.receipt.json` |
|
||
|
||
Context files cited by the primary receipt:
|
||
|
||
- `6-Documentation/docs/distilled/ArithmeticSpec_Corrected_2026-05-11.md`
|
||
- `shared-data/data/stack_solidification/prime_gap_k21_rerun_receipt_2026-05-11.md`
|
||
- `0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/EntropyCollapseDetector.lean`
|
||
|
||
The review covers:
|
||
|
||
- **Arithmetic verification** of the canonical spec (crossing counts, `D₂`,
|
||
`σ_q`, Kendall SD, and exact tail probabilities at `W=8`).
|
||
- **Statistical interpretation** of threshold selection — confirming `K=7`
|
||
as non-selective (94.57% FPR under random-permutation null) and `K=21` as a
|
||
heuristic ~5% FPR calibration (strict `>21`: 3.05%, inclusive `>=21`: 5.43%).
|
||
- **Failure mode analysis** for the random-permutation null model when
|
||
applied to prime gaps (ties, non-uniform marginal distribution, local
|
||
dependence make the FPR estimates unreliable).
|
||
- **Canonical spec patches** that add HEURISTIC warnings, clarify that
|
||
window-level `σ_q` and `D₂` are deterministic features rather than
|
||
estimators, document the null-model mismatch for prime gaps, and address
|
||
multiple-testing and edge-effect considerations.
|
||
- **Continuation** (deepseek-v4-flash) extending the canonical spec patches
|
||
with predictive-versus-contemporaneous fusion semantics, ground-truth
|
||
caveats, parameter-sensitivity reporting requirements, edge-effect handling,
|
||
reproducibility requirements, and an interpretation caveat for the final
|
||
verdict surface.
|
||
|
||
---
|
||
|
||
## Adding a New Review
|
||
|
||
When emitting new review artifacts:
|
||
|
||
1. Write the answer to
|
||
`shared-data/artifacts/deepseek_review/<topic>_deepseek_<model>_<ISO-Z>.md`.
|
||
The literal `_deepseek_` segment is the **provider tag** and is held
|
||
constant across all DeepSeek-family reviews; `<model>` is the specific
|
||
model identifier (e.g. `deepseek-v3.2`, `deepseek-v4-flash`). The two
|
||
segments are kept separate so future provider-level tooling (rate caps,
|
||
budget accounting, fleet-wide audits) can filter on the provider tag
|
||
without parsing the `<model>` slug. This is why the canonical example
|
||
filenames contain `_deepseek_deepseek-v3.2_` — the doubled `deepseek` is
|
||
intentional and reflects `<provider>_<model>`.
|
||
2. Write the matching receipt alongside it with the same stem and
|
||
`.receipt.json` suffix, using `ollama_deepseek_review_receipt_v1` for the
|
||
primary review and `ollama_deepseek_review_continuation_receipt_v1` for any
|
||
continuation. Use
|
||
`5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py`; do not
|
||
generate these receipts with
|
||
`5-Applications/tools-scripts/llm/deepseek_review_adapter.py`, which emits a
|
||
different Anthropic-compatible schema.
|
||
3. Populate `context_files` with repo-relative paths to every file consumed
|
||
by the prompt so future agents can reproduce the prompt body. Continuation
|
||
receipts omit `context_files`. Continuation receipts MUST populate
|
||
`message_keys` to record the alternate shape of the continuation response
|
||
(e.g. `["role", "content", "reasoning"]`). Consumers reconstruct
|
||
continuation context via the primary receipt indexed by
|
||
`previous_answer_path` (see
|
||
[[#Reconstructing context for continuations]] above).
|
||
4. Record `prompt_sha256` and `answer_sha256` for integrity verification.
|
||
5. Commit the answer and receipt together — the receipt is meaningless
|
||
without the answer it indexes, and the answer is unverifiable without the
|
||
receipt.
|
||
|
||
---
|
||
|
||
## Related
|
||
|
||
- [[Build-System]] — pinned Python interpreter for review tooling and
|
||
prompt-hash reproducibility across runs.
|
||
- `5-Applications/tools-scripts/llm/ollama_deepseek_review_emitter.py` —
|
||
canonical Ollama-compatible emitter for this schema; verifies
|
||
`answer_sha256` after writing.
|
||
- `5-Applications/tools-scripts/llm/deepseek_review_adapter.py` —
|
||
Anthropic-compatible DeepSeek adapter with a different receipt schema; useful
|
||
as related infrastructure, but not the emitter for the current Ollama-style
|
||
artifacts.
|
||
- `6-Documentation/docs/distilled/` — canonical specs that reviews validate
|
||
against (e.g. `ArithmeticSpec_Corrected_2026-05-11.md`).
|
||
- `0-Core-Formalism/lean/Semantics/Semantics/HCMMR/Kernels/` — Lean kernels
|
||
cross-referenced as review context.
|