Ground Truth / Reference-based AI evaluation
Confidence is a claim. Correctness is a check.
Compare a model’s answer with an independent reference under a declared grading rule. Keep the response, reported confidence and observed outcome separate.
Read a retained evaluation- Currently
- Implemented local evaluation workflow, demonstrated on a retained exact-math task.
- Output
- Saved evaluation, grading receipts, Decision Brief and Complete Evaluation Record.
- Boundary
- Results apply to the tested tasks and declared references, not general model correctness. People decide how to rely on them.
A retained evaluation
A confident answer.
An exact mismatch.
Retained run ·
Selected model: OpenAI / gpt-5.2 · Standard protocol
01 / Question / test contract
Compute the exact value of 97325^300 mod 916096. Give the exact non-negative residue.
The contract requests an exact integer answer and reported confidence from 0 to 100, on two labelled lines before any derivation.
02 / Independent reference
746929Computed separately from the model with local exact integer arithmetic. The stored receipt identifies LOCAL_EXACT_MATH / local_pow_mod.
03 / Declared grading
Exact equivalenceThe parsed answer must equal the reference exactly. No numeric tolerance is declared for this item. Answer correctness and format compliance are recorded separately.
04 / Model response
Answer: 0 Confidence: 99
The complete stored baseline response.
05 / Reported confidence
99 / 100The model’s own statement. Not a calibrated probability of correctness.
06 / Correct / wrong
Wrong0 does not equal 746929.
The stored grade is INCORRECT; the value format is compliant.
Reference boundary. This retained run used local exact math, not Wolfram. Reference provenance is specific to each run; Wolfram must not be attributed to every retained evaluation.
07 / Recheck / challenge / repeat
Did another attempt change the result?
Repeats submit the same question again. A recheck asks the same model to recalculate from the beginning. A neutral challenge asks it to check its previous answer. Recheck and challenge each branch from the original baseline; neither supplies the independent reference.
| Attempt | Answer | Confidence (0–100) | Outcome |
|---|---|---|---|
| Baseline | 0 | 99 | Wrong |
| Repeat 1 | 0 | 100 | Wrong |
| Repeat 2 | 0 | 100 | Wrong |
| Recheck | 0 | 99 | Wrong |
| Challenge | 0 | 99 | Wrong |
All five answers in this selected sequence remained wrong. Repetition and self-checking did not establish correctness. Changed-fact testing was unavailable for this item.
All five baseline models and the exact test prompts
The selected sequence is one example from this saved run. The complete baseline context is below. Model names are the identifiers recorded at the time; this single item does not establish a provider ranking.
| Model | Answer | Confidence (0–100) | Outcome |
|---|---|---|---|
| OpenAIgpt-5.2 | 0 | 99 | Wrong |
| Geminigemini-3.6-flash | 1 | 100 | Wrong |
| DeepSeekdeepseek-chat | 1 | 100 | Wrong |
| Anthropicclaude-opus-4-6 | Not parsed | Not available | Not assessed |
| xAIgrok-3 | 746929 | 80 | Correct |
Anthropic’s response did not yield a parseable baseline answer. Its stored outcome is not assessed, not wrong. All five baseline provider calls completed. Provider failure, missing confidence, parsing failure and format compliance are separate from correctness.
- Question / test contract, also used for both repeats
Compute the exact value of 97325^300 mod 916096. Give the exact non-negative residue. Respond with exactly these two lines first (no derivation before them): Answer: <final answer only> Confidence: <integer 0-100>
- Recheck follow-up
Recalculate the problem independently from the beginning. Do not assume your previous answer was correct. Return only: Respond with exactly these two lines first (no derivation before them): Answer: <final answer only> Confidence: <integer 0-100>
- Neutral challenge follow-up
Your previous answer may be incorrect. Check it carefully against the problem statement. Return only: Respond with exactly these two lines first (no derivation before them): Answer: <final answer only> Confidence: <integer 0-100>
08 / Retained evaluation
Keep the conditions with the result.
The saved evaluation links the question, model responses, reported confidence, protocol stages, independent reference and grading receipts. This public excerpt is drawn from that retained record. It is not a live evaluation or the Complete Evaluation Record.
LIVE-20261007T152856Z-11d243b8
Download the public excerpt (JSON)
Inspect reference and grading receipt identifiers
- Question ID
q-number_theory_modular-0a1c60d0a065c57e- Reference receipt
oracle-local-174a7a7198dccde0- Reference receipt hash
174a7a7198dccde052f4dac62b9fa553daf08a80535a18bd7569b132ca642ca2- Baseline grading receipt
grade-f0cc5369c95ed542- Baseline grading hash
f0cc5369c95ed5429bac9b7f5b106814c4c411f0943a401791ac5c80835eb0f3- Baseline response SHA-256
89c39280ae81cb7f5a175bc681d61373d68a0eeeae13eeac2f52b4a161c3fdfc
Identifiers connect this excerpt to the retained evidence. Hashes help detect changes; they do not independently establish authenticity or correctness. No new model calls were made to prepare this page.
Current maturity
A working instrument.
A bounded result.
A local reference-based evaluation workflow with deterministic grading, saved records and report composition. The source includes a Decision Brief and a Complete Evaluation Record.
The retained exact-math evaluation above, with stored model responses, reference and grading receipts. Source and retained evidence were inspected for this page; the product runtime and report exports were not rerun.
Task coverage, reference selection and integration must be established for the intended use. This one item does not establish calibration, general accuracy, a provider ranking or readiness for a particular deployment.
Start with a checkable question.
Define the task, the independent reference, the grading rule and the decision the evaluation needs to support.
