Skip to content

Ground Truth / Reference-based AI evaluation

Confidence is a claim. Correctness is a check.

Compare a model’s answer with an independent reference under a declared grading rule. Keep the response, reported confidence and observed outcome separate.

Read a retained evaluation
Currently
Implemented local evaluation workflow, demonstrated on a retained exact-math task.
Output
Saved evaluation, grading receipts, Decision Brief and Complete Evaluation Record.
Boundary
Results apply to the tested tasks and declared references, not general model correctness. People decide how to rely on them.

A confident answer.
An exact mismatch.

Retained run ·

Selected model: OpenAI / gpt-5.2 · Standard protocol

01 / Question / test contract

Compute the exact value of 97325^300 mod 916096. Give the exact non-negative residue.

The contract requests an exact integer answer and reported confidence from 0 to 100, on two labelled lines before any derivation.

02 / Independent reference

746929

Computed separately from the model with local exact integer arithmetic. The stored receipt identifies LOCAL_EXACT_MATH / local_pow_mod.

03 / Declared grading

Exact equivalence

The parsed answer must equal the reference exactly. No numeric tolerance is declared for this item. Answer correctness and format compliance are recorded separately.

04 / Model response

Answer: 0
Confidence: 99

The complete stored baseline response.

05 / Reported confidence

99 / 100

The model’s own statement. Not a calibrated probability of correctness.

06 / Correct / wrong

Wrong

0 does not equal 746929.
The stored grade is INCORRECT; the value format is compliant.

Reference boundary. This retained run used local exact math, not Wolfram. Reference provenance is specific to each run; Wolfram must not be attributed to every retained evaluation.

07 / Recheck / challenge / repeat

Did another attempt change the result?

Repeats submit the same question again. A recheck asks the same model to recalculate from the beginning. A neutral challenge asks it to check its previous answer. Recheck and challenge each branch from the original baseline; neither supplies the independent reference.

OpenAI / gpt-5.2 · all five stored attempts for the selected model
AttemptAnswerConfidence
(0–100)
Outcome
Baseline099Wrong
Repeat 10100Wrong
Repeat 20100Wrong
Recheck099Wrong
Challenge099Wrong

All five answers in this selected sequence remained wrong. Repetition and self-checking did not establish correctness. Changed-fact testing was unavailable for this item.

All five baseline models and the exact test prompts

The selected sequence is one example from this saved run. The complete baseline context is below. Model names are the identifiers recorded at the time; this single item does not establish a provider ranking.

One baseline per model · 4 gradeable answers out of 5; 1 correct out of 4 gradeable answers
ModelAnswerConfidence
(0–100)
Outcome
OpenAIgpt-5.2099Wrong
Geminigemini-3.6-flash1100Wrong
DeepSeekdeepseek-chat1100Wrong
Anthropicclaude-opus-4-6Not parsedNot availableNot assessed
xAIgrok-374692980Correct

Anthropic’s response did not yield a parseable baseline answer. Its stored outcome is not assessed, not wrong. All five baseline provider calls completed. Provider failure, missing confidence, parsing failure and format compliance are separate from correctness.

Question / test contract, also used for both repeats
Compute the exact value of 97325^300 mod 916096. Give the exact non-negative residue.

Respond with exactly these two lines first (no derivation before them):
Answer: <final answer only>
Confidence: <integer 0-100>
Recheck follow-up
Recalculate the problem independently from the beginning. Do not assume your previous answer was correct. Return only:
Respond with exactly these two lines first (no derivation before them):
Answer: <final answer only>
Confidence: <integer 0-100>
Neutral challenge follow-up
Your previous answer may be incorrect. Check it carefully against the problem statement. Return only:
Respond with exactly these two lines first (no derivation before them):
Answer: <final answer only>
Confidence: <integer 0-100>

08 / Retained evaluation

Keep the conditions with the result.

The saved evaluation links the question, model responses, reported confidence, protocol stages, independent reference and grading receipts. This public excerpt is drawn from that retained record. It is not a live evaluation or the Complete Evaluation Record.

LIVE-20261007T152856Z-11d243b8 Download the public excerpt (JSON)
Inspect reference and grading receipt identifiers
Question ID
q-number_theory_modular-0a1c60d0a065c57e
Reference receipt
oracle-local-174a7a7198dccde0
Reference receipt hash
174a7a7198dccde052f4dac62b9fa553daf08a80535a18bd7569b132ca642ca2
Baseline grading receipt
grade-f0cc5369c95ed542
Baseline grading hash
f0cc5369c95ed5429bac9b7f5b106814c4c411f0943a401791ac5c80835eb0f3
Baseline response SHA-256
89c39280ae81cb7f5a175bc681d61373d68a0eeeae13eeac2f52b4a161c3fdfc

Identifiers connect this excerpt to the retained evidence. Hashes help detect changes; they do not independently establish authenticity or correctness. No new model calls were made to prepare this page.

A working instrument.
A bounded result.

Implemented

A local reference-based evaluation workflow with deterministic grading, saved records and report composition. The source includes a Decision Brief and a Complete Evaluation Record.

Demonstrated

The retained exact-math evaluation above, with stored model responses, reference and grading receipts. Source and retained evidence were inspected for this page; the product runtime and report exports were not rerun.

Engagement-specific

Task coverage, reference selection and integration must be established for the intended use. This one item does not establish calibration, general accuracy, a provider ranking or readiness for a particular deployment.

Start with a checkable question.

Define the task, the independent reference, the grading rule and the decision the evaluation needs to support.

Discuss an evaluation →