Skip to content

Diagnostics / AI evaluation

Evaluate AI behaviour before relying on it.

Repeated runs, controlled conditions and independently verifiable references make instability, divergence and support gaps visible before people rely on an AI system.

Evaluate a system
Interactive explainerFictional records · not live product output

Select a task to compare the authored sample answers.

0102030Reference: 25 %25%Sample 0130%Sample 0225%Sample 0325%Sample 04
Known-answer task

Output changes from 80 to 100 units. What is the relative percentage increase?

(100 − 80) ÷ 80 × 100 = 25%

Reference matches / authored sample set3 / 4

Sample 02 reports 30%, but the fixed reference is 25%. Repetition alone would not establish correctness.

Authored samples compared with the calculated reference. Not measured model performance.
SampleAnswerReference check
0125 %Matches reference
0230 %Mismatch
0325 %Matches reference
0425 %Matches reference
Measurement boundaryThe answers are authored fixtures, not model runs. This page calculates the reference and comparisons locally; the results are not an IQAI benchmark or accuracy claim.
Inputs, checks & review limits
Test definition
Specify the task, source packet, model or version, configuration, repetitions and reference. An independent reference may be a calculation, a validated label or another established constraint.
Observation, not mind-reading
Evaluate observable outputs and their changes. An explanation produced by a model is itself an output. It does not provide privileged access to the process that generated it.
Repetition & reproducibility
Retain conditions and outputs so the method can be inspected and repeated. Probabilistic generation and changing provider versions can prevent identical reruns. Read the technical paper (PDF).

A fluent answer can still hide unstable behaviour.

Diagnostics looks beyond the final response and examines how an AI system behaves across controlled tests, repeated runs and model comparisons.

/01

Stability

Does the system behave consistently when the same or closely related task is repeated?

/02

Support

Are important claims adequately supported, or does confidence exceed the available evidence?

/03

Divergence

How differently do models, versions or configurations respond to the same controlled task?

/04

Boundaries

Does the system remain within the intended scope, or expand beyond what the task and evidence justify?

When the answer is knowable, test against the answer, not another model’s opinion.

REFERENCEKnown answer

Use exact calculations, labels, deterministic conditions or externally established reference where available.

BEHAVIOURModel output

Observe what the model actually returns across repetitions, versions and configurations.

DIAGNOSTICGap

Measure mismatch, variance, unsupported confidence and other signals without pretending one metric explains everything.

Expressed confidence is observable behaviour.Correctness requires an independent reference when one exists.

From a test case to a reviewable diagnostic record.

Explore the 5-step workflow
/01DefineSpecify the task, conditions, expected constraints and reference where available.
/02RunExecute repeated tests under controlled conditions.
/03MeasureCapture output, support, confidence and other observable signals.
/04CompareCompare runs, models, versions or configurations without hiding divergence.
/05RecordRetain conditions and outputs so the method can be inspected and repeated; identical model outputs are not guaranteed.

Evaluation workflows today.

Working nowCurrent

Controlled prompt runs, repeated evaluation, multi-model comparison, ground-truth checks where available and reviewable diagnostic records.

Product development & longer-term direction
Product developmentExpanding

Larger evaluation suites, stronger test management, richer scoring and reusable organizational benchmarks.

Longer-termDirection

A continuous AI evaluation environment that makes behavioural change visible before it becomes operational risk.

Current scope: a diagnostic research implementation and controlled evaluation workflows. Results are bounded by the tasks, references and configurations actually tested.

How we scope, test and deploy a system

For teams deciding whether and how an AI system should be relied on.

Diagnostics is most useful when adoption decisions require more than a polished demo or a single benchmark score.

AI adoption

Test behaviour against the organization’s actual tasks before deployment.

Model selection

Compare models and versions without reducing the decision to one generic leaderboard.

Quality assurance

Detect regressions, instability and support gaps across releases or configuration changes.

Governance

Create a repeatable evidence record around consequential AI evaluation.

Evaluate an AI system.

Start with the task, the behaviour that matters, the conditions under which it will be used and any independently verifiable reference available.

Contact IQAI →