Stability
Does the system behave consistently when the same or closely related task is repeated?
Diagnostics / AI evaluation
Repeated runs, controlled conditions and independently verifiable references make instability, divergence and support gaps visible before people rely on an AI system.
Evaluate a systemSelect a task to compare the authored sample answers.
(100 − 80) ÷ 80 × 100 = 25%
Sample 02 reports 30%, but the fixed reference is 25%. Repetition alone would not establish correctness.
| Sample | Answer | Reference check |
|---|---|---|
| 01 | 25 % | Matches reference |
| 02 | 30 % | Mismatch |
| 03 | 25 % | Matches reference |
| 04 | 25 % | Matches reference |
Diagnostics looks beyond the final response and examines how an AI system behaves across controlled tests, repeated runs and model comparisons.
Does the system behave consistently when the same or closely related task is repeated?
Are important claims adequately supported, or does confidence exceed the available evidence?
How differently do models, versions or configurations respond to the same controlled task?
Does the system remain within the intended scope, or expand beyond what the task and evidence justify?
Ground truth
Use exact calculations, labels, deterministic conditions or externally established reference where available.
Observe what the model actually returns across repetitions, versions and configurations.
Measure mismatch, variance, unsupported confidence and other signals without pretending one metric explains everything.
Diagnostic flow
Current maturity
Controlled prompt runs, repeated evaluation, multi-model comparison, ground-truth checks where available and reviewable diagnostic records.
Larger evaluation suites, stronger test management, richer scoring and reusable organizational benchmarks.
A continuous AI evaluation environment that makes behavioural change visible before it becomes operational risk.
Current scope: a diagnostic research implementation and controlled evaluation workflows. Results are bounded by the tasks, references and configurations actually tested.
How we scope, test and deploy a systemWhere it fits
Diagnostics is most useful when adoption decisions require more than a polished demo or a single benchmark score.
Test behaviour against the organization’s actual tasks before deployment.
Compare models and versions without reducing the decision to one generic leaderboard.
Detect regressions, instability and support gaps across releases or configuration changes.
Create a repeatable evidence record around consequential AI evaluation.
Start with the task, the behaviour that matters, the conditions under which it will be used and any independently verifiable reference available.