A prompt, model or retrieval change can improve a demo while quietly breaking cases that matter in production.
Evaluation & Observability
OKAI Solutions
Iris AI
bySee what changed before you release.
An evaluation workbench for comparing prompt, model, retrieval and tool changes against a team’s own versioned test cases.
Request a technical walkthrough →Current92.1Candidate94.6+2.5
- Grounded answer+4.2
- Tool recovery+1.8
- Mandatory regressions0
- P95 latency+120 ms
- Used by
- AI engineering teams, QA engineers and technical product owners
- Inputs
- Versioned evaluation cases, reference answers, traces and candidate system configurations
- Produces
- A reproducible comparison report, failure clusters, cost and latency signals and a reviewable release decision
What happens when Iris AI is in the loop.
Iris runs both versions against the same dataset, groups failures and links every score back to the underlying trace.
The release owner sees what improved, what regressed and whether the candidate meets the agreed gate.
Inside Iris AI
Six visible stages connect the source to the operator. Nothing important disappears inside a single model call.
- 01Versioned cases
- 02Candidate configuration
- 03Evaluation runner
- 04Calibrated review
- 05Comparison report
- 06Release gate
Acceptance targets
These are validation criteria for a scoped pilot. They are not presented as observed customer results.
- 01
Every run records dataset and system versions
- 02
Every seeded mandatory regression blocks release
- 03
Every failure links to its source trace
- 04
Latency and model usage are reported for every included request
Start with the workflow
Contact OKAI →