EvalGlass

EvalGlass Intelligence · causal failure intelligence

Understand why it failed, and make the evaluation system improve.

Core runs the evaluations you define; Discovery finds the ones you are missing. Intelligence is the advanced reasoning system for the hard failure — it determines what actually caused it, tests the cause with controlled replay and intervention, verifies the repair, and turns the learning into a durable evaluation you keep.

Investigate a hard failure private preview not part of the Core release

Cause, not coincidence

A plausible explanation is not a cause.

Discovery can tell you a retry path is untested; Core can tell you a metric moved. Intelligence begins after that — once you have an executable suite and a plausible hypothesis, and you need to know why a difficult, intermittent or consequential failure actually happens. It reasons about cause, and it refuses to compress that into a single unexplained number:

Graded evidence

Association is not causation — every conclusion states its level.

A conclusion is never presented as more certain than the evidence supports. Each causal claim carries an explicit evidence level, and a graph edge moves up the ladder only as evidence improves:

Uncertainty and failed interventions are valuable outputs, not failures: a non-discriminating experiment narrows the hypothesis space and is kept.

The method

Diagnose. Intervene. Verify. Learn.

Diagnose & intervene

Build the causal execution graph, form hypotheses, and test candidate causes with controlled replay and intervention — reporting the strongest supported evidence level, never a story.

Verify the repair

Minimize the incident to a reproducer, and once you approve a change, verify the failure is gone with no nearby regressions — a repair-validated cause is the strongest evidence of all.

Learn

Feed the Evaluation Learning Loop: every incident and change becomes a proposed new or stronger evaluation you promote — so the suite gets better after every important failure. Everything durable runs in Core.

The deep layer — without false confidence.

Intelligence improves your evaluations, not your production system autonomously; humans confirm the cause and approve every change, and no proposal enters the active suite by itself. It never calls your AI safe — only what the evidence supports, at a stated level. It exports to your Jira, incident, GRC and audit systems but does not become those systems. The work is bounded and customer-controlled — your application, your evidence window, your environment — with no broad standing access, and every durable artifact stays runnable in open-source Core. Investigate a hard failure →

How you start

A fixed-scope Intelligence Investigation.

Intelligence is a proprietary product, private preview. It begins with an Intelligence Investigation: one hard failure or incident on one application, bounded evidence, a hypothesis set, at least one controlled intervention where feasible, a minimal reproducer, a verified repair or materially narrowed uncertainty, and a durable Core regression artifact. Ongoing work is scoped by covered applications, diagnostic capacity and deployment, may include Discovery for those applications, and offers private or sovereign deployment. Pricing is agreed privately in the preview and is not published; there are no list-price numbers here yet.

Related

EvalGlass Core
the open runtime every durable artifact runs in
EvalGlass Discovery
finds the missing evals and the hypotheses to investigate
The product family
Core executes, Discovery finds, Intelligence explains