Diagnose & intervene
Build the causal execution graph, form hypotheses, and test candidate causes with controlled replay and intervention — reporting the strongest supported evidence level, never a story.
EvalGlass Intelligence · causal failure intelligence
Core runs the evaluations you define; Discovery finds the ones you are missing. Intelligence is the advanced reasoning system for the hard failure — it determines what actually caused it, tests the cause with controlled replay and intervention, verifies the repair, and turns the learning into a durable evaluation you keep.
Discovery can tell you a retry path is untested; Core can tell you a metric moved. Intelligence begins after that — once you have an executable suite and a plausible hypothesis, and you need to know why a difficult, intermittent or consequential failure actually happens. It reasons about cause, and it refuses to compress that into a single unexplained number:
A conclusion is never presented as more certain than the evidence supports. Each causal claim carries an explicit evidence level, and a graph edge moves up the ladder only as evidence improves:
Uncertainty and failed interventions are valuable outputs, not failures: a non-discriminating experiment narrows the hypothesis space and is kept.
Diagnose & intervene
Build the causal execution graph, form hypotheses, and test candidate causes with controlled replay and intervention — reporting the strongest supported evidence level, never a story.
Verify the repair
Minimize the incident to a reproducer, and once you approve a change, verify the failure is gone with no nearby regressions — a repair-validated cause is the strongest evidence of all.
Learn
Feed the Evaluation Learning Loop: every incident and change becomes a proposed new or stronger evaluation you promote — so the suite gets better after every important failure. Everything durable runs in Core.
Intelligence improves your evaluations, not your production system autonomously; humans confirm the cause and approve every change, and no proposal enters the active suite by itself. It never calls your AI safe — only what the evidence supports, at a stated level. It exports to your Jira, incident, GRC and audit systems but does not become those systems. The work is bounded and customer-controlled — your application, your evidence window, your environment — with no broad standing access, and every durable artifact stays runnable in open-source Core. Investigate a hard failure →
Intelligence is a proprietary product, private preview. It begins with an Intelligence Investigation: one hard failure or incident on one application, bounded evidence, a hypothesis set, at least one controlled intervention where feasible, a minimal reproducer, a verified repair or materially narrowed uncertainty, and a durable Core regression artifact. Ongoing work is scoped by covered applications, diagnostic capacity and deployment, may include Discovery for those applications, and offers private or sovereign deployment. Pricing is agreed privately in the preview and is not published; there are no list-price numbers here yet.
Related