EvalGlass

EvalGlass product family

Three products. One evaluation system.

Each answers one question. Core: how do we run it? — Discovery: what are we missing? — Intelligence: why did it fail? Start with open-source Core; add Discovery when the problem is what to test, and Intelligence when the problem is why it failed. You adopt the next only when the problem changes — never a Free / Pro / Enterprise ladder.

Core

open source

open source · available today

The evaluation execution substrate: define, run, compare and retain project-fitted evaluations as code in your repo. Your coding agent helps you author them; you keep the authority.

  • A bounded scorecard you can stand behind in a PR
  • Metrics you author for your app, not a generic benchmark
  • You keep the authority over what gates — nothing self-approves

Discovery

private preview

private preview

Find the evaluations your AI application is missing — the consequential behaviors your current suite does not represent.

  • Reads a bounded representation of your app and proposes the evals it needs
  • Application map, failure hypotheses and candidate evals you review and accept
  • Accepted proposals become portable, Core-runnable artifacts you own

Intelligence

private preview

private preview

Understand why the hard failure happened — and make the evaluation system improve.

  • Diagnoses cause with controlled replay and intervention, at an explicit evidence level
  • Verifies the repair; association is never presented as proof
  • Turns the learning into a durable, Core-runnable regression you own

Use only what you need

A value progression, not a forced bundle.

Core measures. Discovery proposes. Intelligence tests causes. People approve.

No product grants itself the authority to decide. Core measures — the host approves what gates. Discovery proposes candidate evals — you review and accept them. Intelligence tests causes at an explicit evidence level — people confirm the cause and approve the change. Core stays open and locally useful on its own; generated artifacts remain customer-owned. Our honesty charter →

See it in practice

Read a scorecard
Core’s primary proof artifact
How it works
the loop your agent runs
Get Core
two commands in Claude Code