Core
Define, run, compare and retain the evaluations you author — owned as code in your repo, with the scorecards, rubrics, calibration and authority records.
Project-fitted evaluation · open source
In an AI system, what counts as correct no longer lives in the code — it lives in your evaluations. EvalGlass makes them project-fitted and owned: your coding agent helps you author and run them against your system, as code you keep and gate with statistical honesty.
The whitepaper
A generic catalog is someone else’s spec standing in for yours. Project-Fitted Evaluation makes the case in full — evaluation derived from the system under test, reported no greener than the evidence, grounded in decades of measurement science.
Version 1.0 · Apache-2.0 · free to read
Two commands in your coding agent, then plain language. Your agent helps you author the checks for your app, runs them locally in your repo, and hands back a scorecard like this — informational until you promote a gate. Nothing leaves your machine; no keys.
/plugin marketplace add Evalglass/evalglass-core /plugin install evalglass-core@evalglass
All the ways in → What every scorecard means → How it works →
A blocked metric shows its state, never a fabricated 0.0. Deltas
appear only because the two runs are comparable. No threshold is approved, so the verdict is
informational and nothing gates. Illustrative example, not a measured result.
EvalGlass measures and reports; you decide and gate. A fresh or example run is informational, never a false green — it doesn’t fail your build until you promote a metric to a gate on a threshold you approve. A regression isn’t a claim until two runs are genuinely comparable. It reads what your AI actually did; it never edits your app, and never certifies it. Local-first, repo-owned truth, no telemetry, no false green.
One brand, three products under one doctrine — not a Free / Pro / Enterprise ladder. Core executes. Discovery finds. Intelligence explains and learns. Start with Core; add Discovery when you need to know what the suite is missing, and Intelligence when you need to know why the application failed. Each answers a different question.
Define, run, compare and retain the evaluations you author — owned as code in your repo, with the scorecards, rubrics, calibration and authority records.
Find the important application behaviors your current suite does not represent; review candidate evals and keep them as portable Core artifacts.
Test why the hard failure occurs with controlled replay and intervention, verify the repair, and preserve the learning as a durable Core regression.
Discovery and Intelligence are separate, paid, pre-alpha products — neither is generally available, and neither is shown running here. Compare all three →
Make AI answerable
You own your evaluations and decide what gates. Core runs what you define; Discovery finds what you’re missing; Intelligence explains what failed.
Built for your agent to read