EvalGlass
← Trust

Reading a result

What a green result does not mean

No false green means a passing scorecard asserts only that measured checks cleared the thresholds you approved — never that your AI application is correct, safe, or production-ready.

A passing scorecard is a narrow, honest signal: the metrics that ran, on the examples you gave, met the thresholds you approved. It is not a verdict on your product. The value of EvalGlass is that it never lets a green run say more than the run actually proved.

A green run — what it does and doesn't license

Read a pass for exactly what it is. The left column is everything a green run earns you; the right column is everything it stays silent about.

A green run meansA green run does not mean
The specific metrics that ran actually produced valid measurements. Your output is correct.
They ran on the examples you supplied — not on everything your app does. Your output is safe, or free of harm you didn't measure.
Those measurements met the thresholds you approved. Your output is unbiased, fair, or complete.
Every gating metric was comparable where comparability was required. Your app is production-ready.

A perfect score is not a pass experimental

The sharpest proof that green never overclaims is structural, not a slogan: on the epistemic core the default gate reads the lower confidence bound of a metric’s interval, not the hopeful point — so a lucky handful of examples cannot clear it.

A perfect 3/3 does not clear the default gate

Three-for-three is a point of 1.0 — but its 95% lower bound is ≈ 0.44, so against a 0.8 threshold it does not pass. It passes only under a named point-smoke policy, a deliberate and recorded choice. Gather more evidence and the interval tightens until the lower bound clears the bar. Illustrative example, not a measured result.

Built in the framework; not yet in the version the plugin installs — the shipped path still gates on the point threshold. Why a perfect score can fail → · Confidence & intervals →

Non-scored states are not zero

When a metric can't honestly run, EvalGlass records why — it never invents a number to fill the gap. A blocked, non-evaluable, skipped, or errored result is the absence of a measurement, not a measurement of bad quality. None of these states is ever encoded as 0.0.

Informational is not gating

Scoring something and gating on it are two different acts. A run can measure plenty while deciding nothing, and EvalGlass keeps that line visible.

Read your scorecard responsibly

Every caveat above maps to a typed field in the Scorecard JSON — so you read the data, not the prose. When in doubt, check the field that proves the claim.

The cardinal question

Could this green scorecard be misread as proof the output is correct — when the run is only informational, unvalidated, uncalibrated, non-comparable, non-evaluable, or partly blocked? If it ever could, the result is overclaiming, and that is exactly the gap EvalGlass closes.

Related

Trust model
the six promises that keep an AI-run evaluation honest
Threat model
the ways a green run could lie, and what stops each
Honesty charter
the commitments EvalGlass holds itself to