EvalGlass uses a fixed vocabulary so a term means the same thing on every page. This is the
single definition source; other pages link here on first use of a proper noun.
The effect-free centre: contracts, score semantics, aggregation, provenance, authority, and the Verdict Engine. Standard-library only; no I/O.
Runtime Harness
The effectful layer: CLI, config, adapters, replay, judge calls, persistence, reports, and CI-exit mapping. Owns effects, not meaning.
EvalGlass Skill
The integration-time installer skill (discover / plan / install / re-vendor), now delivered as part of the EvalGlass plugin. Grants no authority; not needed at runtime.
The declared slice of behavior evaluated: call, step, trajectory, or session.
Example
The evaluator-ready item: input, output, unit, optional reference, context, metadata, provenance.
Gold
Host-owned, validated reference data a reference metric compares against — never synthetic by default, never self-authorizing; only a host validation makes it gating-grade. (Marketing/trust pages say validated reference data instead.)
EvidenceBundle
References, source material, judge evidence, verifier evidence, runtime errors, and trace fragments passed into the core.
Measurement & meaning
Term
Definition
MetricSpec
A metric’s declared meaning: lens, score type, direction, profile, evidence needs, prerequisites, aggregation, threshold, authority, data policy.
The optional example_id/unit_id each Score carries (framework slice F1): which Example/EvalUnit it measured. Additive provenance, not meaning — it lets a reader group scores per call; it is never authority or source-function attribution.
An opt-in, deletable integration attached through an existing port.
Vendoring
Installing managed runtime source into a host repo under evals/_evalglass/. The skill copies only core/harness/adapters; the runtime then runs independently of the plugin and any agent.
The umbrella the plugin exposes: setup, connect, run, view, explain, compare, baseline, ci (plus the v1.1 authoring verbs add-metric/add-judge/calibrate). Bare /evalglass is an honest status overview that runs nothing. There is no gate/approve/certify verb — the absence is the identity.
view granularities
per-metric (default) · per-call (--by-call, grouped by Score subject identity, never list order) · per-source-function (advanced, not yet — needs a trace↔call-site correlation that does not exist).
Codex second runtime
EvalGlass also packages for Codex from the same canonical skills/ tree via a thin manifest; the vendored runtime is identical across runtimes. A public Codex marketplace listing is a maintainer step, not yet published.
Banned terms
Never write kernel, test kernel, or pure kernel for the
effect-free centre. Use Evaluation Core, effect-free core, or
core isolation. Likewise, never describe a result as proven-correct, warranted,
or endorsed — a Scorecard is a bounded claim,
not a stamp of approval.