EvalGlass Discovery · find the missing evaluations
Find the evaluations your AI application is missing.
Core runs the evaluations you define. Discovery reads a bounded representation of your application — code, prompts, tools and selected traces — and proposes the evaluation system it actually needs: the consequential behaviors your current suite does not represent, turned into candidate evals you review, edit and keep as Core artifacts.
The gap Discovery closes
The failures that matter are defined by your application.
A generic metric catalog — faithfulness, relevancy, toxicity — is someone else’s spec standing in for yours. The behaviors that actually break an agentic application are specific to its workflow, tools, state and real traces. Core executes the evaluations you define; it deliberately does not inspect your application to infer what should be tested. That is Discovery’s job. It answers the questions a blank “write your own metric” box leaves to you:
- What does this application actually do? An application map reconstructed from what the artifacts support — entry points, routing, prompts, retrieval, tools, permissions, side effects, state and terminal outcomes — every relationship labelled by confidence.
- Where can behavior become consequential? The action surfaces where it matters: external writes, financial actions, permission changes, disclosures, escalation, irreversible tool chains, multi-agent handoffs.
- How might those paths fail, and which are untested? Failure hypotheses ranked by consequence, checked against your current suite to find the coverage gaps — not just the metrics that are easy to generate.
- What is the smallest credible evaluation, and what evidence is missing? The cheapest honest instrument for each gap, plus an explicit evidence-readiness request when the data, labels or instrumentation are not there yet.
Core gives the team the evaluation laboratory. Discovery decides which experiments the application needs.
The method
An evaluation-design compiler: inspect · map · hypothesize · design.
Discovery turns application artifacts, the current suite and any explicit requirements into an application map, then failure hypotheses, then coverage analysis, then candidate evaluations with instruments, scenarios, rubrics and evidence needs — all reviewed and edited by a human before anything is exported to your repo.
- Inspect. Read a bounded set of inputs you authorize — code and structure, prompts and instructions, tool and output schemas, selected traces, datasets and references, and your stated requirements. No standing access; scoped, time-bounded, customer-approved.
- Map. Reconstruct the application and its consequential action surfaces, with a confidence label on every relationship. The map is a structure for deciding what to evaluate — not automatically a causal graph.
- Hypothesize. Generate failure hypotheses, rank them by consequence, and run a failure critic that removes duplicates, unreachable paths, low-consequence noise, and claims that need Intelligence rather than Discovery.
- Design. For each surviving gap, propose the smallest honest instrument, failure-directed scenarios and contrast sets, a rubric with a deterministic–semantic split, a calibration plan, and an evidence-readiness verdict — emitted as editable Core files.
What Discovery proposes
Ten inspectable artifacts — every one a proposal you review.
- Application map. The reconstructed model of what the application does, with confidence labels (declared, observed, statically reachable, inferred, hypothesized, unknown).
- Consequential action surface. A place where behavior becomes material — a write, an action, a disclosure, an escalation — that the team may confirm as in-scope.
- Failure hypothesis. A trigger, workflow, manifestation, consequence, reachability evidence, current coverage and a proposed evaluation objective. A hypothesis is not a finding.
- Coverage gap. A consequential behavior the current suite does not represent — the thing you are not testing.
- Evaluation candidate. A proposed eval: assertion, scope, population, required evidence, instrument, aggregation, threshold or review rule, limitations and cost.
- Scenario proposal. Test data with execution context — state, memory, tool availability, mocked or delayed responses, approvals, ordering, expected invariants — often as minimal-pair contrast sets.
- Rubric proposal. A rubric split into its deterministic and genuinely-semantic parts, with anchors and hard negatives, and a recommended evaluator.
- Calibration plan. The labels, segments, tolerances and robustness probes a judge needs before it can gate. Core executes the plan.
- Evidence-readiness request. What evidence is available, derivable, or missing — missing evidence is a useful output, not a failed scan.
- Core export diff. Exactly which portable Core artifacts get added or updated on acceptance — nothing lands silently.
Discovery proposes. You accept.
A hypothesis is not a finding, and a candidate eval is not an active gate. Every proposal is inspectable — you can see why it was proposed and what evidence it requires — and human acceptance is mandatory; Discovery never silently creates a release gate. Everything you accept becomes a portable EvalGlass Core artifact, so if you stop the engagement your suite keeps running in Core. The work is bounded and customer-controlled — scoped repository and trace access, no standing identity — and it never claims causal proof: that is where Intelligence begins. Run a Discovery Proof →
How you start
A bounded Discovery Proof on one application.
Discovery is a proprietary product, private preview. It begins with a fixed-scope Discovery Proof: one application, its current suite, one material change or concern, bounded repository and trace access, a reviewed set of candidate evaluations, accepted Core artifacts and one rerun. Ongoing work is priced by active applications — never by seats, runs, metrics or traces — with local and private execution available. Pricing is agreed privately in the preview and is not published; there are no list-price numbers here yet.
Related