EvalGlass

Use case · Prompt optimization

You changed the prompt. Did behavior actually improve?

A prompt edit that reads better can still regress real answers. Ask your coding agent to compare the new prompt to the last run on the same examples, and to surface the failure clusters — so “improved” means a measured delta, not a hunch.

Compare this prompt change against the last run. Did it improve the actual product behavior?
You Compare this prompt change against the last run. Did it improve the actual product behavior?
Your coding agent · Claude Code plugin Compared. On the examples you supplied, answer quality improved +0.05 and the “missing citation” failure cluster shrank, but two refusal cases regressed. The delta is comparable (same dataset, same metrics). Verdict informational — I measured the change; I did not tune the prompt, and nothing gates until you approve it.

What it checks

answer quality · faithfulness · failure clusters · refusal behavior. “Improved” is bounded to the examples you supplied and shown only when the runs are comparable — it is a measured delta on your checks, never a certification of the prompt.

The scorecard

prompt v3 · vs. prompt v2 informational
Answer quality0.88 +0.05
Missing-citation cluster3 cases −4
Refusal (should-answer)2 cases +2

Verdict informational — comparable on the supplied examples; no gate approved. Illustrative example, not a measured result.

What it will not claim

EvalGlass is not a prompt auto-tuner. It measures the behavior your edit produced and reports the delta; it never rewrites your prompt, and “better on these examples” is never “better in general.” No false green →

Ask your coding agent.

Evaluate my agentic app using EvalGlass.

Related

All use cases
every change moment
Scorecards
what each run hands back
Get the plugin
two commands in Claude Code
Docs
the detailed reference