EvalGlass
← All extensions

Review and improve truth

Annotation workflow experimental

Route uncertain examples to human review and bring validated truth back into your repo.

What it does

You review the uncertain examples your app produced and mark each right or wrong; the validated labels return to your repo as host-owned truth to score against. You decide every call — EvalGlass just collects them. The governance to return validated truth exists today; a polished annotation UI does not.

Say this to your agent

Let me label these examples so I can score against them.

You decide what counts as right or wrong.

Experimental — the workflow that returns validated labels to host-owned data exists; a polished UI does not. Either way, you decide every label, and nothing becomes validated reference data without your review.

In the docs

Read the detail →
what your agent reads for this