Review and improve truth
Annotation workflow experimental
Route uncertain examples to human review and bring validated truth back into your repo.
What it does
You review the uncertain examples your app produced and mark each right or wrong; the validated labels return to your repo as host-owned truth to score against. You decide every call — EvalGlass just collects them. The governance to return validated truth exists today; a polished annotation UI does not.
Say this to your agent
Let me label these examples so I can score against them.
You decide what counts as right or wrong.
Experimental — the workflow that returns validated labels to host-owned data exists; a polished UI does not. Either way, you decide every label, and nothing becomes validated reference data without your review.
In the docs
Read the detail →
what your agent reads for this