Use case · Prompt optimization
You changed the prompt. Did behavior actually improve?
A prompt edit that reads better can still regress real answers. Ask your coding agent to compare the new prompt to the last run on the same examples, and to surface the failure clusters — so “improved” means a measured delta, not a hunch.
Compare this prompt change against the last run. Did it improve the actual product behavior?
What it checks
answer quality · faithfulness · failure clusters · refusal behavior. “Improved” is bounded to the examples you supplied and shown only when the runs are comparable — it is a measured delta on your checks, never a certification of the prompt.
The scorecard
Verdict informational — comparable on the supplied examples; no gate approved. Illustrative example, not a measured result.
What it will not claim
EvalGlass is not a prompt auto-tuner. It measures the behavior your edit produced and reports the delta; it never rewrites your prompt, and “better on these examples” is never “better in general.” No false green →
Ask your coding agent.
Evaluate my agentic app using EvalGlass.
Related