EvalGlass

Trust

Your coding agent runs the evaluation. It never decides what the scores are worth.

EvalGlass stays honest by design: local-first, host-owned truth, bounded scorecards, and no false green — a pass is a bounded claim, never proof your AI is correct.

Letting an agent set up your evaluation is only safe if it can’t quietly inflate what a score means. EvalGlass is built so it can’t — not by policy, but by construction. Six promises hold whether a human or an agent is at the keyboard.

Human or agent

An honest evaluation

no false confidence

You decide what countsA score gates CI only after you approve the bar — never the agent.
Green never overclaimsA pass means only what the run checked — not that your output is correct.
AI judges earn the gateAn LLM grader can’t fail your build until it’s calibrated to your labels.
Every score shows its evidenceEach result traces to records you can inspect — never just prose.
It measures, never editsIt reads your outputs; it never changes your prompts, models, or code.
Nothing hidden, nothing hostedVendored in your repo — no platform, no keys; it outlives the agent.
An agent can run every step — but it can’t inflate what a score means. Six promises hold the line, whether a human or an agent is at the keyboard.

Each promise — and where it’s enforced

The question that defines it

Could a green scorecard be read as proof your output is correct — when the run only scored your outputs, with no threshold approved, no judge calibrated, and no baseline compared? If it ever could, that’s the gap EvalGlass closes.

Go deeper on the safety story

What a green result doesn’t mean
what a pass does and doesn’t license
Statistical honesty
why a perfect score can fail the lower-bound gate
Threat model
the adversary is false confidence — attack by attack
Honesty charter
the trust signals we refuse to fake