EvalGlass
← Trust

Our covenant

We won't manufacture trust — not in a result, not on this page.

EvalGlass exists to give you an honest, narrow signal: a green run claims only what it actually checked. That same rule binds this website. If we'd refuse to let a scorecard overstate a run, we won't let our marketing overstate the project.

The covenant

No false confidence is the rule the product is built around. We extend it to ourselves.

A green or non-failing EvalGlass result never implies more evidence, authority, calibration, comparability, or safety than the run actually earned. A pass is a bounded, auditable claim — not a verdict on whether your output is correct.

And this site never implies more about the project than is true. EvalGlass is pre-alpha, version 0.1.0, with no tagged releases yet. It is open source under Apache-2.0, local-first, runs on your machine and in CI, sends no telemetry, needs no provider keys, and has no hosted platform and no funding channel wired. We state that plainly wherever it matters, rather than dressing the project up as more finished than it is.

This site articulates the full product vision; the four-state badge shows exactly how far the code has actually reached — vision-forward, never false-confident. A capability marked next or experimental is built in the framework but not yet in the version the plugin installs, and nothing marked so is ever shown executing as if it shipped. The ahead-of-code register is where we keep that promise auditable.

The verb we don't ship

The strongest honesty signal isn't a sentence on this page — it's a command that doesn't exist.

EvalGlass is delivered as a plugin you operate through your coding agent, with a small verb surface: /evalglass setup, connect, run, view, explain, compare, baseline, ci. There is deliberately no gate, approve, certify, pass, verify, or validate verb — and that absence is the identity, not an omission. The plugin is a typist and a reader, never a participant: it asks, the host validates, and a single Verdict Engine decides.

So a fresh run is informational, never "passing." Activating a gate is a host YAML edit you own (metric_status: gating, threshold_approval: approved), guided by a skill — the agent can type it at your direction; it can never grant the authority. A tool that could quietly make its own result pass would be exactly the false confidence EvalGlass exists to refuse. Where the claim stops →

Marketing patterns we refuse

Call us out

This covenant is only worth something if you can hold us to it. If any claim on this site overstates reality — a feature shown as shipped that isn't, a signal implied that the project hasn't earned — open an issue. We treat an overclaim on the website as the same class of bug as an overclaim in a scorecard.

This page is dated

A commitment you can't audit over time isn't a commitment. This page carries a date so you can see when it last changed and judge whether the rest of the site has kept up with it.

Last updated: 2 June 2026.

This page only earns trust if the whole site obeys it.

A charter is just more prose unless every other page holds to it. So the test isn't whether this page sounds honest — it's whether you can find a single claim across the site that the project hasn't actually earned. If you can, that's the bug. Tell us.

Related

Trust model
the six structural promises that keep an evaluation honest
What a green result does not mean
the exact boundary of a passing run
Governance
how the project is run and how to take part