EvalGlass

Project-fitted evaluation · open source

Your evaluations are your AI’s real spec. Make them fit your system.

In an AI system, what counts as correct no longer lives in the code — it lives in your evaluations. EvalGlass makes them project-fitted and owned: your coding agent helps you author and run them against your system, as code you keep and gate with statistical honesty.

Your evals, fitted Runs locally A suite you own No false green

The whitepaper

When the logic moves into weights nobody wrote, the definition of “correct” doesn’t vanish — it moves to the evaluations.

A generic catalog is someone else’s spec standing in for yours. Project-Fitted Evaluation makes the case in full — evaluation derived from the system under test, reported no greener than the evidence, grounded in decades of measurement science.

Version 1.0 · Apache-2.0 · free to read

The plugin

Install it. Point it at your repo. Read the scorecard.

Two commands in your coding agent, then plain language. Your agent helps you author the checks for your app, runs them locally in your repo, and hands back a scorecard like this — informational until you promote a gate. Nothing leaves your machine; no keys.

/plugin marketplace add Evalglass/evalglass-core
/plugin install evalglass-core@evalglass
model-switch run · vs. last baseline informational
Retrieval faithfulness0.91 +0.07
Refund-policy adherence0.96 held
Workflow-policy answers0.74 −0.11
Citation groundingblocked · judge not calibrated
Cost / 1k calls$2.4 −18%

A blocked metric shows its state, never a fabricated 0.0. Deltas appear only because the two runs are comparable. No threshold is approved, so the verdict is informational and nothing gates. Illustrative example, not a measured result.

A green scorecard never means more than the run measured.

EvalGlass measures and reports; you decide and gate. A fresh or example run is informational, never a false green — it doesn’t fail your build until you promote a metric to a gate on a threshold you approve. A regression isn’t a claim until two runs are genuinely comparable. It reads what your AI actually did; it never edits your app, and never certifies it. Local-first, repo-owned truth, no telemetry, no false green.

The product family

Start with Core. Add the next product only when the problem changes.

One brand, three products under one doctrine — not a Free / Pro / Enterprise ladder. Core executes. Discovery finds. Intelligence explains and learns. Start with Core; add Discovery when you need to know what the suite is missing, and Intelligence when you need to know why the application failed. Each answers a different question.

Core

open source

Define, run, compare and retain the evaluations you author — owned as code in your repo, with the scorecards, rubrics, calibration and authority records.

Discovery

private preview

Find the important application behaviors your current suite does not represent; review candidate evals and keep them as portable Core artifacts.

Intelligence

private preview

Test why the hard failure occurs with controlled replay and intervention, verify the repair, and preserve the learning as a durable Core regression.

Discovery and Intelligence are separate, paid, pre-alpha products — neither is generally available, and neither is shown running here. Compare all three →

Make AI answerable

Give your AI system a spec it can be held to.

You own your evaluations and decide what gates. Core runs what you define; Discovery finds what you’re missing; Intelligence explains what failed.

Built for your agent to read

Whitepaper
the case for project-fitted evaluation, in full
Docs
the detailed reference your agent reads
Read a scorecard
what every score means
What green doesn’t mean
what a pass may and may not say
The config it writes
the evaluation setup your agent generates