EvalGlass

Extensions

Add the evaluation team when you need it.

Extensions deepen the same repo-native loop — bring behavior in, turn it into checks, review uncertainty, and send scorecards where your team works. Each is an optional, removable specialist move, not a hosted product suite, and none of them decides your scores.

Every add-on keeps the same promise.

Optional, removable, and never authoritative — an extension can bring behavior in or send a scorecard out, but it can never grant a gate or change a verdict. You say what counts.

Bring behavior in

Get your AI application’s real behavior into EvalGlass.

Turn behavior into checks

Shape project-specific checks around your product’s behavior — they land proposed, never authoritative.

Review and improve truth

Bring human judgment in where uncertainty matters.

Use results

Act on your scores — you decide what gates.

GitHub CI
now

Fail the build only when an approved gate is breached — never on a fresh, informational run.

Dashboard sink
experimental

Send scorecards outward — local export now, a hosted view is next. No telemetry; the repo stays the source of truth.

Diagnostic clusters
next

Group a run's failing and non-scored cases by shared cause (Scorecard.clusters) — “the 18% that failed are all missing-citation cases,” not one flat number. A non-scored case is grouped by cause, never coerced to 0.0.

Metrics explorer
planned

A separate, still-planned view: browse deltas and distributions across runs by call identity — a different axis from the clusters above.

Prompt optimizer
next

Send a copy of your scorecard to an optimizer — EvalGlass never edits your prompts.

Related

How it works
the loop these deepen
Scorecards
the artifact they act on
Roadmap & status
what is now, next, planned, experimental
Get the plugin
start with the core loop