EvalGlass

Claude Code plugin

Evaluate your agentic app from Claude Code.

EvalGlass is the Claude Code plugin for AI quality control: your coding agent builds and runs project-specific checks for your AI application, locally and in CI. Codex is supported as a second runtime.

Install once in Claude Code, then ask in natural language — evaluate this app, compare a model, inspect a drift, wire CI. The plugin scaffolds, runs, and explains; it grants no gate, approve, or certify authority of its own.

Evaluate my AI application and show what improved and what regressed.

Claude Code first — not Claude-only

Claude Code is the primary plugin and launch surface. Codex is supported as a second coding-agent runtime through the same evaluation model; its public marketplace listing is a separate next step. Same repo-native loop, either runtime.

Two commands, then ask

Say this to your agent

Set up EvalGlass in this repo, then evaluate my AI application after the model switch.

It runs locally — no hosted platform, no provider keys on the run path — and hands back a bounded scorecard.

Next

Get the plugin →
install in Claude Code
Quickstart →
first run in minutes
Codex runtime →
second runtime, accurately scoped