EvalGlass

EvalGlass Discovery · find the missing evaluations

Find the evaluations your AI application is missing.

Core runs the evaluations you define. Discovery reads a bounded representation of your application — code, prompts, tools and selected traces — and proposes the evaluation system it actually needs: the consequential behaviors your current suite does not represent, turned into candidate evals you review, edit and keep as Core artifacts.

Run a Discovery Proof See what Discovery would inspect private preview not part of the Core release

The gap Discovery closes

The failures that matter are defined by your application.

A generic metric catalog — faithfulness, relevancy, toxicity — is someone else’s spec standing in for yours. The behaviors that actually break an agentic application are specific to its workflow, tools, state and real traces. Core executes the evaluations you define; it deliberately does not inspect your application to infer what should be tested. That is Discovery’s job. It answers the questions a blank “write your own metric” box leaves to you:

Core gives the team the evaluation laboratory. Discovery decides which experiments the application needs.

The method

An evaluation-design compiler: inspect · map · hypothesize · design.

Discovery turns application artifacts, the current suite and any explicit requirements into an application map, then failure hypotheses, then coverage analysis, then candidate evaluations with instruments, scenarios, rubrics and evidence needs — all reviewed and edited by a human before anything is exported to your repo.

What Discovery proposes

Ten inspectable artifacts — every one a proposal you review.

Discovery proposes. You accept.

A hypothesis is not a finding, and a candidate eval is not an active gate. Every proposal is inspectable — you can see why it was proposed and what evidence it requires — and human acceptance is mandatory; Discovery never silently creates a release gate. Everything you accept becomes a portable EvalGlass Core artifact, so if you stop the engagement your suite keeps running in Core. The work is bounded and customer-controlled — scoped repository and trace access, no standing identity — and it never claims causal proof: that is where Intelligence begins. Run a Discovery Proof →

How you start

A bounded Discovery Proof on one application.

Discovery is a proprietary product, private preview. It begins with a fixed-scope Discovery Proof: one application, its current suite, one material change or concern, bounded repository and trace access, a reviewed set of candidate evaluations, accepted Core artifacts and one rerun. Ongoing work is priced by active applications — never by seats, runs, metrics or traces — with local and private execution available. Pricing is agreed privately in the preview and is not published; there are no list-price numbers here yet.

Related

EvalGlass Core
the open runtime that runs what Discovery proposes
EvalGlass Intelligence
explains why the hard failure happens
The product family
Core executes, Discovery finds, Intelligence explains