EvalGlass

Roadmap & status

What is built, and what is not

EvalGlass is pre-alpha. This is the honest register of what runs today and what is still coming — a precise status on every capability (now, next, planned, experimental), and no dates we cannot keep.

The two status axes

Capability maturity — how far the code has actually reached:

Offer availability — whether a customer can obtain it under a real offer, shown by a separate squared badge, never a maturity chip:

EvalGlass Core open source

The open product you can install today — Apache-2.0, local-first, offer availability open source. Everything in this section is Core; the capability chips below are Core’s live maturity register, unchanged.

Maturity

Read this first. EvalGlass is pre-alpha and should be treated as such.

Built now

These pieces ship in v0.2.1 — they run on your machine and in CI today.

Ahead of code — every unshipped capability, with its status

Documented but not shipped to you. Each carries a precise status; none is shown executing as if it ran, and we name no date we cannot keep.

Non-goals & source of truth

EvalGlass Discovery private preview

Proprietary evaluation discovery, offered privately — it reads a bounded representation of your application and proposes the evaluations your suite is missing. Pre-alpha: nothing here is now — every capability is planned, and none is shown executing or producing a live result. A proposal is not a finding, and candidates never silently gate.

CapabilityOffer availability
Application map & consequential-action surfacesprivate preview
Failure hypotheses & coverage-gap analysisprivate preview
Candidate metrics & minimal-sufficient instrument designprivate preview
Application-directed scenario & rubric proposalsprivate preview
Calibration plans & evidence-readiness — exported as portable Core artifactsprivate preview

Begins with a Discovery Proof; availability is confirmed in an engagement scope, not a self-serve grid. Explore Discovery →

EvalGlass Intelligence private preview

Proprietary causal failure intelligence, offered privately — it diagnoses why a hard failure happens and makes the evaluation system learn. Pre-alpha: nothing here is now — every capability is planned, and none is shown executing or producing a live result. Association is not causation; every conclusion states its evidence level.

CapabilityOffer availability
Causal execution graph & hypothesis testingprivate preview
Controlled replay & intervention (graded evidence levels)private preview
Minimal reproducer & root-cause recordprivate preview
Repair verificationprivate preview
Evaluation Learning Loop — verified learning → durable Core regressionprivate preview

Begins with an Intelligence Investigation; may include Discovery for covered applications. Explore Intelligence →

No false confidence, applied to our own status

Every capability above shows its status honestly — Core’s shipped work is separated from what is still planned, nothing in Discovery or Intelligence reads now, and none is shown executing as if it ran. Nothing here implies production-readiness. Last updated: 15 August 2026.

See GitHub milestones

Related

Changelog
what changed, when
Sustainability
how the project keeps going
How it works
the path from evidence to verdict