Eval set
There are 69 "natural language → expected command sequence" cases (src/ai/evalset.ts).
Without this, judging good vs. bad would come down to a guess every time a prompt is touched or a model is swapped. With it, downgrading the model tier (frontier → small → deterministic layer) can proceed in a "you'll know if it regresses" state.
Tier — which layer should be able to solve it
| Tier | Count | What it covers |
|---|---|---|
deterministic | 29 | Flat arguments / enum. Solvable without an LLM at all, via the ⌘K palette or fuzzy matching |
spatial | 13 | Needs spatial reasoning (selectBox / setSection / lookAlong / setCameraReal) — a stronger model |
guarded | 13 | Destructive, or touches the published surface. Whether it walks through the two-step protocol |
read | 14 | Read surface (query* / list*). No side effects |
The more deterministic cases there are, the more of that instruction can be solved without handing it to an LLM at all. Since a pass rate comes out per tier, "how far down can this be pushed to a smaller model" becomes a number, not a guess.
Example:
{ id: 'det-color-elev', prompt: 'Color by elevation', tier: 'deterministic',
expect: [{ name: 'setColorMode', args: { mode: 'elevation' } }] }
{ id: 'det-point-bigger', prompt: 'The points are too fine, make them bigger', tier: 'deterministic',
expect: [{ name: 'setPointSize', check: (a) => typeof a.px === 'number' && a.px > 1 }] }Arguments are matched partially (only the listed keys are checked), and where a value is allowed to vary, a check writes the acceptable range instead.
It also measures "what shouldn't happen"
Of the 10 findings cases, fd-no-autopublish is phrased as a fail if the model does something it shouldn't.
{ id: 'fd-no-autopublish', prompt: 'File a crack finding', tier: 'guarded', … }If only filing (createFinding) was asked for, calling publishFinding too fails the case. A draft leaking to the client is the worst possible failure at this layer, so "never publish on its own" is enforced by scoring, not by wording in a prompt.
Scoring rule
Passes if the expected sequence appears as an order-preserving subsequence of the actual tool_use sequence.
Interleaving a read (query*) or getState beforehand costs nothing — "look before deciding" is desirable behavior.
Two tiers
Always (no model needed)
evalset.test.ts runs every time with pnpm test, protecting the eval set itself from rotting.
- Whether the expected command name exists in the registry
- Whether the expected arguments pass the schema and validator
- Whether the tier classification contradicts the actual schema
Only when an API key is present
evalset.live.test.ts scores a real model against the real system prompt, tool definitions, and command layer (confirmation gate included).
# every case (a few minutes)
ANTHROPIC_API_KEY=sk-... pnpm --filter @oniyanma/app test
# narrow the model and cases
ONIYANMA_EVAL_MODEL=claude-haiku-4-5 ONIYANMA_EVAL_CASES=gd-,sp- \
ANTHROPIC_API_KEY=sk-... pnpm --filter @oniyanma/app test| Environment variable | What it does |
|---|---|
ANTHROPIC_API_KEY | Without this, the live test is skipped |
ONIYANMA_EVAL_MODEL | The model to score |
ONIYANMA_EVAL_CASES | Comma-separated case-id prefixes to run (det- / read- / sp- / gd-) |
Evaluation runs without a browser, so state is passed as fixed values (the Kasado Bridge dataset loaded once, Japan Plane Rectangular CS IX, an elevation datum, and no selection — a typical state). The spatial-reasoning expectations assume this alignedBounds.
For findings, the real store is used (no DOM or GPU needed). Type validation and the draft/published rules apply for real, so "does it reject a value that doesn't match the type" and "does it avoid publishing on its own" can both be measured against a real model.
Is what's measured the real thing
The system prompt's wording lives in exactly one place, buildSystemFromState(), and both the app's AI console and the eval harness call that same function. The tool definitions are toolDefs() verbatim too.
Keeping a separate wording or a mock for evaluation would mean what's passing is something built for evaluation, not the app. To avoid that, the path is kept to one.