Evals
How I measure the work
Evals are infrastructure, not a scoreboard I curate. This is the registry: every system I can measure, the method I measure it with, and the honest current state.
- DBWhisper73% (101/139, dev)Benchmark / method
Spider
MetricExecution accuracy
73% (101/139, dev)DateJul 2026
NotesScoped Spider dev-split run of DBWhisper’s generation model (qwen/qwen3-32b at temp 0.1) with each database’s schema in context: 139 questions across 18 databases. Generated SQL is executed against the real Spider SQLite databases and compared by result set — ordered when the gold query has ORDER BY, multiset otherwise (standard execution match). 101/139 correct; malformed generations count as incorrect. 9 of the 148 sampled questions could not be scored after repeated provider throttling and are excluded, not counted either way. Measures the NL→SQL generation core; the deployed agent adds schema retrieval and a read-only validator on top.
method - DBWhisper82% exact · 100% fail-closedBenchmark / method
Custom golden-query set
MetricExecution accuracy · fail-closed refusals
82% exact · 100% fail-closedDateJul 2026
Notes22 natural-language golden queries + 4 unsafe/out-of-scope prompts over a read-only Postgres store, run end-to-end through DBWhisper’s live pipeline (schema retrieval → generation → read-only validator → execute). 82% exact result-set match (18/22); 95% (21/22) when crediting correct answers that returned an extra column. All 4 destructive or out-of-scope prompts were refused fail-closed (4/4).
method - CrownWager65.2% ± 0.8%Benchmark / method
NBA moneyline · 15,115 games
MetricModel accuracy (cross-validated)
65.2% ± 0.8%DateJul 2026
Notes5-fold stratified cross-validation on 15,115 NBA games — replacing an earlier post-hoc “best of 300 random splits” 68%. The base home-win rate is 57.5%, so this is a real ~8-point edge (ROC-AUC 0.685). Published picks are then graded against the real final score, and the track record is flagged “insufficient” below 20 settled picks.
method - TradePulse0 leaks · byte-identicalBenchmark / method
Look-ahead canary + parity suite
MetricLook-ahead leakage
0 leaks · byte-identicalDateJul 2026
NotesThe engine decides on closed bar i and fills at bar i+1’s open. A canary multiplies every future bar by 3× and asserts the past equity curve is byte-identical; any look-ahead leak changes the curve and fails the build. Correctness is proven structurally, not scored on returns — the platform deliberately publishes no flattering performance number.
method
“The model felt right” is not a number. For text-to-SQL that means execution accuracy — does the generated query return the correct rows when run against the real database — not string-matching against a reference.
When a result isn’t in yet, the row says so plainly rather than borrowing a number from somewhere else. When a run completes, its result, date, and method land here — and where a provider limit forces a partial run, the excluded count is stated, not hidden.