Note
For text-to-SQL, measure execution accuracy
- ~1 min read
- #evals#sql
- Source
Two SQL queries can be spelled completely differently and return exactly the same rows. So string-matching generated SQL against a reference — exact-match accuracy — punishes correct answers for cosmetic reasons.
The metric that matters is execution accuracy: run both queries against the database and compare result sets. That is what DBWhisper's golden-query set scores, and it is why the benchmark run uses Spider, whose harness scores execution rather than text.
Both runs have since completed. The scoped Spider dev run scored 73% (101/139), and the end-to-end golden-query set scored 82% exact result-set match — 95% if you credit answers that returned a correct extra column. That gap between 82% and 95% is the point: exact match is a deliberately harsh reading, and publishing the harsh one is the only version worth anything.
"The model felt right" is not a number. Both figures, their methods, and the nine questions a provider rate-limit forced out of the sample are on Evals.