We're finding out in public. Chapter 1 asked whether AI could build a real SQL engine at all. Chapter 2 asks for more: proof that its answers are right.
Step 0 was simple: is this worth doing, and can it be done at all? About 36 hours after our first recorded commit, the answer was yes, and we started chapter 2 the next afternoon. The fast engine passed 79.98% of all frozen SQLite test records (80.04% non-SQLite maximum), or 99.93% of applicable records, with zero observed wrong answers in that run.
But passing tests only shows that no wrong answer turned up. It can't show none exists, and a database that gets an answer wrong doesn't crash. It hands you a number that looks right. So step 1 is more audacious: prove the answers are right, with machine-checked proofs against a written spec, and grow the share of real SQL that's covered until it reaches the ceiling too.
Each pilot record that doesn't count yet stopped at the first thing the proven core can't handle. Fixing one item can uncover the next, so these don't simply add up. They do show where the biggest gains are.
| Date | Change | Owner | Tools | Result |
|---|
Today this log is updated with each export of the code. Once the repository is public, rows will come from merged pull requests and their checks.
Raises the number. It counts only when all of these hold, rechecked from scratch:
Makes the test harder and the number more meaningful. It may go down, and that's fine.