Now
What I'm doing now
A running snapshot of what I'm building, exploring, and reading — updated by hand, not scraped. Inspired by Derek Sivers' /now page.
Last updated Jul 2026
Building now
- DBWhisper evals — the natural-language-to-SQL agent is now measured on execution accuracy, not vibes: a golden-query harness (82% exact, 100% fail-closed) plus a scoped Spider dev run (73%, 101/139). Numbers and method live on Evals.
- This site as a product — turning the portfolio into a running notebook: Notes, a public eval registry, and the grounded, cited "Ask this site" assistant, all on the same $0 stack.
- Keeping four products live — DBWhisper, TradePulse, CrownWager, and LLM Studio are deployed and maintained, not screenshots.
Reading
- The Model Context Protocol specification — typed tool-use as a protocol, not per-agent glue.
- Spider (Yu et al., 2018) — the cross-domain text-to-SQL benchmark I'm measuring DBWhisper against.
- Agent memory beyond flat conversation summaries — what actually survives a long session, and what should.
Currently exploring
Golden-query evals for text-to-SQL
Execution accuracy against a real database, not string-matching against a reference query.
Model Context Protocol for typed tool-use
Tool contracts as a protocol between agents, rather than bespoke glue per integration.
Long context vs. retrieval
When a bigger window still loses to a small, scoped prompt — and how to tell in advance.
Open to
Open to senior AI engineering and applied AI roles — remote or Ahmedabad, India.