Open to AI engineering roles

Now

What I'm doing now

A running snapshot of what I'm building, exploring, and reading — updated by hand, not scraped. Inspired by Derek Sivers' /now page.

Last updated Jul 2026

Building now

  • DBWhisper evals — the natural-language-to-SQL agent is now measured on execution accuracy, not vibes: a golden-query harness (82% exact, 100% fail-closed) plus a scoped Spider dev run (73%, 101/139). Numbers and method live on Evals.
  • This site as a product — turning the portfolio into a running notebook: Notes, a public eval registry, and the grounded, cited "Ask this site" assistant, all on the same $0 stack.
  • Keeping four products live — DBWhisper, TradePulse, CrownWager, and LLM Studio are deployed and maintained, not screenshots.

Reading

  • The Model Context Protocol specification — typed tool-use as a protocol, not per-agent glue.
  • Spider (Yu et al., 2018) — the cross-domain text-to-SQL benchmark I'm measuring DBWhisper against.
  • Agent memory beyond flat conversation summaries — what actually survives a long session, and what should.

Currently exploring

  • Golden-query evals for text-to-SQL

    Execution accuracy against a real database, not string-matching against a reference query.

  • Model Context Protocol for typed tool-use

    Tool contracts as a protocol between agents, rather than bespoke glue per integration.

  • Long context vs. retrieval

    When a bigger window still loses to a small, scoped prompt — and how to tell in advance.


Open to

Open to senior AI engineering and applied AI roles — remote or Ahmedabad, India.