Case studies
Verifiable builds. Real numbers. Public source.
These are systems I designed and shipped myself — one that measures whether AI agents are safe to trust in regulated work, and one that turns a real problem into a live product in a weekend. Both are open to inspect: read the code, open the demo, check the numbers.
Applied AI research
Industrial Agent Reliability Benchmark
A reproducible benchmark that scores AI agents on the behavior single-turn evals miss — tool-call accuracy, grounding, escalation, and hallucination — across 18 synthetic aerospace workflows.
- Aerospace scenarios
- 18
- Reliability dimensions
- 6
- Reference vs. flawed composite
- 100 / 33
Aerospace scenarios
Reliability dimensions
Reference vs. flawed composite
Product build
My Bag Claim
A live baggage-claim product that turns aviation regulation into filing-deadline tracking and a demand-letter generator — designed, built, and shipped with payments and accounts in 72 hours.
- Repo to shipped MVP
- 72 hr
- Production features
- 6
- Published paid tiers
- $29–$149
Repo to shipped MVP
Production features
Published paid tiers
On client work
Client engagement write-ups get added here as clients opt in.
The two studies above are founder-built proof, not client results. That distinction stays explicit. As engagements wrap and clients opt in, named or anonymized write-ups will add the starting baseline, scope, system shipped, evaluation method, and measured result—including trade-offs and anything that was out of scope.
Want to see this applied to your workflows?
Start with the 30-Day Revenue Follow-Through Pilot when the workflow is clear. Use the Assessment first when it is not.