Case studies
Verifiable builds. Real numbers. Public source.
These are systems I designed and shipped myself — one that measures whether AI agents are safe to trust in regulated work, and one that turns a real problem into a live product in a weekend. Both are open to inspect: read the code, open the demo, check the numbers.
Applied AI research
Industrial Agent Reliability Benchmark
A reproducible benchmark that scores AI agents on the behavior single-turn evals miss — tool-call accuracy, grounding, escalation, and hallucination — across 18 synthetic aerospace workflows.
- Aerospace scenarios
- 18
- Reliability dimensions
- 6
- Reference vs. flawed composite
- 100 / 33
Aerospace scenarios
Reliability dimensions
Reference vs. flawed composite
Product build
My Bag Claim
A live baggage-claim product that turns aviation regulation into filing-deadline tracking and a demand-letter generator — designed, built, and shipped with payments and accounts in 72 hours.
- Repo to shipped MVP
- 72 hr
- Production features
- 6
- Free + paid tiers
- $0 / $19
Repo to shipped MVP
Production features
Free + paid tiers
On client work
Client engagement write-ups get added here as clients opt in.
The two studies above are my own builds — the clearest way to show the work when a firm is young and client engagements are still under NDA. As Assessments and Builds wrap and clients agree to share, named and anonymized write-ups will join them: starting-state diagnosis, the tier chosen, the system shipped, and the measured result — including honest trade-offs and anything that was out of scope.
Want to see this applied to your workflows?
Book the AI Readiness Assessment. If I can’t find $5K/month in recoverable time, I refund it.