📍 SF Bay Area, CA | 🧭 Principal PM | 🚀 Building full-cycle AI products
I turn human judgment into measurable product quality. After a decade shipping 0→1 and at-scale AI/ML products, I'm now building full-cycle AI tools end to end — from the underlying architecture to the practical apps people actually use.
I'm a Principal Product Manager in Walmart's Data & Ads Global Product Org, where I run the human-in-the-loop evaluation program behind AI-driven products — designing eval frameworks, scorecards, and annotator operations that make models measurably better.
My background blends Computer Science, Cognitive Science, and an MBA (Cornell University) — technical fluency paired with consulting-grade strategy. I care equally about the low-level architecture and the real-world application.
What I work on
- 🧪 LLM & AI evaluation — eval harnesses, scorecards, quality benchmarks, and human-in-the-loop pipelines
- 🧭 Annotator operations at scale — onboarding, calibration, and quality governance across global vendor teams
- 🎯 Recommendation & personalization — ranking, relevance, and retention-driving ML products
- 🛠️ Full-cycle product building — 0→1 prototypes through scaled launches, hands-on as both planner and builder
- 📊 Quantitative impact — every initiative tied to revenue, retention, quality, or efficiency
Twenty-nine repos is a pile, not a portfolio. Here is the same material sorted by what each piece is actually evidence of.
| Project | What it demonstrates |
|---|---|
| LLM Eval Scorecard · live | Side-by-side human eval with the parts most scorecards skip: randomized pane order with a position-bias test, bootstrap confidence intervals on the score gap, Krippendorff's α and weighted κ across raters, and a sample-size readout. Answers can you act on this yet?, not just who won? |
| AI Trackers | Three scheduled trackers whose entire logic lives in natural-language SKILL.md files rather than code — an agent reads them, gathers live data, and delivers a bilingual digest. A bet on prompts-as-programs. |
| Project | What it demonstrates |
|---|---|
| Ad Creative Optimizer · live | Predicts creative fatigue across Google, Meta and TikTok, then auto-rotates. Decisioning and pacing logic made inspectable. |
| Offsite Ads Demo · live | The off-site flow end to end — objective, budget, bid type, audience, multi-placement preview, 7-day KPIs — with Amazon DSP and Walmart Connect side by side. The $500 floor and the $10,000 floor select for different advertisers, and the whole build follows from that. |
| FarmVend · live | Vendor management for farmers markets: inventory prediction, payments, dynamic pricing. A whole two-sided marketplace at small scale. |
Each one picks a narrow problem and takes a position on it. The README in each explains the problem, the positions it defends, and what it deliberately is not.
| Project | What it demonstrates |
|---|---|
| Glean Delivery Intelligence · live | Proactive beats search — but only if every pushed insight can show the weight and measured precision of the signals beneath it. |
| Glean Compass · live | Priority drift, detected. Precision over recall on purpose: two flags, not forty, and dismissal is training data rather than a delete button. |
| Glean EI · full-stack · live | The same thesis with a real Express backend — confidence genuinely recomputes server-side, agents stream over SSE. |
| NerdWallet Money Next Steps · live | Sequencing beats ranking: clear the 24% APR balance before the rewards card. Plus a full-stack build on Next.js + SQLite with a live funnel. |
| Lyft Verticals · live | Airport, scheduled rides and teens as one system rather than three features. |
| Stitch Fix Household · live | Extending personal styling to a household — cross-category browse and shared profiles. |
| Midi Health · live | Five specific conversion changes to a telehealth funnel, each tied to the anxiety it removes. |
| Project | What it is |
|---|---|
| Sotto | Privacy-first on-device voice dictation. Speak softly, never type. |
| Notch 日迹 | Local-first macOS work journal that traces your day into AI summaries. |
| Meeting Scribe | Menu-bar app that notices a Zoom/Meet/Teams call has started and handles the notes. |
| LinkedIn Schedule Send | The button LinkedIn should have shipped. |
| Project | What it demonstrates |
|---|---|
| Adaptation Radar · live | Ranks books by screen-adaptation potential from four live public APIs, with every weight a slider so the model is something you argue with. A scheduled harvest records history, so the board shows the slope, not just the level. |
| StreamScope · live | Unifies viewing history across Netflix, HBO Max, Hulu, Disney+ and Prime. |
Animatics from Romance of the Three Kingdoms, and a song.
| Project | What it is |
|---|---|
| Red Cliffs 赤壁 · watch | The battle that split the empire three ways |
| Guandu 官渡 · watch | Cao Cao outnumbered ten to one |
| Three Visits 三顧茅廬 · watch | Liu Bei at the thatched cottage, three times |
| Six Sorties 六出祁山 · watch | Zhuge Liang's northern expeditions |
| 天不再借 · listen | Lyrics, arrangement, and three locally generated vocal takes |
- Walmart — Principal PM, Data & Ads · lifted model response quality 42% and helpfulness 38% across 10,000+ enterprise users; scaled eval ops to 600+ annotators across 4 vendor partners
- iHerb — Lead PM · improved recommendation precision 33%, cut churn 22%, raised inter-annotator agreement from 71% → 94%
- Wish — Senior PM · boosted engagement 38% and grew active users 147% on the personalized recommendation engine
- Amazon — Senior PM · lifted forecast accuracy 24% and cut stockouts 18% across millions of SKUs
- Building a full-cycle AI product — owning architecture and application end to end
- Open to remote Product Manager roles — AI/ML, evaluation, personalization, and platform products
- Sharing what I learn — practical patterns for human-in-the-loop evaluation and AI product development
"Ship measurable quality." — I build evaluation systems and AI products that turn human judgment into outcomes you can prove.

