I turn ambiguous business problems into products that can be tested, evaluated and shipped.
My work sits at the intersection of product strategy, enterprise AI, rapid prototyping, AI evaluation and human-in-the-loop systems.
I am particularly interested in the part of AI product development that happens after the demo works:
How do we know it is good enough to ship?
- 0→1 Product Development: problem discovery → product thesis → user journeys → prototype → launch criteria → iteration
- Enterprise AI Products: document intelligence, RAG, agentic workflows, decision-support systems and AI-assisted enterprise workflows
- AI Evaluation: golden datasets, error taxonomies, product-level metrics, launch thresholds, false-positive controls and human review
- Product Prototyping: building enough of the product to test assumptions before committing engineering capacity
- Product Measurement: KPI trees, instrumentation, experiments, operational metrics and adoption signals
A guided property-inspection product for Indian homebuyers taking possession of apartments and villas.
The product helps buyers move from an unstructured walkthrough to a guided inspection, evidence capture, issue tracking and a builder-ready action list.
Product questions explored
- How much guidance should a non-expert buyer receive?
- How do apartment and villa inspection journeys differ?
- How should evidence be stored and associated with issues?
- What should persist when an inspection is interrupted?
- Where should the product explicitly avoid making expert or structural claims?
- How can a large developer/project directory remain usable on mobile?
Current implementation includes
- Context-aware guided inspection
- Apartment and villa-specific inspection paths
- Evidence capture and issue register
- Save and resume
- Private user evidence
- Searchable developer and project discovery
- Mobile-first interaction
- Production and branch-preview QA
Stack used to prototype and ship: React, TypeScript, Supabase, Cloudflare and Playwright.
A public, executable collection of product-level evaluation patterns for AI systems.
The objective is not to benchmark models in isolation. It is to answer:
Would I ship this AI capability to users, and what evidence would justify that decision?
Initial evaluation packs cover:
- Document extraction and cross-document validation
- RAG and grounded-answer evaluation
- AI agent task-completion evaluation
Each pack connects:
Business objective → evaluation dataset → metrics → failure taxonomy → thresholds → human review → launch decision
View AI Product Evaluation Lab
A strong model metric can still produce an unacceptable product experience.
Every material AI capability should have an evaluation dataset, failure taxonomy, quality threshold and release decision.
Human-in-the-loop should be intentionally designed around confidence, risk, evidence and escalation.
A useful prototype should reduce uncertainty about user behavior, feasibility, value or risk.
The objective is a product users can trust and a business can operate.
Discover → Frame → Prioritize → Prototype → Evaluate → Ship → Measure → Iterate
I use AI-assisted development heavily for implementation speed and exploration. I remain accountable for the product problem, requirements, user journeys, prioritization, evaluation criteria, trade-offs and acceptance decisions.
This GitHub is not intended to be a collection of coding exercises. It is a growing set of working products, product experiments, AI evaluation systems and decision artifacts that demonstrate how I approach product problems.
Enterprise AI · AI Agents · RAG · AI Evaluation · Human-in-the-Loop · Document Intelligence · Product Analytics · Rapid Prototyping
Portfolio · India