I design and ship production-grade AI systems: real-time voice agents, self-reflective RAG pipelines,
and cloud-deployed ML, where the LLM stays a language interface and deterministic engineering does the heavy lifting.
|
• Sub-500ms TTFT full-duplex conversational voice agents • Pipeline: Silero VAD · Whisper STT · LangGraph · LiveKit / Orpheus • Protocols: Pure WebSockets, low-latency streaming & dynamic barge-in |
• Self-Reflective RAG: Query rewriting, hallucination grading & abstention • Vector Engines: pgvector, Qdrant, multi-hop semantic graph routing • Benchmarking: RAGAS evaluation across 200k+ document biomedical corpuses |
|
• Zero Race Conditions: PostgreSQL row-level locks for concurrent booking • State Machines: LangGraph state graphs with strict Pydantic validation • Resilience: Redis distributed state caching, retry circuits & rate limiting |
• Containerization: Dockerized microservices & Kubernetes orchestration • AWS Ecosystem: S3, SageMaker, EC2, ECS, Lambda, IAM & VPC • CI/CD: Automated linting, test suites & deployment via GitHub Actions |
|
|
|
|
|
|