DevOps Engineer · Platform & Observability · AI Infrastructure
CS graduate from NIT Goa. Currently building production monitoring infrastructure at Persistent Systems (embedded with IBM) — replaced Sysdig with a Prometheus/Grafana platform across 20+ microservices, cutting monitoring costs by 40%. I also build AI agent systems on the side and write about LLMOps on Substack.
Platform & Infra — Kubernetes · Docker · Helm · Tekton · GitHub Actions · IaC
Observability — Prometheus · Grafana · PromQL · Langfuse · Structured Logging · Locust
Cloud — IBM Cloud · AWS · Azure
AI / Backend — Python · FastAPI · LangChain · LangGraph · Groq API · REST APIs
Databases — PostgreSQL · MySQL · MongoDB
- Observability migration — Automated 45+ Sysdig→Grafana panel migrations with PromQL translation; migrated all alerting rules; 40% cost reduction across IBM internal engineering environments
- CI/CD reliability — Stabilized pipelines for 20+ microservices; resolved deployment, dependency, compliance, and secret-management failures
- Product Strategy Copilot — Production multi-agent AI backend (FastAPI + LangGraph + Groq) with full observability: Langfuse tracing, Prometheus metrics, structured logging, flamegraph profiling, and rate-limit resilience under Locust load testing
- AI document platform (Turtlemint internship) — Backend for LLM-based insurance document analysis; reduced review time by 30%
I write about LLMOps and AI agent observability on Substack — multi-agent architecture, production failure analysis using Langfuse traces, and platform engineering patterns for LLM backends.