SDE-1 at Concentrix Catalyst, building agentic workflows in production. On my own time I build systems where an agent's mistake would be expensive, and I publish how I checked them.
| Project | What it is, and the proof |
|---|---|
| Argus showcase |
UI-testing agent: an LLM writes each test once, then it replays deterministically and heals UI changes 0 LLM calls over 10 runs · 5/5 hidden bugs caught |
| RaceLab demo |
What an agent should do when the data it reasoned about changes before it commits 0/50 bad commits vs 45–48/50 for blind retry, over 5,000 decisions |
| Continuity | Release checks for dubbed and subtitled films; Grafana decides whether each market can ship 5 markets · 87 checks · agents fix assets, never the verdict |
| VERDICT live |
Self-hostable hackathon judging with judge-bias correction Rank agreement τ 0.669 → 0.877 · 1,144 tests |
| Video anomaly page |
Traffic and CCTV anomaly detection across 11 classes, with event timestamps 38× faster than real time on a 6 GB laptop GPU |
| gpu-fleet-operator | Go control plane for a GPU inference fleet (Kubernetes, Temporal) 56 tests that need no cluster to run |
- Measure, then claim. Every number above comes from a run you can repeat, like Argus's recorded results.
- Keep the failures in the write-up. RaceLab lists the predictions it falsified, and the anomaly detector documents the model I didn't ship.
- Agents read the verdict; they don't write it. In Continuity, agents reach Grafana through a write-disabled MCP server.
1st place, DigiPay Pro (NPCI) competition, IIT Bombay Techfest 2024, out of 200+ teams (code)
Merged fixes in Eclipse Theia, Eclipse Thing-Web and rust-ffmpeg-sys
Writing: Teaching AI to read financial tables: fine-tuning LayoutLMv3
Python · Go · TypeScript · PyTorch · Playwright · FastAPI · Django · PostgreSQL · CockroachDB · Docker · GCP · AWS