I build and run distributed systems at scale: reliability, cloud infrastructure, and the platform work that keeps teams shipping. These days I'm building LLM agent infrastructure in the open.
Open to Infrastructure Engineer / Senior SWE / Staff Engineer / SRE roles. Remote or Hybrid or On-site
- 💰 Migrated a high-throughput distributed system across 100s of nodes to an open-source licensing alternative, saving $1.6M/year. Zero downtime, start to finish
- ☁️ Identified and eliminated idle compute waste across all Kubernetes clusters through scheduled scaling, cutting $100K/year in cloud spend
- ⚡ Migrated all team repos to GitHub Actions in 3 days, 2 months ahead of deadline
- 📊 Full-stack observability: OpenTelemetry → Grafana SLI/SLO for production GRC platform
- 🔐 Zero-downtime S3 SDK migration under FedRAMP + SOX compliance constraints
- 🚀 ChatOps deployment workflow that increased daily deploys 5x
ripe, planefs, geminon (see above)
- 🔐 FIPS 140-2 compliance test suites (Kotlin/Java)
- GitHub Actions migration: all repos in 3 days, 2 months early
- Production runbooks, ChatOps deploys, monitoring metrics
- 💰 $1.6M/year: zero-downtime distributed system migration across 100s of nodes
- $100K/year: fleet-wide scheduled scaling to eliminate idle compute
- Cassandra/MySQL backends, Terraform, Go/Java/Kotlin stack
- Multi-datacenter Linux infrastructure
- KVM prototyping; MongoDB/Python NOC messaging system