Skip to content
View aivinay's full-sized avatar

Block or report aivinay

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
aivinay/README.md

header

Vinay Gupta

AI Infrastructure · Distributed Systems · Serverless Platforms

Senior Member of Technical Staff @ Oracle — OCI Functions

12+ years building hyperscale cloud infrastructure · prev. Microsoft Azure & Bloomberg · IEEE Senior Member

Website LinkedIn ORCID Email

🛰 What I do

I build and operate hyperscale cloud platforms. Currently I define technical strategy and architecture for OCI Functions — Oracle's Kubernetes-based serverless platform — including its multi-architecture (x86 + ARM) runtime migration. Daily tools: Go, Python, Java, Kubernetes, Terraform, and gRPC.

Previously: Microsoft Azure (real-time Kafka ingestion pipelines; security infrastructure for Azure Networking), Bloomberg (0→1 cloud-native alerting platform — Go, Python, AWS Lambda, PostgreSQL, OpenTelemetry, Kubernetes), and OCI Telemetry (Kafka-based routing for Oracle's observability data plane).

On the side, I build open-source AI-infrastructure tooling — model routing, agent observability, cluster guardrails, and reproducible ML data pipelines — each shipped with a citable preprint (below).

AI infrastructure · LLM inference & agents · Kubernetes · serverless · platform engineering · distributed systems · Kafka / streaming · observability

🚀 Projects — code ⇄ papers

Each project pairs working code with a citable preprint — design decisions and claims written down where they can be checked.

Project What it is Paper
switchboard Privacy-aware, local-first router for CLI coding agents (Codex, Claude Code) and local LLM inference — keeps sensitive prompts on-device. Benchmarked in the preprint: 62% fewer premium-agent calls at 4.1/5 vs 4.6/5 always-premium quality, with zero privacy leaks DOI
agent-tracebench Reproducible observability, replay, and regression checks for LLM agents DOI
kube-clusterguard Static guardrails for Kubernetes AI/ML compute clusters DOI
shardflow-ml Deterministic manifest, planning, and checkpoint layer for reproducible ML data pipelines DOI

Also: ai-spend-cap-tracker — a sourced, public tracker of organizations capping or cutting employee AI-coding spend; the cost pressure switchboard is built for.

🔧 Open source

🎤 Talks, publications & service

⚙️ Stack

Popular repositories Loading

  1. switchboard switchboard Public

    Privacy-aware, local-first router for your CLI coding agents (Codex, Claude Code) and local LLMs (Ollama) — keeps sensitive prompts on-device and cuts premium-model usage.

    Python 6 1

  2. agent-tracebench agent-tracebench Public

    Python 1

  3. shardflow-ml shardflow-ml Public

    Python 1

  4. kube-clusterguard kube-clusterguard Public

    Python 1

  5. predict_housing_price predict_housing_price Public

    Python

  6. stock_portfolio_recommender stock_portfolio_recommender Public

    Python