Senior Cloud Platform Engineer
Building GPU/AI infrastructure at scale · CNCF Golden Kubestronaut · Open Source Builder
Senior Cloud Platform Engineer at W.W. Grainger, Inc. with deep expertise in cloud-native GPU/AI infrastructure, Kubernetes ecosystems, and platform engineering. I build open-source tools for GPU workload autoscaling, observability, and topology-aware incident response.
Actively contributing to CNCF projects with 31+ merged PRs across 17+ projects. Building GPU autoscaling, observability, and scheduling tools for Kubernetes.
Recognized as a CNCF Golden Kubestronaut — one of the elite professionals holding all five Kubernetes certifications. Elected Technical Lead of the CNCF TAG Workloads Foundation (2026–2028), a volunteer community role guiding workload scheduling, autoscaling, and batch processing across CNCF projects. Community member of the Dragonfly project and active contributor to Volcano, KEDA, OpenTelemetry, and more.
Volcano GPU NUMA-aware scheduler (3-repo PR), KEDA GPU Scaler, Kube Topology Agent, Dragonfly Community Member
OTel GPU Receiver, OpenTelemetry docs contributions, Kubernetes website docs PRs
OpenColorIO release signing & Vulkan tests, OpenCue subscription recalculation, OpenImageIO bug fix, RAWtoACES docs, xSTUDIO links fix
Achieved all 16 CNCF certifications including CKS, CKA, CKAD, KCNA, KCSA plus 11 Golden-tier certs
One of the elite professionals who have earned all CNCF Kubernetes and Cloud Native certifications — demonstrating comprehensive expertise across the entire cloud-native ecosystem. A highly selective professional designation held by fewer than 400 practitioners globally.
Certified Kubernetes Security Specialist
Certified Kubernetes Administrator
Certified Kubernetes Application Developer
Kubernetes & Cloud Native Associate
Kubernetes & Cloud Native Security Associate
Prometheus Certified Associate
Certified GitOps Associate
Certified Cilium Associate
Certified Argo Project Associate
Istio Certified Associate
Kyverno Certified Associate
OpenTelemetry Certified Associate
Cloud Native Platform Associate
Cloud Native Platform Engineer
Certified Backstage Associate
Linux Foundation Certified SysAdmin
Elected volunteer community Technical Lead (2026–2028) — guiding workload scheduling, autoscaling, and batch processing standards across CNCF projects
Elected via community governance vote — contributing to AI/ML model distribution, Helm charts, and dragonfly-injector
Active contributor across Volcano, Dragonfly, KEDA, Kubernetes, OpenTelemetry, and more
Contributing to HAMi (Heterogeneous AI Computing Virtualization Middleware) — GPU sharing and virtualization for Kubernetes
Recognized by Oracle for strong technical expertise and community contribution in cloud infrastructure and Kubernetes
Active contributor to CNCF foundation projects — 31+ PRs across 17+ repos
Cloud-native batch scheduling for AI/HPC
P2P file distribution & image acceleration
Production-grade container orchestration
Distributed transactional key-value database
Kubernetes event-driven autoscaling
Observability framework
Bare metal host provisioning for K8s
K8s-native packaging & resource management
Heterogeneous AI Computing Virtualization Middleware
Open source tools for GPU autoscaling, observability, and topology-aware infrastructure
Independent repository developing an event-driven GPU autoscaler using KEDA’s External gRPC Scaler interface. Native NVML metrics, DaemonSet deployment, pre-built scaling profiles for vLLM, Triton, and training workloads. Not yet merged into the KEDA core repository.
Referenced in KEDA #7538
NVIDIA GPU metrics as MCP tools for AI agents. Utilization, memory, temperature, power, and MIG instance metrics. Published on MCP Registry and ghcr.io.
Published on MCP Registry | ghcr.io
OpenTelemetry Collector receiver for NVIDIA GPU metrics. GPU utilization, memory, temperature via NVML. Standard OTLP export with built-in Prometheus exporter.
Kubernetes knowledge graph & automated root-cause analysis. Real-time resource topology, graph-based incident investigation, AlertManager webhook integration.
Kubernetes-native autoscaler for AI inference workloads. Custom scaling algorithms, GPU-focused policies, latency SLA enforcement, Prometheus metrics.
Comprehensive Kubernetes certification study guides covering all CNCF certifications. Interactive quizzes, flashcards, lab exercises, and PDF generation.
Industry publications, foundation blogs, and personal technical writing
How peer-to-peer networking fixes the 26 TB download problem when scaling LLMs across Oracle Kubernetes Engine and multicloud clusters — native hf:// and modelscope:// backends.
The GPU visibility gap in Kubernetes — NVML direct reads, scaling profiles, and the cost impact of scheduling blind to GPU utilization.
A deep dive into Kubernetes networking at the transport and network layers — how packets actually traverse the cluster data path.
Why autonomous AI agents need zero-trust guardrails at the infrastructure layer — moving beyond RBAC to real-time action evaluation with policy-as-code.
How to eliminate GPU budget waste with KEDA external scalers — native NVML metrics, DaemonSet architecture, and scaling profiles for vLLM, Triton, and training workloads.
How P2P mesh architecture eliminates registry bottlenecks in enterprise CI/CD — Dragonfly's distributed caching, bandwidth optimization, and multi-datacenter image distribution at scale.
Why standard HPA fails for LLM inference — token-aware autoscaling, KV cache pressure, GPU memory headroom, and building KEDA scalers for production serving.
Why autonomous infrastructure needs execution guardrails — policy-as-code with Kyverno, blast-radius containment, and building trust boundaries for AI agents in production.
Multi-cluster Argo CD on Oracle Kubernetes Engine — ApplicationSets, cluster secrets, RBAC delegation, and progressive delivery patterns for enterprise GitOps.
Running Docker-based AI agents on Oracle Cloud — Docker Model Runner, GPU shapes, container orchestration, and agent sandboxing on OKE.
How to restore the Golden Path for ML engineers by pushing GPU scaling complexity down the stack — edge-native NVML telemetry, KEDA External Scaler architecture, and eliminating the Prometheus latency trap.
The models are ready but the pipes aren’t — how CI/CD pipelines, GPU scheduling, model distribution, and governance are killing enterprise AI deployments.
Implementing zero-trust security on Oracle Kubernetes Engine with Terraform — IAM policies, network security groups, workload identity, and confidential computing.
Formal verification of ArgoCD manifests — resource invariants, temporal logic, and rollback safety for mission-critical deployments.
Dynamic resource allocation, in-place vertical scaling, and immutability improvements in Kubernetes v1.35 for AI/ML workloads and FinOps.
How AI agents, eBPF, and LLMs are transforming SRE from reactive incident management to autonomous self-healing infrastructure.
Why internal developer platforms need to prioritize the Java ecosystem — bridging enterprise reality with platform engineering ideals.
How ArgoCD v3 evolves from a sync tool to the backbone of modern platform engineering — multi-tenancy, scalability, and GitOps at enterprise scale.
Journey from Java developer to earning all CNCF certifications — what it takes to unlearn and re-learn in the cloud-native world.
How Dragonfly’s P2P architecture accelerates large AI model downloads — HuggingFace and ModelScope integration.
Conference presentations on cloud-native infrastructure, GitOps, and HPC
Pavan Madduri, W.W. Grainger
Applying ArgoCD, Kubernetes, and GitOps workflows to HPC environments — bridging the gap between cloud-native DevOps and scientific computing.
Pavan Madduri, W.W. Grainger
CI/CD pipelines, testing strategies, and automation for scientific and research software development — making open source science reproducible and maintainable.
Pavan Madduri (Grainger), Rohit Dhawan (Amazon), Alina Astapovich (Storytel), Goutham Rao (NeuBird) · Moderated by Renato Losio (InfoQ)
How AI agents and generative models are being used for incident detection, root cause analysis, and automated remediation — reducing MTTR and operational load at scale.