Pavan Madduri

Pavan Madduri

Senior Cloud Platform Engineer

Building GPU/AI infrastructure at scale · CNCF Golden Kubestronaut · Open Source Builder

About Me

Senior Cloud Platform Engineer at W.W. Grainger, Inc. with deep expertise in cloud-native GPU/AI infrastructure, Kubernetes ecosystems, and platform engineering. I build open-source tools for GPU workload autoscaling, observability, and topology-aware incident response.

Actively contributing to CNCF projects with 31+ merged PRs across 17+ projects. Building GPU autoscaling, observability, and scheduling tools for Kubernetes.

Recognized as a CNCF Golden Kubestronaut — one of the elite professionals holding all five Kubernetes certifications. Elected Technical Lead of the CNCF TAG Workloads Foundation (2026–2028), a volunteer community role guiding workload scheduling, autoscaling, and batch processing across CNCF projects. Community member of the Dragonfly project and active contributor to Volcano, KEDA, OpenTelemetry, and more.

0
Open Source PRs
0
Projects Contributed
0
Featured Projects
0
CNCF Certifications

GitHub Activity

GitHub Contribution Heatmap
GitHub Streak GitHub Stats

Open Source Journey

2026 — Present

GPU NUMA Topology & AI Infrastructure

Volcano GPU NUMA-aware scheduler (3-repo PR), KEDA GPU Scaler, Kube Topology Agent, Dragonfly Community Member

VolcanoKEDADragonfly
2025

Cloud-Native Observability & Platform Engineering

OTel GPU Receiver, OpenTelemetry docs contributions, Kubernetes website docs PRs

OpenTelemetryKubernetes
2024

GPU/AI Infrastructure Contributions

OpenColorIO release signing & Vulkan tests, OpenCue subscription recalculation, OpenImageIO bug fix, RAWtoACES docs, xSTUDIO links fix

OpenColorIOOpenCueOpenImageIORAWtoACESxSTUDIO
2023

Golden Kubestronaut & Certification Journey

Achieved all 16 CNCF certifications including CKS, CKA, CKAD, KCNA, KCSA plus 11 Golden-tier certs

CNCFKubestronautCertifications

Achievements

CNCF Golden Kubestronaut

One of the elite professionals who have earned all CNCF Kubernetes and Cloud Native certifications — demonstrating comprehensive expertise across the entire cloud-native ecosystem. A highly selective professional designation held by fewer than 400 practitioners globally.

Kubestronaut Core Certifications

CKS

Certified Kubernetes Security Specialist

CKA

Certified Kubernetes Administrator

CKAD

Certified Kubernetes Application Developer

KCNA

Kubernetes & Cloud Native Associate

KCSA

Kubernetes & Cloud Native Security Associate

Golden Kubestronaut Certifications

PCA

Prometheus Certified Associate

CGOA

Certified GitOps Associate

CCA

Certified Cilium Associate

CAPA

Certified Argo Project Associate

ICA

Istio Certified Associate

KCA

Kyverno Certified Associate

OTCA

OpenTelemetry Certified Associate

CNPA

Cloud Native Platform Associate

CNPE

Cloud Native Platform Engineer

CBA

Certified Backstage Associate

LFCS

Linux Foundation Certified SysAdmin

Community Recognition

CNCF TAG Workloads Foundation Tech Lead

Elected volunteer community Technical Lead (2026–2028) — guiding workload scheduling, autoscaling, and batch processing standards across CNCF projects

Dragonfly Community Member

Elected via community governance vote — contributing to AI/ML model distribution, Helm charts, and dragonfly-injector

CNCF Contributor

Active contributor across Volcano, Dragonfly, KEDA, Kubernetes, OpenTelemetry, and more

HAMi Contributor

Contributing to HAMi (Heterogeneous AI Computing Virtualization Middleware) — GPU sharing and virtualization for Kubernetes

Oracle ACE Associate

Recognized by Oracle for strong technical expertise and community contribution in cloud infrastructure and Kubernetes

Original Open Source Contributions

Active contributor to CNCF foundation projects — 31+ PRs across 17+ repos

CNCF (Cloud Native Computing Foundation)

Volcano

Cloud-native batch scheduling for AI/HPC

Dragonfly

P2P file distribution & image acceleration

Community Member

Kubernetes

Production-grade container orchestration

  • #53891 Document deployment.kubernetes.io/* annotations
  • #53892 kubectl apply view-last-applied docs

TiKV

Distributed transactional key-value database

  • #19225 Add AGENTS.md for AI agent guidance

KEDA

Kubernetes event-driven autoscaling

  • keda-docs#1658 Remove deprecated metricName from docs
  • keda-docs#1769 Fix datadog scaler typos across all versions
  • #7538 GPU/AI inference scaler architectural analysis

OpenTelemetry

Observability framework

  • #8632 Add .NET troubleshooting page

Metal³

Bare metal host provisioning for K8s

  • #624 Fix redirect links in tryit.md

kpt

K8s-native packaging & resource management

  • #4278 Fix kpt fn doc for KRM functions

HAMi

Heterogeneous AI Computing Virtualization Middleware

  • #1893 Add unit tests for nvinternal info, mig, and watch packages

traceAI

Open-source LLM observability SDK

  • #165 Fix exporter shutdown and thread safety in Python SDK
  • #166 Add Go SDK with OpenAI instrumentor

Featured Projects

Open source tools for GPU autoscaling, observability, and topology-aware infrastructure

KEDA GPU Scaler Independent Repository

Independent repository developing an event-driven GPU autoscaler using KEDA’s External gRPC Scaler interface. Native NVML metrics, DaemonSet deployment, pre-built scaling profiles for vLLM, Triton, and training workloads. Not yet merged into the KEDA core repository.

GogRPCNVMLKubernetesHelm

Referenced in KEDA #7538

GPU MCP Server MCP Registry

NVIDIA GPU metrics as MCP tools for AI agents. Utilization, memory, temperature, power, and MIG instance metrics. Published on MCP Registry and ghcr.io.

GoMCP SDKNVMLDockerHelm

Published on MCP Registry | ghcr.io

OpenTelemetry GPU Receiver

OpenTelemetry Collector receiver for NVIDIA GPU metrics. GPU utilization, memory, temperature via NVML. Standard OTLP export with built-in Prometheus exporter.

GoOpenTelemetryNVMLPrometheus

Kube Topology Agent

Kubernetes knowledge graph & automated root-cause analysis. Real-time resource topology, graph-based incident investigation, AlertManager webhook integration.

GoKubernetes APIKnowledge GraphHelm

KubeAI Autoscaler

Kubernetes-native autoscaler for AI inference workloads. Custom scaling algorithms, GPU-focused policies, latency SLA enforcement, Prometheus metrics.

GoKubernetesCRDHelm

Golden Kubestronaut Learning

Comprehensive Kubernetes certification study guides covering all CNCF certifications. Interactive quizzes, flashcards, lab exercises, and PDF generation.

MkDocsPythonKubernetesEducation

Ingress2Gateway

Convert Kubernetes Ingress resources to Gateway API. Supports ALB, GCE, Nginx annotations with automated migration and validation.

PythonKubernetesGateway APIHelm

Technical Expertise

Container Orchestration & GitOps

KubernetesArgoCDDockerCrossplaneHelmFlux

Cloud Platforms

AWSAzureEKSEC2S3IAM

Observability

PrometheusGrafanaOpenTelemetrySplunkDatadog

Policy & Security

KyvernoOPAZero-TrustRBACNetwork Policies

CI/CD

GitHub ActionsJenkinsFluxUrbanCode Deploy

Languages & Tools

GoPythonRustTerraformgRPCBash

GPU / AI Infrastructure

NVIDIA NVMLCUDAvLLMTritonKEDAVolcano

Big Data

PrestoDBTrinoApache SupersetAlluxioJupyter

Technical Writing

Industry publications, foundation blogs, and personal technical writing

Industry Writing

Academic & Foundation Writing

Personal Blogs

Conferences & Speaking

Conference presentations on cloud-native infrastructure, GitOps, and HPC

HPSF Conference 2026 Productivity, Performance & the HPC Pipeline

GitOps for HPC: Bringing Cloud-Native DevOps Practices to High Performance Computing Environments

Pavan Madduri, W.W. Grainger

Applying ArgoCD, Kubernetes, and GitOps workflows to HPC environments — bridging the gap between cloud-native DevOps and scientific computing.

Chicago River Ballroom A-D Intermediate
HPSF Conference 2026 Building & Sustaining Community

DevOps for Scientific Software: Tools, Practices, and Automation Strategies

Pavan Madduri, W.W. Grainger

CI/CD pipelines, testing strategies, and automation for scientific and research software development — making open source science reproducible and maintainable.

Chicago River Ballroom A-D Beginner
InfoQ Live · Apr 21, 2026 Roundtable Panel

AI-Powered SRE for Autonomous Incident Response

Pavan Madduri (Grainger), Rohit Dhawan (Amazon), Alina Astapovich (Storytel), Goutham Rao (NeuBird) · Moderated by Renato Losio (InfoQ)

How AI agents and generative models are being used for incident detection, root cause analysis, and automated remediation — reducing MTTR and operational load at scale.

Online Panel Discussion

Let's Connect

Always open to connecting with fellow engineers in the cloud-native and AI/ML space

© 2026 Pavan Madduri. Built with passion for open source.