👋 Hi, I'm Hitesh Sahu
🚀 AI Infrastructure Engineer • GPU Systems • Cloud Native • Open Source
🌍 View Portfolio: https://hiteshsahu.com
I build the infrastructure behind modern AI systems.
My work focuses on GPU clusters, Slurm, Kubernetes, distributed systems, observability, and * developer tooling* that makes AI workloads easier to run, benchmark, and debug.
I'm currently building an open-source ecosystem for AI infrastructure, including local Slurm clusters, GPU observability, HPC developer tools, and LLM benchmarking.
These projects complement each other and solve a piece of the puzzle in the AI workflow from training, inference, deployment to monitoring
| Project | Description | Tech |
|---|---|---|
| 🏋️ Model Gym |
A fitness center for AI models. Import, export, optimize, benchmark, and report LLM inference performance across engines, runtimes, and hardware platforms. | Next.js • TypeScript • Python • PyTorch • Hugging Face |
| 🦆 RAG Factory |
Transforms chaotic PDFs, documents, websites, databases, and APIs into trusted answers using embeddings, retrieval, reranking, and large language models. | Python • FastAPI • LangChain • Vector Databases • OpenAI • Ollama |
| 🐸 NVIDIA SuperPod |
GPU Infrastructure Lab for building an AI supercomputer from commodity GPU servers. Explore multi-node training, networking, storage, scheduling, observability, and large-scale AI infrastructure. | Go • Kubernetes • Slurm • NVIDIA GPUs • InfiniBand • Prometheus |
| 🏴☠️ GhostFleet |
Simulates a 1,000-node / 8,000-GPU Kubernetes cluster on a laptop using KWOK, loads it with ClusterLoader2 and a custom GPU scheduling workload, and measures control-plane behavior against upstream scalability SLOs. | Go • Kubernetes • KWOK • ClusterLoader2 • Prometheus • Docker |
| ֎ GPU Lens |
Drop-in GPU + scheduler observability for clusters you already have. Get instant visibility into GPU health, utilization, memory, temperatures, ECC errors, XID faults, scheduler activity, and queue health. | Go • Prometheus • Grafana • DCGM Exporter • Kubernetes • Slurm |
| 🦝🐾 Squint |
A GPU-aware Slurm monitor for your terminal. Read-only, zero-config, and runs anywhere. Visualize jobs, nodes, GPUs, queue health, and pending reasons through a fast terminal UI. | Go • Bubble Tea • Lip Gloss • Slurm • TUI |
| 🐪 Caravan |
Spin up a complete local Slurm cluster with a single command. Develop, test, and submit HPC and AI workloads on Docker or Podman with GPU scheduling, making local experimentation fast and reproducible. | Go • Slurm • Docker • Podman • Cobra • HPC |
| 🛜 GPU-Fabric-Bench |
Reproducible RDMA fabric benchmarking suite for NCCL GPU collective communications on AWS EFA. Maps InfiniBand concepts to cloud-native HPC networking and visualizes latency, bandwidth, topology, and scaling behavior. | NCCL • AWS EFA • RDMA • MPI • NVIDIA GPUs • Python |
| GitHub Stats | Streak |
|---|---|
```