Run more RL experiments. Wait less for GPUs.
-
Updated
Aug 30, 2026 - Python
Run more RL experiments. Wait less for GPUs.
Fully Autonomous AI Research System with Self-Evolution, built natively on Claude Code
Tensor Fusion is a state-of-the-art GPU virtualization and pooling solution designed to optimize GPU cluster utilization to its fullest potential.
A tool for examining GPU scheduling behavior.
CPU + GPU scheduler optimized for modern-day desktop interactivity
SLURM-native software GPU slicing for NVIDIA clusters using memory limits and compute time-slicing.
A programmable distributed training system for PyTorch
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
PipelineScheduler optimizes workload distribution between servers and edge devices, setting optimal batch sizes to maximize throughput and minimize latency amid content dynamics and network instability. It also addresses resource contention with spatiotemporal inference scheduling to reduce co-location interference.
Persistent, quality-gated, resource-aware execution plans for DeepSeek Harness agents
Topology-aware Kubernetes scheduler for multi-tenant, heterogeneous clusters
hags gaming optimizer windows 11 — HAGS Gaming Optimizer for Windows 11 & 10. Direct repair download and step-by-step fix guide.
Seven hands-on labs for running AI workloads on Kubernetes: GPU scheduling, distributed training, model serving, and agents. Runs on a local kind cluster, no GPU required. Companion to a KubeCon EU 2026 talk.
A lightweight execution-control plane for AI workloads. It focuses on admission, durable job state, scheduling policy, worker ownership, recovery, artifact flow, and observability.
Human-approved, evidence-aware AutoDL GPU scheduling for Codex research workflows.
The GPU Optimizer for ML Models enhances GPU performance for machine learning. It offers advanced scheduling, real-time monitoring, and efficient resource management through a user-friendly web interface and robust API, integrating big data technologies for seamless data processing and model optimization. @NVIDIA
Topology- and workload-aware GPU scheduling with Ray and NVIDIA Dynamo for distributed LLM inference across heterogeneous GPU clusters, evaluated using normalized Job Completion Time.
Generate reproducible deep-learning workload traces by mapping production GPU-cluster traces to profiled executable workloads.
Simulation-based research prototype for fairness-aware, scarcity-adaptive GPU scheduling in a shared multi-institution compute federation, framed around African research infrastructure.
Reinforcement learning for LLM inference scheduling. DQN agent learns to balance throughput, TTFT, latency, and memory pressure vs FIFO/SJF/priority baselines.
To associate your repository with the gpu-scheduling topic, visit your repo's landing page and select "manage topics."