Machine Learning Engineering Open Book
-
Updated
Sep 23, 2026 - Python
Machine Learning Engineering Open Book
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Slurm: A Highly Scalable Workload Manager
A DSL for data-driven computational pipelines
A unified orchestration layer for heterogeneous AI compute. It standardizes how to manage compute and run training and inference on GPU clouds, Kubernetes, VMs, or bare-metal clusters.
A WDL, CWL and Python API supporting easy-to-use workflow engine. It is scalable, efficient and cross-platform (Linux/macOS).
Best practices & guides on how to write distributed pytorch training code
A Slurm cluster using docker-compose
Lightweight fast function pipeline (DAG) creation in pure Python for scientific (HPC) workflows 🕸️🧪
Best practices, reference architectures, and examples for distributed AI training and inference on AWS.
A scheduler for GPU/CPU tasks
Run Slurm in Kubernetes
TorchX is a universal job launcher for PyTorch applications. TorchX is designed to have fast iteration time for training/research and support for E2E production ML pipelines when you're ready.
Run Slurm on Kubernetes. A Slinky project.
Prometheus exporter for performance metrics from Slurm.
To associate your repository with the slurm topic, visit your repo's landing page and select "manage topics."