Machine Learning Engineering Open Book
-
Updated
Sep 12, 2026 - Python
Machine Learning Engineering Open Book
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Slurm: A Highly Scalable Workload Manager
A DSL for data-driven computational pipelines
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
A WDL, CWL and Python API supporting easy-to-use workflow engine. It is scalable, efficient and cross-platform (Linux/macOS).
Best practices & guides on how to write distributed pytorch training code
A Slurm cluster using docker-compose
Lightweight fast function pipeline (DAG) creation in pure Python for scientific (HPC) workflows 🕸️🧪
Best practices, reference architectures, and examples for distributed AI training and inference on AWS.
A scheduler for GPU/CPU tasks
Run Slurm in Kubernetes
TorchX is a universal job launcher for PyTorch applications. TorchX is designed to have fast iteration time for training/research and support for E2E production ML pipelines when you're ready.
Run Slurm on Kubernetes. A Slinky project.
Prometheus exporter for performance metrics from Slurm.
To associate your repository with the slurm topic, visit your repo's landing page and select "manage topics."