Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
EXO Infrastructure as Code
- Survey of EXO deployment configurations: single-node, multi-node, and RDMA clusters
- Automating dependency installation (Xcode, uv, Node.js, Rust) via configuration management
- Leveraging Nix flakes for reproducible EXO builds and developer environments
- Developing Ansible playbooks or shell scripts for autonomous cluster provisioning
Reproducible Builds and CI Integration
- Fixing dependencies and constructing the dashboard within CI pipelines
- Executing EXO smoke tests on GitHub Actions or GitLab CI runners
- Generating golden images and snapshot-based rollback workflows for macOS and Linux VMs
- Versioning custom model cards in parallel with application code
Cluster Discovery and Networking Automation
- Configuring mDNS and static DNS to ensure dependable libp2p node discovery
- Automating network profile generation and Thunderbolt bridge management on macOS
- Utilizing custom namespaces (EXO_LIBP2P_NAMESPACE) to segregate dev, staging, and prod clusters
- Establishing firewall rules and network segmentation for multi-tenant setups
Storage and Model Lifecycle Management
- Structuring EXO_MODELS_DIRS and EXO_MODELS_READ_ONLY_DIRS strategies
- Connecting NFS or SAN shares as read-only model repositories for rapid provisioning
- Managing garbage collection for outdated caches and versioned weight retention policies
- Automating model pre-downloads and health checks prior to rolling updates
Monitoring and Alerting
- Transmitting EXO logs to centralized logging systems (ELK, Loki, or Splunk)
- Developing Grafana dashboards based on EXO_TRACING_ENABLED output
- Triggering alerts for cluster membership shifts, OOM incidents, and inference latency surges
- Connecting macmon hardware telemetry with model performance degradations
Update, Rollback, and Disaster Recovery
- Testing EXO binary updates on a canary node before full fleet deployment
- Model-level rollback: toggling between quantized versions without requiring re-downloads
- Backing up and restoring cluster state, custom namespaces, and cached weights
- Compiling recovery runbooks for total cluster reconstruction scenarios
Security Hardening and Compliance
- Applying TLS at the reverse proxy layer (nginx, traefik) for the dashboard and API
- Implementing API rate limiting and IP whitelisting for EXO endpoints
- Isolating clusters using VLANs and zero-trust network policies
- Auditing access rights and maintaining an inventory of deployed models and versions
Requirements
- Proficiency in DevOps methodologies (CI/CD, IaC, container orchestration)
- Knowledge of macOS or Linux system administration and package management
- Comprehension of networking, DNS, and storage principles
Target Audience
- DevOps engineers
- Infrastructure architects
- SREs tasked with managing on-premise AI workloads
21 Hours
Testimonials (2)
Craig was extremely involved in the training, always making sure we are paying attention, adapted the examples to our day-to-day activities and always provided an answer when asked, even if the information was not added in the presentation.
Ecaterina Ioana Nicoale - BOOKING HOLDINGS ROMANIA SRL
Course - DevOps Foundation®
High level of commitment and knowledge of the trainer