Get in Touch

Course Outline

EXO Infrastructure as Code

  • Survey of EXO deployment configurations: single-node, multi-node, and RDMA clusters
  • Automating dependency installation (Xcode, uv, Node.js, Rust) via configuration management
  • Leveraging Nix flakes for reproducible EXO builds and developer environments
  • Developing Ansible playbooks or shell scripts for autonomous cluster provisioning

Reproducible Builds and CI Integration

  • Fixing dependencies and constructing the dashboard within CI pipelines
  • Executing EXO smoke tests on GitHub Actions or GitLab CI runners
  • Generating golden images and snapshot-based rollback workflows for macOS and Linux VMs
  • Versioning custom model cards in parallel with application code

Cluster Discovery and Networking Automation

  • Configuring mDNS and static DNS to ensure dependable libp2p node discovery
  • Automating network profile generation and Thunderbolt bridge management on macOS
  • Utilizing custom namespaces (EXO_LIBP2P_NAMESPACE) to segregate dev, staging, and prod clusters
  • Establishing firewall rules and network segmentation for multi-tenant setups

Storage and Model Lifecycle Management

  • Structuring EXO_MODELS_DIRS and EXO_MODELS_READ_ONLY_DIRS strategies
  • Connecting NFS or SAN shares as read-only model repositories for rapid provisioning
  • Managing garbage collection for outdated caches and versioned weight retention policies
  • Automating model pre-downloads and health checks prior to rolling updates

Monitoring and Alerting

  • Transmitting EXO logs to centralized logging systems (ELK, Loki, or Splunk)
  • Developing Grafana dashboards based on EXO_TRACING_ENABLED output
  • Triggering alerts for cluster membership shifts, OOM incidents, and inference latency surges
  • Connecting macmon hardware telemetry with model performance degradations

Update, Rollback, and Disaster Recovery

  • Testing EXO binary updates on a canary node before full fleet deployment
  • Model-level rollback: toggling between quantized versions without requiring re-downloads
  • Backing up and restoring cluster state, custom namespaces, and cached weights
  • Compiling recovery runbooks for total cluster reconstruction scenarios

Security Hardening and Compliance

  • Applying TLS at the reverse proxy layer (nginx, traefik) for the dashboard and API
  • Implementing API rate limiting and IP whitelisting for EXO endpoints
  • Isolating clusters using VLANs and zero-trust network policies
  • Auditing access rights and maintaining an inventory of deployed models and versions

Requirements

  • Proficiency in DevOps methodologies (CI/CD, IaC, container orchestration)
  • Knowledge of macOS or Linux system administration and package management
  • Comprehension of networking, DNS, and storage principles

Target Audience

  • DevOps engineers
  • Infrastructure architects
  • SREs tasked with managing on-premise AI workloads
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories