Get in Touch

Course Outline

Preparing Machine Learning Models for Production Deployment

  • Packaging models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Considering versioning and storage strategies

Serving Models on Kubernetes

  • Introduction to inference server options
  • Deploying TensorFlow Serving and TorchServe
  • Configuring and managing model endpoints

Optimizing Inference Performance

  • Implementing effective batching strategies
  • Handling concurrent requests efficiently
  • Tuning for optimal latency and throughput

Scaling ML Workloads

  • Utilizing the Horizontal Pod Autoscaler (HPA)
  • Using the Vertical Pod Autoscaler (VPA)
  • Implementing Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Control

  • Setting up GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Strategic Model Rollout and Release

  • Applying Blue/green deployment patterns
  • Executing canary rollout strategies
  • Conducting A/B testing for model validation

Monitoring and Observability in Production

  • Tracking key metrics for inference workloads
  • Establishing best practices for logging and tracing
  • Configuring dashboards and alerting systems

Security and Reliability Best Practices

  • Securing model endpoints
  • Implementing network policies and access control mechanisms
  • Ensuring high availability for critical services

Summary and Future Directions

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Familiarity with fundamental Kubernetes concepts

Intended Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams
 14 Hours

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories