Get in Touch

Course Outline

Foundations of Performance: Concepts and Metrics

  • Analyzing latency, throughput, energy consumption, and resource usage
  • Distinguishing between system-wide and model-specific bottlenecks
  • Profiling strategies for inference versus training phases

Profiling Workflows on Huawei Ascend

  • Leveraging CANN Profiler and MindInsight
  • Diagnosing kernel and operator performance
  • Analyzing offload patterns and memory mapping

Performance Analysis on Biren GPUs

  • Utilizing Biren SDK monitoring features
  • Exploring kernel fusion, memory alignment, and execution queues
  • Profiling with consideration for power and thermal constraints

Optimization on Cambricon MLUs

  • Employing BANGPy and Neuware performance utilities
  • Interpreting kernel-level visibility and system logs
  • Integrating MLU profilers with deployment frameworks

Graph and Model-Level Refinements

  • Strategies for graph pruning and quantization
  • Restructuring computational graphs through operator fusion
  • Standardizing input sizes and optimizing batch processing

Memory and Kernel Enhancements

  • Improving memory layout and data reuse
  • Managing buffers efficiently across different chipsets
  • Applying platform-specific kernel tuning methods

Best Practices Across Platforms

  • Ensuring performance portability through abstraction strategies
  • Developing shared tuning pipelines for multi-chip setups
  • Case study: Optimizing an object detection model across Ascend, Biren, and MLU architectures

Conclusion and Forward-Looking Steps

Requirements

  • Practical experience in AI model training or deployment pipelines
  • Comprehensive understanding of GPU/MLU compute principles and model optimization
  • Foundational knowledge of performance profiling tools and key metrics

Target Audience

  • Performance Engineers
  • Machine Learning Infrastructure Teams
  • AI System Architects
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories