Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Performance: Concepts and Metrics
- Analyzing latency, throughput, energy consumption, and resource usage
- Distinguishing between system-wide and model-specific bottlenecks
- Profiling strategies for inference versus training phases
Profiling Workflows on Huawei Ascend
- Leveraging CANN Profiler and MindInsight
- Diagnosing kernel and operator performance
- Analyzing offload patterns and memory mapping
Performance Analysis on Biren GPUs
- Utilizing Biren SDK monitoring features
- Exploring kernel fusion, memory alignment, and execution queues
- Profiling with consideration for power and thermal constraints
Optimization on Cambricon MLUs
- Employing BANGPy and Neuware performance utilities
- Interpreting kernel-level visibility and system logs
- Integrating MLU profilers with deployment frameworks
Graph and Model-Level Refinements
- Strategies for graph pruning and quantization
- Restructuring computational graphs through operator fusion
- Standardizing input sizes and optimizing batch processing
Memory and Kernel Enhancements
- Improving memory layout and data reuse
- Managing buffers efficiently across different chipsets
- Applying platform-specific kernel tuning methods
Best Practices Across Platforms
- Ensuring performance portability through abstraction strategies
- Developing shared tuning pipelines for multi-chip setups
- Case study: Optimizing an object detection model across Ascend, Biren, and MLU architectures
Conclusion and Forward-Looking Steps
Requirements
- Practical experience in AI model training or deployment pipelines
- Comprehensive understanding of GPU/MLU compute principles and model optimization
- Foundational knowledge of performance profiling tools and key metrics
Target Audience
- Performance Engineers
- Machine Learning Infrastructure Teams
- AI System Architects
21 Hours