Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren and its primary use cases.
  • Hardware configuration: cores, memory, and compute clusters.
  • Comparative analysis with NVIDIA and AMD GPUs.

Configuring the Biren Programming Environment

  • Installation of the Biren SDK and runtime.
  • Insight into the toolchain and compiler model.
  • Basic project structure and build workflows.

GPU Programming within the Biren Stack

  • Thread and block models.
  • Memory management and data transfer mechanisms.
  • Kernel development and launch strategies.

Migration from CUDA to Biren

  • Techniques for translating CUDA code.
  • Common API mappings and necessary adaptations.
  • Hands-on code conversion labs and practice sessions.

Debugging and Profiling

  • Utilizing Biren’s debugger and profiler tools.
  • Detecting performance bottlenecks.
  • Analyzing memory access patterns for optimization.

Optimization Strategies

  • Thread scheduling and instruction pipelining.
  • Loop unrolling and leveraging shared memory.
  • Advanced kernel tuning for enhanced throughput.

Case Studies and Practical Applications

  • Training a model using Biren accelerators.
  • Porting and profiling a vision or NLP model.
  • Performance comparison against CUDA/NVIDIA solutions.

Conclusion and Next Steps

Requirements

  • Proficiency in GPU architecture and parallel processing concepts.
  • Practical experience with CUDA, OpenCL, or comparable GPU programming environments.
  • Familiarity with deep learning frameworks such as PyTorch or TensorFlow.

Target Audience

  • HPC developers.
  • AI infrastructure engineers.
  • Performance optimization specialists.
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories