Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren and its key use cases.
  • Hardware composition: cores, memory structures, and compute clusters.
  • Comparative analysis with NVIDIA and AMD GPUs.

Configuring the Biren Programming Environment

  • Installation of Biren SDK and runtime components.
  • Exploring the toolchain and compiler architecture.
  • Fundamental project structures and build workflows.

GPU Programming with the Biren Stack

  • Thread and block management models.
  • Memory management strategies and data transfer mechanisms.
  • Kernel development processes and launch patterns.

Migrating from CUDA to Biren

  • Methodologies for translating CUDA code.
  • Mapping common APIs and required adaptations.
  • Hands-on labs focused on code conversion practice.

Debugging and Profiling

  • Utilizing Biren’s integrated debugger and profiler.
  • Diagnosing performance bottlenecks.
  • Optimizing memory access patterns.

Optimization Techniques

  • Thread scheduling and instruction pipelining strategies.
  • Loop unrolling and effective shared memory utilization.
  • Advanced kernel tuning to maximize throughput.

Case Studies and Application Examples

  • Training models using Biren accelerators.
  • Porting and profiling vision or NLP models.
  • Benchmarking performance against CUDA and NVIDIA platforms.

Conclusion and Future Directions

Requirements

  • A solid understanding of GPU architecture and parallel processing concepts.
  • Prior experience with CUDA, OpenCL, or equivalent GPU programming frameworks.
  • Familiarity with deep learning frameworks like PyTorch or TensorFlow.

Audience

  • HPC developers.
  • AI infrastructure engineers.
  • Performance optimization specialists.
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories