Tracel Technologies’ cover photo
Tracel Technologies

Tracel Technologies

Technology, Information and Internet

Québec, Quebec 737 followers

High-performance computing to bring intelligence everywhere

About us

At Tracel, we’re solving the challenge of AI model portability and efficiency across platforms, making deployment simpler and freeing teams from hardware constraints. We believe compute is central to AI’s evolution, powering smarter, more adaptable models. Through open-source tools, we’re building a foundation that drives AI innovation and scalability without the need for fragmented implementations.

Website
https://tracel.ai/
Industry
Technology, Information and Internet
Company size
2-10 employees
Headquarters
Québec, Quebec
Type
Privately Held
Founded
2023
Specialties
Artificial Intelligence, Machine Learning, ML Ops, Deep Learning, GPGPU, and Rust

Locations

Employees at Tracel Technologies

Updates

  • Tracel Technologies reposted this

    Jetson + Copper-rs + Burn + CUDA Oxide + Nix Deterministic operating systems for Physical AI, Aerospace, Robotics, Space and Manufacturing are going to matter a lot. Copper-rs https://lnkd.in/gQ9AdEW8 NVIDIA CUDA Oxide https://lnkd.in/g-PBXx9j Burn https://burn.dev/ Jetpack-nixos for Jetson https://lnkd.in/gNx4tAMm Esp32 Wireless Offense https://lnkd.in/gV9QyQmE Huge respect to Anduril Industries , Copper Robotics Inc., founder Guillaume Binet , Tracel Technologies , "justcallmekoko" and NVIDIA for meaningfully contributing to open source. I’m testing and reading a ton while I wait for my Jetson to arrive, and I’m genuinely excited for the Physical AI boom to accelerate. This feels like the perfect opportunity to combine robotics, embedded systems, deterministic software, GPU acceleration, Rust, and real HPC experience into something useful. Physical AI is going to need more than existing models

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
      +5
  • Burn 0.21.0 brings 4 months of improvements that make the framework up to 8 times faster and more reliable across the board. The gains span distributed workflows for training large models all the way down to small-model inference, where the reduced framework overhead becomes especially noticeable. We rethought our distributed computing stack around differentiable collective operations. Kernel selection is now more reliable thanks to better autotuning and a new validation layer, and a project-level burn.toml file lets you tweak those internals (and many others) without recompiling. A reworked device handle reduces framework overhead, and a new burn-dispatch crate simplifies backend selection while paving the way for faster compile times. The release also ships burn-flex, a lightweight eager CPU backend for WebAssembly and embedded targets that replaces burn-ndarray. Finally, we added early off-policy reinforcement learning support and a fresh round of kernel work on GEMV, top-k, and FFT. For more details about the release, you can read our blog post: https://lnkd.in/e36h_2yr

  • 𝗕𝘂𝗿𝗻 𝟬.𝟮𝟬.𝟬 𝗥𝗲𝗹𝗲𝗮𝘀𝗲: 𝗨𝗻𝗶𝗳𝗶𝗲𝗱 𝗖𝗣𝗨 & 𝗚𝗣𝗨 𝗣𝗿𝗼𝗴𝗿𝗮𝗺𝗺𝗶𝗻𝗴 𝘄𝗶𝘁𝗵 𝗖𝘂𝗯𝗲𝗖𝗟 𝗮𝗻𝗱 𝗕𝗹𝗮𝗰𝗸𝘄𝗲𝗹𝗹 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻𝘀 It’s been an intense few months of development, and we’re ready to release Burn 0.20.0. Our goal was to solve a classic challenge in HPC: achieving peak performance on diverse hardware without maintaining a fragmented codebase. By unifying CPU and GPU kernels through CubeCL, we’ve managed to squeeze maximum efficiency out of everything from NVIDIA Blackwell GPUs to standard consumer CPUs. 𝗖𝘂𝗯𝗲𝗖𝗟 𝗖𝗣𝗨 𝗢𝘃𝗲𝗿𝗵𝗮𝘂𝗹 The CubeCL CPU backend received a major update. It now features proper lazy execution and the same multi-stream support as our WGPU runtime. We’ve also added support for kernel fusion, which was a missing piece in our previous CPU backends. In addition, by focusing on cache line alignment and memory coalescing, our kernels are now outperforming established libraries like libtorch in several benchmarks. On the figure, you can see that CubeCL achieves up to a 4x speedup over LibTorch CPU, with even larger margins compared to SIMD-enabled ndarray. The real win here is that CubeCL kernels are designed to adapt their computation based on launch arguments. By selecting the optimal line size (vectorization), cube dimensions, and cube counts specifically for the CPU, we can control exactly how threads map to data without touching the kernel code. We increased the line size to ensure optimal SIMD vectorization and tuned the cube settings so that data ranges respect physical cache line boundaries. This automatically eliminates cache contention, preventing multiple cores from fighting over the same memory segments, and keeps the underlying logic fully portable and optimal across both GPU and CPU. 𝗕𝗹𝗮𝗰𝗸𝘄𝗲𝗹𝗹 𝗢𝗽𝘁𝗶𝗺𝗶𝘇𝗮𝘁𝗶𝗼𝗻 On the high-end GPU side, this release adds support for the Tensor Memory Accelerator (TMA) and inlined PTX for manual Matrix-Multiply Accumulate (MMA) instructions. This allows us to get closer to the theoretical peak of modern silicon. We’ve adapted our matmul engine to combine TMA with warp specialization, specifically targeting Blackwell-based hardware like the RTX 5090. These improvements also benefit NVIDIA’s Ada and Hopper architectures. New benchmarks show our kernels reaching state-of-the-art performance, matching the industry-standard CUTLASS and cuBLAS libraries found in LibTorch. This release also packs several other enhancements, ranging from zero-copy weight loading to a more streamlined training API. For a deep dive into all the new features and performance gains, check out the full release post here: https://lnkd.in/eQNrNXAP We’re excited to see what you build with these new capabilities. As always, feel free to reach out on Discord or GitHub with your feedback!

    • Benchmark: Max Pool 2D (Lower is Better)
  • 🔥 End of the Year Review and Burn Central Announcement This year, we made Burn and CubeCL even more reliable, faster, flexible, and multiplatform. A notable achievement has been showing that GPU and CPU programming can be abstracted within a single framework without sacrificing efficiency. By leveraging comptime to specialize kernels for specific plane and line sizes, CubeCL achieves high performance on both architectures. While we did work a lot on our high-performance compute engine, it is only one part of the story. As the year reaches its end, we are proud to introduce Burn Central, a cloud platform that brings these capabilities together. Whether you want to track your experiments, share them with others, manage models in production, use your own hardware, or scale with cloud GPUs, Burn Central makes it simple and efficient. We will release the Burn Central Alpha early next year and are opening a waiting list for early access: https://lnkd.in/e6gB4fpj. Our goal is to get feedback and continue improving the platform for the community. We are pushing the Rust ecosystem for AI and HPC further, and we look forward to 2026. For more information, read the full post here: https://lnkd.in/e3s66Ssz Happy holidays! ☃️ https://lnkd.in/egMN6W5N

    Burn Central - Alpha

    https://www.youtube.com/

  • C'est un plaisir de supporter l'écosystème IA de la Ville de Québec! C'est impressionnant ce que le Club d'Intelligence Artificielle de l'Université Laval a réussi à accomplir en mobilisant des centaines d'étudiants, bravo!

    Tracel Technologies rejoint à son tour l’aventure du Club d’Intelligence Artificielle de l’ULaval comme partenaire 🤖 💙 TraceL AI, c’est une start-up en IA qui développe Burn, un framework d’apprentissage profond qui permet de créer, d’entraîner et de déployer des modèles d’IA sur tout type de hardware. En gros : ils construisent les outils qui rendent l’IA puissante et utilisable dans le monde réel🔥 Et le plus cool dans tout ça? Ils viennent parrainer notre projet Drone 🚁 On va pouvoir pousser encore plus loin l’IA embarquée, les tests en vrai terrain et toutes les idées un peu folles qu’on avait sur le tableau depuis longtemps. Un énorme merci à Tracel Technologies de croire en notre club et en nos projets 💙 #CIA #IA #TraceLAI #ULaval ⁠#Partenariat #Drone

    • No alternative text description for this image
  • Tracel Technologies reposted this

    Tracel Technologies rejoint à son tour l’aventure du Club d’Intelligence Artificielle de l’ULaval comme partenaire 🤖 💙 TraceL AI, c’est une start-up en IA qui développe Burn, un framework d’apprentissage profond qui permet de créer, d’entraîner et de déployer des modèles d’IA sur tout type de hardware. En gros : ils construisent les outils qui rendent l’IA puissante et utilisable dans le monde réel🔥 Et le plus cool dans tout ça? Ils viennent parrainer notre projet Drone 🚁 On va pouvoir pousser encore plus loin l’IA embarquée, les tests en vrai terrain et toutes les idées un peu folles qu’on avait sur le tableau depuis longtemps. Un énorme merci à Tracel Technologies de croire en notre club et en nos projets 💙 #CIA #IA #TraceLAI #ULaval ⁠#Partenariat #Drone

    • No alternative text description for this image
  • 🔥 Burn 0.19.0 is out! This release brings us closer to the goals we set for 2024: large-scale distributed training and quantized model deployment. 𝗗𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝗧𝗿𝗮𝗶𝗻𝗶𝗻𝗴 To make true multi-GPU parallelism possible, we reimagined several core systems: - Multi-Stream Execution: concurrent compute and data transfers on the same GPU. - Redesigned Locking: a global device lock shared across CubeCL, fusion, and autotuning—deadlock-free by design. - Distributed Infrastructure: burn-collective now powers gradient synchronization and multi-device training. 𝗤𝘂𝗮𝗻𝘁𝗶𝘇𝗮𝘁𝗶𝗼𝗻 Quantization and persistent memory optimization now allow Burn to run large models with dramatically reduced memory usage. On a LLAMA-1B model, memory dropped from 4.37 GB (no persistent memory, FP16) to 3.01 GB with persistent memory alone. When combined with quantization, the footprint shrinks even further — to 1.51 GB using q8t tensor quantization, 758 MB with q4t, 851 MB with q4b32 block quantization, and just 570 MB using q2b16.  𝗡𝗲𝘄 𝗖𝗣𝗨 𝗕𝗮𝗰𝗸𝗲𝗻𝗱 (𝗟𝗟𝗩𝗠 + 𝗠𝗟𝗜𝗥) We’ve also introduced a CPU backend with JIT compilation, autotuning, and fusion, powered by LLVM and MLIR — bringing CubeCL’s GPU-level capabilities to CPU execution. Fun fact: this effectively makes CubeCL an alternative Rust compiler specialized for tensor workloads. This release is one of our biggest steps yet toward unifying deep learning across devices, data types, and scales — all in pure Rust. Read the full post and migration guide here 👇 🔗 https://lnkd.in/euKpfGXe

  • Tracel Technologies reposted this

    Cette année nous avons reçu 50 000 $ de la Ville de Québec afin de propulser notre entreprise. Le soutien de la Ville va nous permettre de soutenir le lancement de notre produit commercial Burn-Central! Stay tuned 🔥 Informez-vous dès maintenant: https://lnkd.in/eUPNetTF  Faites comme nous, soumettez votre projet au Défi-Québec d'ici le mercredi 29 octobre 2025.

    • No alternative text description for this image
  • Nathaniel Simard made a talk at RustConf last month where he showed why Rust is an ideal choice for accelerated computing and AI development. The talk dives into the design of Burn and CubeCL, demonstrating how Rust’s robust type system and ownership model enable flexible, high-performance, and portable solutions for cutting-edge AI applications. RustConf 2025 was a fantastic event filled with interesting talks and people. See you there next year in Montreal! Link: https://lnkd.in/e-UNv6fy #rustconf #rustlang #rustconf25

Similar pages

Browse jobs