🤖️ Optimized CUDA Kernels for Fast MobileNetV2 Inference
-
Updated
Dec 28, 2021 - Cuda
🤖️ Optimized CUDA Kernels for Fast MobileNetV2 Inference
Jetson Orin-tuned LLM inference runtime for gemma 4, qwen 3.5 — memory-first, power-aware, zero-allocation. C++17 + CUDA.
Add a description, image, and links to the inference-optimization topic page so that developers can more easily learn about it.
To associate your repository with the inference-optimization topic, visit your repo's landing page and select "manage topics."