Reconstruct 3D point clouds from images using this zero-dependency CUDA inference engine for the DVLT model.
-
Updated
Sep 25, 2026 - Cuda
Reconstruct 3D point clouds from images using this zero-dependency CUDA inference engine for the DVLT model.
🧠 Generate dynamic narratives with LMD, a framework for episodic memory that evolves, feels, and tells stories through creative language grounding.
Radio-Frequency Engineering Modeling Toolkit (RF-EMT). This project provides a high-fidelity simulation framework for Radar,Telecommunication and the other Radio-frequency engineering systems. The goal is to achieve realistic, design-level system modeling and simulation.
🚀 Optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels using reinforcement learning, surpassing cuBLAS and other benchmarks with superior performance.
CUDA Core Compute Libraries
LLM speculative inference server for heterogeneous hardware & consumer GPUs
Custom CUDA GEMM kernels (tiled + TensorCore WMMA) as a PyTorch extension for transformer inference
Crazy Fast Qwen3.8–27b With a 420k Context on an RTX 5090
Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
Minimalistic llm-inference tool in Rust
Safe rust wrapper around CUDA toolkit
OpenQASM VS Code Extension
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Numba based GPU block matching
Physical AI inference runtime: deadline-aware, memory-bounded inference for streaming Vision-Language-Action (VLA) models on edge GPUs. Hand-written CUDA/C++ transformer (RMSNorm, RoPE, GQA, SwiGLU) under a hard 33ms control-loop deadline and 8GB VRAM ceiling, with lock-free async perception and semantic KV-cache eviction.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
Integrate Hermes Agent skills directly into Claude Code for native agentic operations without external processes or locking.
A C++ Autograd Engine written without ANY external frameworks. It can train a pretty big GPT locally. Built to maximize performance.
Qwen3.8-27B on a single RTX 4090 (sm_89): Cinference/NInfer port with MTP speculative decoding and K8V4 KV cache
To associate your repository with the cuda-kernels topic, visit your repo's landing page and select "manage topics."