- Silicon Valley, California
- https://leimao.github.io/
- @matchaleimao
- dukeleimao
Highlights
- Pro
Stars
Silvertorch is a high-performance open-source GPU retrieval engine designed for large-scale recommender systems, generative retrieval (RAG), and vector search — optimized for latency, throughput, a…
A library for efficient similarity search and clustering of dense vectors.
CUDA Tile IR is an MLIR-based intermediate representation and compiler infrastructure for CUDA kernel optimization, focusing on tile-based computation patterns and optimizations targeting NVIDIA te…
Facebook's branch of Apache Thrift, including a new C++ server.
Puzzles for learning Triton
This repository contains companion software for the Colfax Research paper "Categorical Foundations for CuTe Layouts".
AddressSanitizer, ThreadSanitizer, MemorySanitizer
The official Python SDK for Model Context Protocol servers and clients
C++ tensors with broadcasting and lazy computing
Deep Learning tools and applications for NVIDIA AGX platforms.
Training materials associated with NVIDIA's CUDA Training Series (www.olcf.ornl.gov/cuda-training-series/)
FlashInfer: Kernel Library for LLM Serving
CUDA Templates and Python DSLs for High-Performance Linear Algebra
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
[ECCV2022] PETR: Position Embedding Transformation for Multi-View 3D Object Detection & [ICCV2023] PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
Code examples for running V4L2 USB Cameras on NVIDIA Jetson Developer Kits
JetsonHacksNano / USB-Camera
Forked from jetsonhacks/USB-CameraCode examples for running V4L2 USB Cameras on NVIDIA Jetson Developer Kits
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
NVIDIA DLA-SW, the recipes and tools for running deep learning workloads on NVIDIA DLA cores for inference applications.
Generate diagrams from textual description
This is an online course where you can learn and master the skill of low-level performance analysis and tuning.
Waydroid uses a container-based approach to boot a full Android system on a regular GNU/Linux system like Ubuntu.