-
University of Edinburgh
- Shenzhen, China
-
18:11
(UTC +08:00) - https://dengyangshen.netlify.app
Highlights
Lists (5)
Sort Name ascending (A-Z)
Starred repositories
Companion code for the global workspace interpretability paper
FlashMLA: Efficient Multi-head Latent Attention Kernels
TokenSpeed is a speed-of-light LLM inference engine.
A high-performance, universal serving framework for any-to-any models.
A copy of the Cursor AI editor theme from cursor.sh, repackaged for use in Visual Studio Code without needing the Cursor editor.
Homepage for ProLong (Princeton long-context language models) and paper "How to Train Long-Context Language Models (Effectively)"
ElasticMM: Elastic and Efficient MLLM Serving System
Can LLMs Write Correct and Efficient GPU Communication Code?
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.
TurboServe: Serving Streaming Video Generation Efficiently and Economically
vklimkov-nvidia / Speech
Forked from erastorgueva-nv/NeMoNeMo: a toolkit for conversational AI
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
Sparse Decoupled Attention for Efficient Long-Context LLM Inference
An unofficial cuda assembler, for all generations of SASS, hopefully :)
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
high-performance inference and serving library for interactive autoregressive video and world models
MapLibre GL JS - Interactive vector tile maps in the browser
cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
Distributed Compiler based on Triton for Parallel Systems
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone