Skip to content
View TKONIY's full-sized avatar
🌋
Working on Data x AI
🌋
Working on Data x AI

Organizations

@DBGroup-SUSTech

Block or report TKONIY

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Companion code for the global workspace interpretability paper

Python 1,560 221 Updated Jul 17, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,767 1,103 Updated Apr 30, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,656 197 Updated Jul 24, 2026

Kernels, of the mega variety :)

Python 786 64 Updated May 26, 2026

A high-performance, universal serving framework for any-to-any models.

Python 56 10 Updated Jul 24, 2026

A copy of the Cursor AI editor theme from cursor.sh, repackaged for use in Visual Studio Code without needing the Cursor editor.

37 4 Updated Sep 17, 2025

High-Throughput Batch Inference

C++ 13 1 Updated Jul 14, 2026

Homepage for ProLong (Princeton long-context language models) and paper "How to Train Long-Context Language Models (Effectively)"

Python 261 15 Updated Sep 12, 2025

ElasticMM: Elastic and Efficient MLLM Serving System

Python 44 2 Updated May 10, 2026

Can LLMs Write Correct and Efficient GPU Communication Code?

Python 47 2 Updated Jul 7, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,761 626 Updated Jul 9, 2026

Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.

Python 122 7 Updated Jul 17, 2026

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Python 37 2 Updated Jul 12, 2026

NeMo: a toolkit for conversational AI

Python 1 Updated Jul 21, 2026

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

C++ 13,182 2,389 Updated Jul 7, 2026
Python 222 22 Updated Jun 8, 2026

Sparse Decoupled Attention for Efficient Long-Context LLM Inference

Python 51 2 Updated Jun 4, 2026

An unofficial cuda assembler, for all generations of SASS, hopefully :)

Python 609 110 Updated Apr 20, 2023

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Python 1,655 257 Updated Apr 24, 2026

UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

C++ 1,470 164 Updated Jul 24, 2026

high-performance inference and serving library for interactive autoregressive video and world models

Python 416 40 Updated Jul 22, 2026

MapLibre GL JS - Interactive vector tile maps in the browser

TypeScript 11,133 1,148 Updated Jul 24, 2026

cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.

Rust 713 50 Updated Jul 23, 2026
Python 313 38 Updated Jun 9, 2026

Distributed Compiler based on Triton for Parallel Systems

Python 1,498 162 Updated Jul 20, 2026

A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.

Python 1,291 91 Updated Jul 14, 2026

A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models

Python 738 19 Updated Jun 15, 2026

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 25,981 2,037 Updated Jul 23, 2026
Next