Skip to content
View HankYe's full-sized avatar
  • Duke University
  • Durham, NC

Block or report HankYe

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Scaling the Horizon, Not the Parameters

Python 537 50 Updated Jul 16, 2026

Preview Code for MARS Paper

Python 6 1 Updated Jun 5, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,207 693 Updated Aug 13, 2026

A hybrid GPU cluster simulator for ML system performance estimation

Rust 42 10 Updated Jul 15, 2026

TPU inference for vLLM, with unified JAX and PyTorch support.

Python 406 284 Updated Aug 13, 2026

[DAC 2026] FlashFPS

Python 15 Updated Jun 1, 2026

Ultra-light Harness scaffolding for AI agents, a mini version of claude code

Python 953 358 Updated Jun 10, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,070 109,175 Updated Aug 6, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,787 13,300 Updated Mar 26, 2026
Python 115 9 Updated Mar 14, 2026

A lightweight inference engine supporting speculative speculative decoding (SSD).

Python 987 78 Updated May 10, 2026

This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning

Python 24 1 Updated Mar 5, 2026

[CVPR 2026 Highlight] ForeAct: Steering Your VLA with Efficient Visual Foresight Planning

Python 86 5 Updated May 1, 2026

[ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

Python 328 33 Updated Jul 1, 2026

MLEvolve is an open-source autonomous system for end-to-end machine learning algorithm design and optimization powered by progressive search and experience-driven memory.

Python 417 58 Updated Jul 14, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,165 81,167 Updated Aug 13, 2026

Accelerate FLUX.2 inference from 18s to 12s (33% speedup) using SADA in H200

Python 10 Updated Feb 3, 2026

🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence…

TypeScript 67,760 8,112 Updated Aug 13, 2026

[HPCA 2026 Best Paper Candidate] Official implementation of "Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models"

Python 61 8 Updated Feb 8, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

Python 5,612 403 Updated May 10, 2026

MoBA: Mixture of Block Attention for Long-Context LLMs

Python 2,162 156 Updated Apr 3, 2025
Python 9 Updated Dec 30, 2025

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,742 787 Updated May 17, 2026

cuTile is a programming model for writing parallel kernels for NVIDIA GPUs

Python 2,126 143 Updated Aug 12, 2026

Fast, memory-efficient attention column reduction (e.g., sum, mean, max)

Python 50 3 Updated Feb 10, 2026

[HPCA 2026] FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing

Python 22 4 Updated Apr 21, 2026

[ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems

Python 1,085 165 Updated Jun 18, 2026

[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"

Python 428 57 Updated Mar 13, 2026

a high-performance Block Sparse Attention kernel in Triton

Python 4 Updated Nov 14, 2025
Next