Stars
Scaling the Horizon, Not the Parameters
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
A hybrid GPU cluster simulator for ML system performance estimation
TPU inference for vLLM, with unified JAX and PyTorch support.
Ultra-light Harness scaffolding for AI agents, a mini version of claude code
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
AI agents running research on single-GPU nanochat training automatically
A lightweight inference engine supporting speculative speculative decoding (SSD).
This is Official implementation for T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning
[CVPR 2026 Highlight] ForeAct: Steering Your VLA with Efficient Visual Foresight Planning
[ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
MLEvolve is an open-source autonomous system for end-to-end machine learning algorithm design and optimization powered by progressive search and experience-driven memory.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
Accelerate FLUX.2 inference from 18s to 12s (33% speedup) using SADA in H200
🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence…
[HPCA 2026 Best Paper Candidate] Official implementation of "Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models"
DFlash: Block Diffusion for Flash Speculative Decoding
MoBA: Mixture of Block Attention for Long-Context LLMs
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
cuTile is a programming model for writing parallel kernels for NVIDIA GPUs
Fast, memory-efficient attention column reduction (e.g., sum, mean, max)
[HPCA 2026] FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
[ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems
[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
a high-performance Block Sparse Attention kernel in Triton