Stars
Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power.
Code for the NPJ AI paper "How Large Language Models Encode Theory-of-Mind: A Study on Sparse Parameter Patterns"
A PyTorch native platform for training generative AI models
Train transformer language models with reinforcement learning.
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
[EMNLP'25] Code for the paper "DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic"
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
PubMedQA: A Dataset for Biomedical Research Question Answering
Vortex: Programmable Sparse Attention for Agents as Algorithm Designers
🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 OpenAI, Anthropic, Gemini, Ollam…
Optimize prompts, code, and more with AI-powered Reflective Optimization
Open-source implementation of AlphaEvolve
[ICML 2026] Decoding Tree Sketching (DTS): a training-free & model agonistic & plug-in framework for LLM parallel reasoning.
This guide provides a setup for running `vLLM` with `Triton` on NVIDIA GH200.
RapidIn: Scalable Influence Estimation for Large Language Models (LLMs). The implementation for paper "Token-wise Influential Training Data Retrieval for Large Language Models" (Accepted on ACL 2024).
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
H-Net: Hierarchical Network with Dynamic Chunking
A lightweight python-only library for reading and writing SMILES strings
Code repository for ICLR 2025 paper "LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid"
Code Repo for EMNLP paper: Do LLMs Know to Respect Copyright Notice
A game theoretic approach to explain the output of any machine learning model.
[NeurIPS 2024] Simple and Effective Masked Diffusion Language Model
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
A novel medical large language model family with 13/70B parameters, which have SOTA performances on various medical tasks
FlashMLA: Efficient Multi-head Latent Attention Kernels
MoBA: Mixture of Block Attention for Long-Context LLMs