Lists (3)
Sort Name ascending (A-Z)
Stars
This repo is for code that supports talks, blogs, and other activities the Google Cloud Developer Relations team engages in.
A gallery that showcases on-device ML/GenAI use cases and allows people to try and use models locally.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex β developed and maintained with no human intervention.
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
PyTorch building blocks for the OLMo ecosystem
Modeling, training, eval, and inference code for OLMo
A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
250+ Fine-tuning & RL Notebooks for text, vision, audio, embedding, TTS models.
"RAG-Anything: All-in-One RAG Framework"
The official implementation for [NeurIPS2025 Oral] Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode
This repository is a read-only mirror of https://gitlab.arm.com/kleidi/kleidiai
Official PyTorch implementation for Hogwild! Inference: Parallel LLM Generation with a Concurrent Attention Cache
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrβ¦
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning
NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
SGLang is a high-performance serving framework for large language models and multimodal models.
Paper list for Personal LLM Agents
Official code repo for the O'Reilly Book - "Hands-On Large Language Models"