Lists (4)
Sort Name ascending (A-Z)
Stars
A suite of plugins for legal workflows
[ICML2026] From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
🦋 An Infographic Generation and Rendering Framework, bring words to life with AI!
Create beautiful slides on the web using a coding agent's frontend skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Talk with Reachy Mini!
An open-source simulator framework for neural processing units
A collection of various llm pruning implementations, training code for GPUs & TPUs, and evaluation script.
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
🤗 smolagents: a barebones library for agents that think in code.
Official repository for the paper VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices
Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.
A framework for efficient model inference with omni-modality models
FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Trae, Traycer AI…
Examples and tutorials to help developers build AI systems
A step by step implementation of a complex RAG pipeline to solve real world situations
LiteRT and LiteRT-LM sample apps, model recipes, agent skills and utilities.
Maximizing the Performance of a Simple RAG using RL
Implementation of a GPT-4o like Multimodal from Scratch using Python
In this blog, we will build a small scale text-to-video model from scratch. We will input a text prompt, and our trained model will generate a video based on that prompt.
A straightforward method for training your LLM, from downloading data to generating text.
Building LLaMA 4 MoE from Scratch
Building DeepSeek R1 from Scratch
Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
Fast and memory-efficient exact attention
DeepSeek Coder: Let the Code Write Itself