Stars
An end-to-end production workspace for AI-generated short dramas. From script input to structured storyboarding, consistency management, shot preparation, video generation, and export.
From Pipeline to Council — A lightweight agentic framework for search & recommendation. 五个 Agent 重构搜广推管线。
[SIGIR 2024 perspective] The implementation of paper "On Generative Agents in Recommendation"
GPT-Image-2 PPT Generator Skill for Creating Image-Based PowerPoint Presentations in Codex and Other Skill-Compatible Agents
A lightweight, flexible, and easy-to-use Python tool for generating and exporting CapCut drafts to build fully automated video editing/remix pipelines! Another similar project: https://github.com/G…
The open-source CapCut alternative
Experimental implementation of DeepSeek v4 flaash in llama.cpp
Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local
The official Lark/Feishu CLI tool, maintained by the larksuite team — built for humans and AI Agents. Covers core business domains including Messenger, Docs, Base, Sheets, Calendar, Mail, Tasks, Me…
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with C…
分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
how to optimize some algorithm in cuda.
About Awesome things towards foundation agents. Papers / Repos / Blogs / ...
A collection of research and application papers of (uncertainty) calibration techniques.
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
A high-throughput and memory-efficient inference and serving engine for LLMs
📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉
Fast and memory-efficient exact attention
Awesome-LLM-KV-Cache: A curated list of 📙Awesome LLM KV Cache Papers with Codes.
This repository serves as a comprehensive survey of LLM development, featuring numerous research papers along with their corresponding code links.
[ICLR2025 Spotlight🔥] Official Implementation of TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
A book for Learning the Foundations of LLMs
Learning Large Language Model (LLM)(大语言模型学习)
🔥🔥🔥 Latest Advances on Large Recommendation Models
MoBA: Mixture of Block Attention for Long-Context LLMs
An easy PyTorch implementation of "Stabilizing Transformers for Reinforcement Learning"