Stars
Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Curso…
🚀 A curated collection of papers focusing on LLM-based quantitative trading.
Official JAX implementation of End-to-End Test-Time Training for Long Context
武汉大学博士/硕士学位论文latex模板(包含插图索引、表格索引、中英文缩略语对照、主要符号表等,字体格式以及排版修正)
Lightweight coding agent that runs in your terminal
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
An agent-first mobile runtime built on OpenClaw: control apps, learn reusable skills, and connect phones.
A general memory system for agents, powered by deep-research
Official implementation of Accelerating Prefilling for Long-Context Inference via Sparse Pattern Sharing
Unofficial implementations of block/layer-wise pruning methods for LLMs.
Official implementation for LaCo (EMNLP 2024 Findings)
For releasing code related to compression methods for transformers, accompanying our publications
[ICML 2024] Official Implementation of SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Official Repo of "MMBench: Is Your Multi-modal Model an All-around Player?"
🚀 A curated list of awesome resources focusing on Context Compression techniques for Large Language Models(LLMs).
CUDA Python: Performance meets Productivity
[COLM 2025] Official PyTorch implementation of "Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models"
This is the official code for ZLST-Project, Generative Recommendation Benchmark
Official code implementation of Context Cascade Compression: Exploring the Upper Limits of Text Compression
Official repository of RARE: Retrieval-Augmented Reasoning Modeling [KDD 2026 Research Track]
A character-level language diffusion model trained on Tiny Shakespeare
Official PyTorch implementation for "Large Language Diffusion Models"
A quick rundown on each feature and its settings
[ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.