Stars
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
The Generative AI Landscape - A Collection of Awesome Generative AI Applications
Use ChatGPT to summarize the arXiv papers. 全流程加速科研,利用chatgpt进行论文全文总结+专业翻译+润色+审稿+审稿回复
Papers from the computer science community to read and discuss.
Apache Pinot - A realtime distributed OLAP datastore
Jeff Dean's latency numbers plotted over time
Productive, portable, and performant GPU programming in Python.
The Fastest Distributed Database for Transactional, Analytical, and AI Workloads.
⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)
tlaplus / azure-cosmos-tla
Forked from Azure/azure-cosmos-tlaAzure Cosmos TLA+ specifications
Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics
MindSpore is a new open source deep learning training/inference framework that could be used for mobile, edge and cloud scenarios.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Fast SHAP value computation for interpreting tree-based models
由图灵的猫开发,基于开源GPT2.0的初代创作型人工智能 | 可扩展、可进化
DNN connection pruning with Alternating Direction Method of Multipliers (ADMM)
An Introduction to Statistical Learning (James, Witten, Hastie, Tibshirani, 2013): Python code
Implementations of some Android Auto features as unofficial IDrive apps