-
alibaba
- beijing
Stars
FlashMLA: Efficient Multi-head Latent Attention Kernels
SGLang is a high-performance serving framework for large language models and multimodal models.
✨✨Latest Advances on Multimodal Large Language Models
第一个支持中英文双语语音-文本多模态对话的开源可商用对话模型。便捷的语音输入将大幅改善以文本为输入的大模型的使用体验,同时避免了基于 ASR 解决方案的繁琐流程以及可能引入的错误。
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
[TMLR] A curated list of language modeling researches for code (and other software engineering activities), plus related datasets.
Large Language Model Text Generation Inference
Llama中文社区,实时汇总最新Llama学习资料,构建最好的中文Llama大模型开源生态,完全开源可商用
Semantic cache for LLMs. Fully integrated with LangChain and llama_index.
A community-driven way to read and chat with AI bots - powered by chatGPT.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Simple, light-weight and easy-to-use asynchronous components
brpc is an Industrial-grade RPC framework using C++ Language, which is often used in high performance system such as Search, Storage, Machine learning, Advertisement, Recommendation etc. "brpc" mea…
A cheatsheet of modern C++ language and library features.
A collection of resources on modern C++
Collection of various algorithms in mathematics, machine learning, computer science and physics implemented in C++ for educational purposes.
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search
✔️李沐 【动手学深度学习】课程学习笔记:使用pycharm编程,基于pytorch框架实现。
std::experimental::simd for GCC [ISO/IEC TS 19570:2018]
C++ Insights - See your source code with the eyes of a compiler
Scans for potential unported or non-portable code in source code trees.
A tool for use with clang to analyze #includes in C and C++ source files