Stars
你想蒸馏的下一个员工,何必是同事。蒸馏任何人的思维方式——心智模型、决策启发式、表达DNA。Distill how anyone thinks.
[ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models
Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
[EMNLP 2025 Oral] MemoryOS is designed to provide a memory operating system for personalized AI agents.
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
MOA is an open source framework for Big Data stream mining. It includes a collection of machine learning algorithms (classification, regression, clustering, outlier detection, concept drift detecti…
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
“百聆”是一个基于LLaMA的语言对齐增强的英语/中文大语言模型,具有优越的英语/中文能力,在多语言和通用任务等多项测试中取得ChatGPT 90%的性能。BayLing is an English/Chinese LLM equipped with advanced language alignment, showing superior capability in English/Ch…
A book for Learning the Foundations of LLMs
Translation models for 22 scheduled languages of India
A library for preparing data for machine translation research (monolingual preprocessing, bitext mining, etc.) built by the FAIR NLLB team.
Repository accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)
[EMNLP 2023] 💬 Language Identification with Support for More Than 2000 Labels
This project use the Meta NLLB-200 translation model through the Hugging Face transformers library.
Geographically-informed language identification
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Awesome-Multilingual-LLMs-Papers
Retrieval and Retrieval-augmented LLMs
The most accurate natural language detection library for Java and the JVM, suitable for long and short text alike
This code provides word level language identification tool for identifying language for individual words in Code-Mixed text. e.g. The text that includes words from two languages such as Hindi writt…
Mobile-Agent: The Powerful GUI Agent Family
BlueLM(蓝心大模型): Open large language models developed by vivo AI Lab
[ACL 2023] One Embedder, Any Task: Instruction-Finetuned Text Embeddings
Source code for AAAI 2022 paper: Unified Named Entity Recognition as Word-Word Relation Classification
ModelScope: bring the notion of Model-as-a-Service to life.
An Open-Source Package for Neural Relation Extraction (NRE)
AdaSeq: An All-in-One Library for Developing State-of-the-Art Sequence Understanding Models