Skip to content
View litetoooooom's full-sized avatar
💭
I may be slow to respond.
💭
I may be slow to respond.

Block or report litetoooooom

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

论文阅读助手

Python 2 Updated Jan 11, 2024

Using GPT to parse PDF

Python 3,563 260 Updated Apr 17, 2025
Jupyter Notebook 650 80 Updated Jun 2, 2026

Awesome-LLM-RAG: a curated list of advanced retrieval augmented generation (RAG) in Large Language Models

1,343 94 Updated Jul 22, 2026

A Survey of Attributions for Large Language Models

229 9 Updated Jan 14, 2026

This is a Github repository that focuses on articles related to skill-based meta reinforcement learning. The main focus is on skill extraction, combination, and generalization.

7 Updated Nov 3, 2023

论文阅读助手

Mermaid 1 Updated Sep 2, 2023

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

4,261 296 Updated Aug 15, 2026

整理开源的中文大语言模型,以规模较小、可私有化部署、训练成本较低的模型为主,包括底座模型,垂直领域微调及应用,数据集与教程等。

22,738 2,134 Updated May 10, 2026

FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.

Python 3,870 416 Updated Jul 13, 2026

RoFormer升级版

Python 153 15 Updated Aug 11, 2022

本文是Michael Nielson所著的《Neural Networks and Deep Learning》的简体中文翻译版。

93 13 Updated Jan 5, 2023
Python 82 12 Updated Jul 30, 2024

Official implementation of the papers "GECToR – Grammatical Error Correction: Tag, Not Rewrite" (BEA-20) and "Text Simplification by Tagging" (BEA-21)

Python 972 220 Updated May 21, 2024

3000000+语义理解与匹配数据集。可用于无监督对比学习、半监督学习等构建中文领域效果最好的预训练模型

Python 313 39 Updated Oct 11, 2022

Pre-Training with Whole Word Masking for Chinese BERT(中文BERT-wwm系列模型)

Python 10,221 1,381 Updated Apr 19, 2026
JavaScript 436 62 Updated Apr 25, 2025

中文分词 词性标注 命名实体识别 依存句法分析 成分句法分析 语义依存分析 语义角色标注 指代消解 风格转换 语义相似度 新词发现 关键词短语提取 自动摘要 文本分类聚类 拼音简繁转换 自然语言处理

Python 36,477 10,923 Updated Nov 15, 2025

自然语言处理,知识图谱相关语料。按照Task细分,欢迎PR。

Python 734 152 Updated Jan 15, 2021

A Chinese EHR Bert Pretrained Model.

Python 270 43 Updated Jul 14, 2021

使用深度学习方法解析问题 知识图谱存储 查询知识点 基于医疗垂直领域的对话系统

Python 792 212 Updated Sep 7, 2019

Fit data to many distributions

Python 412 60 Updated Mar 7, 2026

A LITE BERT FOR SELF-SUPERVISED LEARNING OF LANGUAGE REPRESENTATIONS, 海量中文预训练ALBERT模型

Python 3,982 742 Updated Nov 21, 2022
Python 1 Updated Apr 9, 2019

Command-line utility to transcribe/translate from video/audio/subtitles to subtitles

Python 1,960 234 Updated Dec 21, 2023

trie树实现多类型匹配

Python 5 Updated Apr 17, 2019

Never see escaped bytes in output.

Python 157 16 Updated Apr 5, 2022

PKU Team Zero's code for participation in ICDAR2019 ArT Recognition track (Champion)

Roff 225 66 Updated Oct 12, 2022
Next