Skip to content
View qiugen's full-sized avatar

Block or report qiugen

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Python 2,151 250 Updated Aug 14, 2026

C++ port of ZXing

C++ 1,965 554 Updated Aug 13, 2026
Python 30 5 Updated Jun 30, 2025
Python 33 7 Updated Mar 13, 2024

Fully open reproduction of DeepSeek-R1

Python 26,431 2,446 Updated Apr 2, 2026

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 74,095 9,067 Updated Aug 13, 2026

记录本人整理的一些数据集

1,092 135 Updated Jun 16, 2022

All-in-one text de-duplication

Python 765 79 Updated Mar 9, 2026

TencentLLMEval is a comprehensive and extensive benchmark for artificial evaluation of large models that includes task trees, standards, data verification methods, and more.

41 1 Updated Mar 16, 2025

The RedPajama-Data repository contains code for preparing large datasets for training large language models.

Python 4,977 375 Updated Jun 3, 2026

Convert WIKI dumped XML (Chinese) to human readable documents in markdown and txt.

Python 8 2 Updated Mar 25, 2020

A tool for extracting plain text from Wikipedia dumps

Python 3,997 1,003 Updated Aug 10, 2026

This is a repository using the Wiki Extractor to build and prepare WIKIPEDIA for use in tensorflow.

Python 1 Updated Jul 21, 2018

We release a dataset based on Wikipedia sentences and the corresponding translations in 6 different languages along with the scores (scale 1 to 100) generated though human evaluations that represen…

81 14 Updated Aug 31, 2021

An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.

Jupyter Notebook 2,012 315 Updated Aug 9, 2025

Ongoing research training transformer language models at scale, including: BERT & GPT-2

Python 2,259 369 Updated Aug 14, 2025

活字通用大模型

Python 393 26 Updated Sep 12, 2024

沉浸式双语网页翻译扩展 , 支持输入框翻译, 鼠标悬停翻译, PDF, Epub, 字幕文件, TXT 文件翻译 - Immersive Dual Web Page Translation Extension

18,462 1,096 Updated Aug 14, 2026

TigerBot: A multi-language multi-task LLM

Python 2,259 189 Updated Dec 28, 2024

12306 订票程序,自动登录,自动下单

Python 27 10 Updated Jan 17, 2023

The agent engineering platform.

Python 144,266 24,024 Updated Aug 14, 2026

BELLE: Be Everyone's Large Language model Engine(开源中文对话大模型)

HTML 8,279 756 Updated Oct 16, 2024

Aligning pretrained language models with instruction data generated by themselves.

Python 1 Updated Mar 10, 2023

Personal short implementations of Machine Learning papers

Jupyter Notebook 251 56 Updated Jan 6, 2024

A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)

Python 4,753 486 Updated Jan 8, 2024

天涯 kkndme 神贴聊房价

19,422 3,845 Updated Jun 4, 2026

Code for "Learning to summarize from human feedback"

Python 1,062 153 Updated Sep 5, 2023

Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"

1,857 160 Updated Jun 17, 2025

ONNX Model Exporter for PaddlePaddle

C++ 940 196 Updated Mar 18, 2026
Next