Skip to content
View baofff's full-sized avatar

Block or report baofff

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Vidu S1: A Real-Time Interactive Video Generation Model

212 4 Updated Jul 13, 2026

A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models

Python 738 19 Updated Jun 15, 2026

TurboDiffusion: 100–200× Acceleration for Video Diffusion Models

Python 3,579 271 Updated Jul 16, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,501 447 Updated Jan 17, 2026

MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集。对标chatGPT训练的40T数据。MNBVC数据集不但包括主流文化,也包括各个小众文化甚至火星文的数据。MNBVC数据集包括新闻、作文、小说、书籍、杂志、论文、台词、帖子、wiki、古诗、歌词、商品介绍、笑话、糗事、聊天记录等一切形式的纯文本中文数据。

4,247 292 Updated Jul 13, 2026

NeurIPS 2025 Spotlight; ICLR2024 Spotlight; CVPR 2024; EMNLP 2024

Python 1,848 78 Updated Nov 27, 2025

Consistency Distilled Diff VAE

Python 2,213 82 Updated Nov 7, 2023

A framework for few-shot evaluation of language models.

Python 13,395 3,435 Updated Jul 13, 2026

Scaling Data-Constrained Language Models

Jupyter Notebook 344 18 Updated Jun 28, 2025

OpenLLaMA, a permissively licensed open source reproduction of Meta AI’s LLaMA 7B trained on the RedPajama dataset

7,531 404 Updated Jul 16, 2023

The RedPajama-Data repository contains code for preparing large datasets for training large language models.

Python 4,971 377 Updated Jun 3, 2026

Tools to download and cleanup Common Crawl data

Python 1,047 154 Updated Apr 25, 2023

A quick guide (especially) for trending instruction finetuning datasets

3,403 235 Updated Nov 28, 2023

Unofficial implementation for [ECCV'22] "Exploring Plain Vision Transformer Backbones for Object Detection"

Python 586 46 Updated Apr 24, 2022

The official GitHub page for the survey paper "A Survey of Large Language Models".

Python 12,194 934 Updated Mar 11, 2025

A series of large language models developed by Baichuan Intelligent Technology

Python 4,089 296 Updated Nov 8, 2024

[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | 基于CPM基础模型的中英双语多模态大模型系列

Python 1,063 88 Updated Jun 13, 2024

Diffusion model papers, survey, and taxonomy

3,362 258 Updated Sep 27, 2025

Generative Agents: Interactive Simulacra of Human Behavior

21,791 3,057 Updated Aug 5, 2024

[IJCV2024] Exploiting Diffusion Prior for Real-World Image Super-Resolution

Python 2,666 171 Updated Jul 12, 2024

Awesome-LLM: a curated list of Large Language Model

27,182 2,650 Updated Jul 31, 2025

Digital Human Resource: 2D/3D/4D Human Modeling, Avatar Generation & Animation, Clothed People Digitalization, Virtual Try-On, etc.

1,964 173 Updated Apr 18, 2026

OpenAI Baselines: high-quality implementations of reinforcement learning algorithms

Python 16,747 4,937 Updated Aug 1, 2024

An elegant PyTorch deep reinforcement learning library.

Python 10,881 1,331 Updated Apr 3, 2026

A curated list of Multimodal Related Research.

Python 1,393 147 Updated Aug 5, 2023

Reading list for research topics in multimodal machine learning

6,910 900 Updated Aug 20, 2024

A large-scale 7B pretraining language model developed by BaiChuan-Inc.

Python 5,650 501 Updated Jul 18, 2024

An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.

Python 39,501 4,786 Updated May 1, 2026

中国大模型

6,459 565 Updated Nov 30, 2024

A GPT-4 AI Tutor Prompt for customizable personalized learning experiences.

29,611 3,293 Updated Sep 30, 2025
Next