Skip to content
View wdlctc's full-sized avatar

Block or report wdlctc

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the BrowseComp and BrowseComp Zh, respectively.

Python 8,366 643 Updated Jul 6, 2026
Jupyter Notebook 30 Updated Jan 16, 2025

The official implementation for [NeurIPS2025 Oral] Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free

Jupyter Notebook 975 61 Updated Dec 20, 2025

A unified inference and post-training framework for accelerated video generation.

Python 3,936 398 Updated Aug 10, 2026

Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)

Python 3,476 280 Updated Sep 12, 2025

MAGI-1: Autoregressive Video Generation at Scale

Python 3,761 238 Updated Jun 17, 2026

A pipeline parallel training script for diffusion models.

Python 2,004 282 Updated Aug 9, 2026

A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training

Python 904 65 Updated Aug 10, 2026

minimal GRPO implementation from scratch

Python 104 13 Updated Mar 14, 2025

Minimal reproduction of DeepSeek R1-Zero

Python 13,224 1,579 Updated Feb 27, 2026

A Lossless Compression Library for AI pipelines

Python 327 39 Updated Apr 11, 2026

📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥

2,160 102 Updated Aug 6, 2026
Python 17 2 Updated Aug 1, 2025

A Python library transfers PyTorch tensors between CPU and NVMe

C++ 125 28 Updated Nov 27, 2024

Mini versions of GPT2, LLama3, .. for pre-training

Python 3 Updated Jul 12, 2025

Everything about the SmolLM and SmolVLM family of models

Python 3,867 303 Updated May 26, 2026

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 73,974 9,050 Updated Aug 10, 2026
Python 63 14 Updated May 16, 2025

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Python 540 41 Updated Feb 10, 2025

[NeurIPS-2024] 📈 Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies https://arxiv.org/abs/2407.13623

Python 112 5 Updated Sep 26, 2024

LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.

Python 3,145 225 Updated May 19, 2025
Python 51 2 Updated Oct 29, 2024

OLMoE: Open Mixture-of-Experts Language Models

Jupyter Notebook 1,057 120 Updated Sep 23, 2025

Linear Attention Sequence Parallelism (LASP)

Python 87 6 Updated Jun 4, 2024

VideoSys: An easy and efficient system for video generation

Python 2,022 130 Updated Aug 27, 2025

Development repository for the Triton language and compiler

MLIR 19,921 3,093 Updated Aug 10, 2026

Latency and Memory Analysis of Transformer Models for Training and Inference

Python 491 59 Updated Apr 19, 2025

RTP: Rethinking Tensor Parallelism with Memory Deduplication

Python 11 Updated Dec 15, 2023
Next