Skip to content
View winglian's full-sized avatar

Sponsors

@narrative-io

Highlights

  • Pro

Block or report winglian

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)

Python 25 2 Updated May 29, 2026

Train Large Language Models on MLX.

Python 407 53 Updated Jul 21, 2026

🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support

Python 796 243 Updated Aug 6, 2026

This repo contains the code for our paper "From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"

2 Updated May 27, 2026
Python 9 Updated Jul 23, 2026

Unified model-distillation harness for code abilities: pi / claude-code / codex teachers, configurable judge, verified agentic traces

Python 11 1 Updated Aug 5, 2026
Python 15 1 Updated Jul 9, 2026

This repository consists of python scripts for LLM finetuning (SFT, LoRA, QLoRA) and LLM synthetic data generation scripts.

Python 4 1 Updated Aug 11, 2025

Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models

Python 45 10 Updated Jan 14, 2025

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

629 20 Updated Aug 6, 2026

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…

Python 625 56 Updated Jul 7, 2026
Python 348 34 Updated Jul 7, 2026

Impact of typos and common misspellings on LLM task performance.

Python 25 2 Updated Mar 22, 2024

Official code for the paper "Sparsely gated tiny linear experts"

Python 7 1 Updated Jul 31, 2026
Python 5 1 Updated Jun 8, 2026

Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.

Python 46 4 Updated Aug 5, 2026

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting

Python 173 6 Updated Jul 31, 2026

the recursive language model (RLM) CLI agent

Python 200 13 Updated Jul 16, 2026

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

Python 16 Updated Jul 27, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,621 114 Updated Aug 6, 2026
Python 4 2 Updated Jul 30, 2026

[ICLR 2026] SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs

Python 138 16 Updated Aug 2, 2026

Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs

Python 151 9 Updated Aug 4, 2026
Python 14 1 Updated May 24, 2026

Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x

Python 150 10 Updated Jun 10, 2026
Python 26 1 Updated May 25, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 238 24 Updated Aug 6, 2026
Next