Skip to content
View winglian's full-sized avatar

Sponsors

@narrative-io

Highlights

  • Pro

Block or report winglian

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A curated list of papers and selected technical blogs on Loop Models.

Python 257 7 Updated Aug 11, 2026

This is the code repository of paper "SFT Conflicts, RL Coexists"

Python 29 2 Updated Aug 4, 2026

Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.

Python 437 40 Updated Aug 11, 2026

AlphaGo Moment for Model Architecture Discovery.

Python 1,180 220 Updated Dec 3, 2025

A pi extension with a single tool: execute, running TypeScript in a persistent Bun evaluator. Files, shell, and subagents are expressed as code rather than as more tools.

TypeScript 72 5 Updated Aug 9, 2026

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Python 29 9 Updated Aug 6, 2026

A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.

192 10 Updated Aug 9, 2026

LeanDojo-v2 is an end-to-end framework for training, evaluating, and deploying AI-assisted theorem provers for Lean 4.

Python 123 20 Updated Aug 10, 2026

🧘 BLISS – a Benchmark for Language Induction from Small Sets

Python 13 Updated Sep 10, 2023

Intern-S2-Mobius

Python 37 Updated Aug 6, 2026

Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)

Python 25 2 Updated May 29, 2026

Train Large Language Models on MLX.

Python 407 52 Updated Aug 10, 2026

🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support

Python 808 246 Updated Aug 12, 2026

This repo contains the code for our paper "From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"

2 Updated May 27, 2026
Python 3 Updated Aug 12, 2026
Python 10 Updated Jul 23, 2026

Unified model-distillation harness for code abilities: pi / claude-code / codex teachers, configurable judge, verified agentic traces

Python 12 1 Updated Aug 5, 2026
Python 15 1 Updated Jul 9, 2026

This repository consists of python scripts for LLM finetuning (SFT, LoRA, QLoRA) and LLM synthetic data generation scripts.

Python 4 1 Updated Aug 11, 2025

Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models

Python 45 10 Updated Jan 14, 2025

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

656 21 Updated Aug 9, 2026

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…

Python 635 56 Updated Aug 7, 2026
Python 350 35 Updated Jul 7, 2026

Impact of typos and common misspellings on LLM task performance.

Python 25 2 Updated Mar 22, 2024

Official code for the paper "Sparsely gated tiny linear experts"

Python 7 1 Updated Jul 31, 2026
Python 5 1 Updated Jun 8, 2026

Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.

Python 47 4 Updated Aug 11, 2026

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting

Python 176 6 Updated Aug 9, 2026
Next