Lists (2)
Sort Name ascending (A-Z)
Stars
A curated list of papers and selected technical blogs on Loop Models.
This is the code repository of paper "SFT Conflicts, RL Coexists"
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.
AlphaGo Moment for Model Architecture Discovery.
A pi extension with a single tool: execute, running TypeScript in a persistent Bun evaluator. Files, shell, and subagents are expressed as code rather than as more tools.
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.
LeanDojo-v2 is an end-to-end framework for training, evaluating, and deploying AI-assisted theorem provers for Lean 4.
🧘 BLISS – a Benchmark for Language Induction from Small Sets
Official PyTorch implementation of "Latent Reasoning in TRMs is Secretly a Policy Improvement Operator" (ICML 2026)
Train Large Language Models on MLX.
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
This repo contains the code for our paper "From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"
Unified model-distillation harness for code abilities: pi / claude-code / codex teachers, configurable judge, verified agentic traces
This repository consists of python scripts for LLM finetuning (SFT, LoRA, QLoRA) and LLM synthetic data generation scripts.
Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
Impact of typos and common misspellings on LLM task performance.
Official code for the paper "Sparsely gated tiny linear experts"
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting