Lists (2)
Sort Name ascending (A-Z)
Stars
This repo contains the code for our paper "From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"
Unified model-distillation harness for code abilities: pi / claude-code / codex teachers, configurable judge, verified agentic traces
This repository consists of python scripts for LLM finetuning (SFT, LoRA, QLoRA) and LLM synthetic data generation scripts.
Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
Impact of typos and common misspellings on LLM task performance.
Official code for the paper "Sparsely gated tiny linear experts"
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting
the recursive language model (RLM) CLI agent
EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
[ICLR 2026] SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs
Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.