Skip to content
View winglian's full-sized avatar

Sponsors

@narrative-io

Highlights

  • Pro

Block or report winglian

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

This repo contains the code for our paper "From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"

2 Updated May 27, 2026
Python 3 Updated Jul 29, 2026
Python 9 Updated Jul 23, 2026

Unified model-distillation harness for code abilities: pi / claude-code / codex teachers, configurable judge, verified agentic traces

Python 10 1 Updated Jul 29, 2026
Python 15 1 Updated Jul 9, 2026

This repository consists of python scripts for LLM finetuning (SFT, LoRA, QLoRA) and LLM synthetic data generation scripts.

Python 4 1 Updated Aug 11, 2025

Generative AI-based CyberSecurity-focused Prompt Dataset for Benchmarking Large Language Models

Python 45 10 Updated Jan 14, 2025

A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

583 20 Updated Jul 29, 2026

LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…

Python 607 56 Updated Jul 7, 2026
Python 334 32 Updated Jul 7, 2026

Impact of typos and common misspellings on LLM task performance.

Python 25 2 Updated Mar 22, 2024

Official code for the paper "Sparsely gated tiny linear experts"

Python 7 1 Updated Jul 6, 2026
Python 5 1 Updated Jun 8, 2026

Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.

Python 45 4 Updated Jul 24, 2026

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Causal Parallel Tree Drafting

Python 171 6 Updated Jun 27, 2026

the recursive language model (RLM) CLI agent

Python 199 12 Updated Jul 16, 2026

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

Python 16 Updated Jul 27, 2026

Tile-Based Runtime for Ultra-Low-Latency LLM Inference

Python 1,600 112 Updated Jul 14, 2026
Python 4 2 Updated Jul 29, 2026

[ICLR 2026] SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs

Python 138 16 Updated May 20, 2026

Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs

Python 150 9 Updated Jun 28, 2026
Python 14 1 Updated May 24, 2026

Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x

Python 148 10 Updated Jun 10, 2026
Python 26 1 Updated May 25, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 235 24 Updated Jul 30, 2026
Python 15 7 Updated Feb 5, 2026

HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.

Python 1,759 167 Updated Jun 17, 2026

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

JavaScript 235,586 35,879 Updated Jul 29, 2026
Next