Skip to content
View NOrangeeroli's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report NOrangeeroli

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Implementation of Poly-attention, a higher-order self-attention proposed by Chakrabarti et al. of Columbia

Python 54 5 Updated Jul 22, 2026

A collection of weight space learning including papers, codes, and datasets.

85 6 Updated May 12, 2026

Awesome papers on weight-space learning

38 3 Updated Jan 16, 2026

[ICLR 2025] NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG Signals

Python 165 22 Updated Sep 19, 2025

Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs

Python 151 9 Updated Aug 12, 2026
Python 37 Updated Jul 2, 2026

Official repository of PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

Python 28 1 Updated Jun 13, 2026

The Cloud Sandbox Built for AI Agents

Python 1,144 59 Updated Jun 8, 2026

Transform geospatial relations into graphs for Graph Neural Networks and spatial network analysis

Python 1,386 147 Updated Aug 11, 2026

A clean implementation based on AlphaZero for any game in any framework + tutorial + Othello/Gobang/TicTacToe/Connect4 and more

Jupyter Notebook 4,498 1,153 Updated Jan 1, 2025

A Survey of Reinforcement Learning for Large Reasoning Models

TeX 2,478 132 Updated Aug 1, 2026

A curated list of awesome exploration RL resources (continually updated)

720 26 Updated May 21, 2026
Slash 554 164 Updated Jul 1, 2023
Python 76 16 Updated Feb 17, 2022

Muon is an optimizer for hidden layers in neural networks

Python 2,779 128 Updated May 24, 2026

A MemAgent framework that can be extrapolated to 3.5M, along with a training framework for RL training of any agent workflow.

Python 1,092 74 Updated May 12, 2026

Awesome In-Context RL: A curated list of In-Context Reinforcement Learning - - —

306 15 Updated Sep 8, 2025

Resources for the Enigmata Project.

Python 83 7 Updated Aug 13, 2025

[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards

Python 1,480 123 Updated Apr 17, 2026

Understanding R1-Zero-Like Training: A Critical Perspective

Python 1,270 61 Updated Aug 27, 2025

Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ.

Python 1,809 92 Updated Feb 4, 2026

Reproduce R1 Zero on Logic Puzzle

Python 2,450 163 Updated Mar 20, 2025

minimal-cost for training 0.5B R1-Zero

Python 817 103 Updated May 14, 2025

Simple RL training for reasoning

Python 3,870 285 Updated Dec 23, 2025

Fully open reproduction of DeepSeek-R1

Python 26,433 2,445 Updated Apr 2, 2026

Official Repo for Open-Reasoner-Zero

Python 2,098 121 Updated Jun 2, 2025

Learning Formal Mathematics from Intrinsic Motivation

Rust 36 17 Updated Jul 10, 2025

An environment for learning formal mathematical reasoning from scratch

Python 72 9 Updated Aug 18, 2024

A high-performance LLM inference API and Chat UI that integrates DeepSeek R1's CoT reasoning traces with Anthropic Claude models.

Rust 1 Updated Feb 4, 2025

A high-performance LLM inference API and Chat UI that integrates DeepSeek R1's CoT reasoning traces with Anthropic Claude models.

Rust 5,363 439 Updated Oct 7, 2025
Next