Skip to content
View pyemma's full-sized avatar

Block or report pyemma

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

🤖 The analysis of Claude Code

TypeScript 3,827 2,047 Updated Apr 2, 2026

🧠 Train a 64M-parameter LLM from scratch in just 2h!

Python 54,864 7,167 Updated Aug 6, 2026

An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models

Python 3,365 306 Updated Aug 20, 2026

🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.

Python 34,345 7,252 Updated Aug 20, 2026

Research on Coding Agents

12,227 19,583 Updated Apr 1, 2026

An educational resource to help anyone learn deep reinforcement learning.

Python 11,903 2,465 Updated Aug 5, 2024

The corresponding codes and dataset for OneSearch series

Python 170 18 Updated May 1, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,640 608 Updated May 23, 2026

A Claude Code skill that acts as your daily 军师 (strategic research advisor).

Shell 109 6 Updated Mar 18, 2026
Python 1,307 135 Updated May 20, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,936 1,003 Updated Aug 13, 2026

Implement a reasoning LLM in PyTorch from scratch, step by step

Jupyter Notebook 5,017 766 Updated Aug 4, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,124 385 Updated Jul 30, 2026

AI agents running research on single-GPU nanochat training automatically

Python 94,223 13,321 Updated Mar 26, 2026

Open-source RL Framework with Online Teacher-Student Distillation

Python 22 1 Updated Mar 5, 2026

Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200

Python 120 10 Updated Feb 28, 2026

Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding (https://arxiv.org/abs/2602.22647)

Python 231 28 Updated Mar 29, 2026

Minimalistic 4D-parallelism distributed training framework for education purpose

Python 2,286 201 Updated Aug 26, 2025

Fast, small, and fully autonomous AI personal assistant infrastructure, any OS, any platform — deploy anywhere, swap anything 🦀

Rust 32,624 4,903 Updated Aug 20, 2026

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

TypeScript 386,843 81,265 Updated Aug 20, 2026

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Python 1,703 256 Updated Apr 24, 2026

The official code for paper "Token-Level Collaborative Alignment for LLM-based Generative Recommendation"

Python 19 Updated Jun 9, 2026

[TMLR 26]: "UniRec: Unified Multimodal Encoding for LLM-Based Recommendations", Zijie Lei, Tao Feng, Zhigang Hua, Yan Xie, Guanyu Lin, Shuang Yang, Ge Liu, Jiaxuan You

Python 15 1 Updated Jul 9, 2026

"Unleashing the Potential of Sparse Attention on Long-term Behaviors for CTR Prediction." In Proceedings of WWW '26.

Python 9 3 Updated Jan 23, 2026

Our first fully AI generated deep learning system

Python 635 50 Updated Feb 2, 2026

Algorithm powering the For You feed on X

Rust 32,067 5,259 Updated Aug 19, 2026

High-quality single file implementation of Deep Reinforcement Learning algorithms with research-friendly features (PPO, DQN, C51, DDPG, TD3, SAC, PPG)

Python 10,295 1,154 Updated Apr 20, 2026

An Open Foundation Model and Benchmark to Accelerate Generative Recommendation

Python 902 132 Updated May 18, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,780 794 Updated May 17, 2026

The official implementation of "ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning"

Python 444 57 Updated Mar 29, 2026
Next