Skip to content
View kir152's full-sized avatar
👾
👾

Block or report kir152

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Log viewer for app development

Rust 220 2 Updated May 14, 2026

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

Python 3,240 451 Updated Aug 15, 2026

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Python 315 65 Updated Nov 21, 2025

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python 2,000 360 Updated Aug 16, 2026

🔥 Clone and recreate any website as a modern React app in seconds

TypeScript 28,271 5,382 Updated Nov 19, 2025

A Python library for extracting structured information from unstructured text using LLMs with precise source grounding and interactive visualization.

Python 38,396 2,693 Updated Aug 11, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,043 1,148 Updated Aug 14, 2026

Revisiting Mid-training in the Era of Reinforcement Learning Scaling

Jupyter Notebook 188 14 Updated Jul 23, 2025

An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models

Python 3,362 305 Updated Aug 16, 2026

LLM checkpointing for DeepSpeed/Megatron

C++ 26 5 Updated Nov 30, 2025

Learning to Retrieve by Trying - Source code for Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval

Python 52 4 Updated Oct 31, 2024

Our library for RL environments + evals

Python 4,519 647 Updated Aug 16, 2026

[ICLR 2025] COAT: Compressing Optimizer States and Activation for Memory-Efficient FP8 Training

Python 263 28 Updated Aug 9, 2025

Get started with building Fullstack Agents using Gemini 2.5 and LangGraph

Jupyter Notebook 18,304 3,070 Updated Jun 14, 2026

Implementing DeepSeek R1's GRPO algorithm from scratch

Python 1,897 99 Updated Apr 18, 2025
Jupyter Notebook 135 16 Updated Nov 11, 2024

[ICLR 2025] Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language Models

Python 75 2 Updated Mar 29, 2025

s1: Simple test-time scaling

Python 6,663 757 Updated Jun 25, 2025

MoBA: Mixture of Block Attention for Long-Context LLMs

Python 2,165 158 Updated Apr 3, 2025

Training-free Post-training Efficient Sub-quadratic Complexity Attention. Implemented with OpenAI Triton.

Python 153 15 Updated Mar 31, 2026

a minimal cache manager for PagedAttention, on top of llama3.

Python 150 12 Updated Aug 26, 2024

Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

Python 344 31 Updated Feb 23, 2025

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,973 4,410 Updated Aug 15, 2026

RAGEN leverages reinforcement learning to train LLM reasoning agents in interactive, stochastic environments.

Python 2,770 229 Updated Jul 24, 2026

[ACL'24 Oral] Analysing The Impact of Sequence Composition on Language Model Pre-Training

Python 24 5 Updated Aug 18, 2024

A Self-adaptation Framework🐙 that adapts LLMs for unseen tasks in real-time!

Python 1,222 142 Updated Jan 30, 2025

Training Large Language Model to Reason in a Continuous Latent Space

Python 1,684 187 Updated Jul 2, 2026
Next