Skip to content
View 0xPabloxx's full-sized avatar
👣
Focusing
👣
Focusing

Highlights

  • Pro

Organizations

@langgenius

Block or report 0xPabloxx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Auditable quantitative research for market structure and time-correct experiments. / 可复现、可审计且避免未来数据的结构与趋势量化研究台

Python 2 1 Updated Aug 10, 2026

A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

Python 168 12 Updated May 1, 2026

a toolkit on knowledge distillation for large language models

Python 447 43 Updated Mar 10, 2026

Official code for "Self-Distilled Agentic Reinforcement Learning"

Python 334 25 Updated Aug 10, 2026

Post-training with Tinker

Python 4,009 507 Updated Aug 11, 2026

Claw Code No Rust No TypeScript Only Python. Easy to work with. Fast to iterate. 🔥 Zero external dependencies 🔥

Python 538 222 Updated Jun 22, 2026

Token efficient Claude Code full Python rebuild. AI Coding Agent in 310K LoC Python. Up to 200X Cost Saving!

TypeScript 860 154 Updated Aug 11, 2026

Uni-Agent is a framework for training long-horizon agents.

Python 504 91 Updated Aug 11, 2026

与其蒸馏别人,不如蒸馏自己。欢迎加入数字永生!Inspired by colleague-skill(同事skill)。

Python 3,281 272 Updated Apr 1, 2026

从 幻觉翻译 获取基于 LaTex 源码翻译的arXiv文章

TypeScript 20 3 Updated Jul 20, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,122 1,586 Updated Aug 11, 2026

AgentFlow: In-the-Flow Agentic System Optimization

Python 2,001 233 Updated Feb 8, 2026

将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶

Python 20,814 2,028 Updated Aug 11, 2026

你想蒸馏的下一个员工,何必是同事。蒸馏任何人的思维方式——心智模型、决策启发式、表达DNA。Distill how anyone thinks.

Python 30,348 4,202 Updated Jul 27, 2026

A project implementing various agentic RL based on the Slime post-training framework

Python 516 36 Updated Apr 11, 2026

OpenSeeker: A search agent with open-source data and models

Python 767 60 Updated Jun 25, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,629 606 Updated May 23, 2026

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,656 574 Updated Aug 12, 2026

A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience

TypeScript 137,220 9,899 Updated Aug 12, 2026

MiroMind Deep Research Skill

23 3 Updated Mar 19, 2026

Agent harness to publish your agent chat history as Huggingface datasets.

Python 2,109 233 Updated Jun 5, 2026

bilibili video course src code

Jupyter Notebook 478 123 Updated Nov 14, 2023

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

Python 1,159 108 Updated Jun 10, 2026

An agentic skills framework & software development methodology that works.

Shell 270,812 24,194 Updated Aug 8, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,860 1,132 Updated Aug 11, 2026

Open Visual Agentic Intelligence

2,298 322 Updated Aug 6, 2026

Scalable toolkit for efficient model reinforcement

Python 1,897 508 Updated Aug 12, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,145 402 Updated Aug 11, 2026
Python 1,306 135 Updated May 20, 2026

A curated list of reinforcement learning (RL) for agents.

111 5 Updated Jul 27, 2026
Next