Skip to content
View 0xPabloxx's full-sized avatar
👣
Focusing
👣
Focusing

Highlights

  • Pro

Organizations

@langgenius

Block or report 0xPabloxx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A Business-Driven Real-World Financial Benchmark for Evaluating LLMs

Python 168 12 Updated May 1, 2026

a toolkit on knowledge distillation for large language models

Python 443 44 Updated Mar 10, 2026

Official code for "Self-Distilled Agentic Reinforcement Learning"

Python 312 24 Updated Jul 23, 2026

Post-training with Tinker

Python 3,950 494 Updated Jul 29, 2026

Claw Code No Rust No TypeScript Only Python. Easy to work with. Fast to iterate. 🔥 Zero external dependencies 🔥

Python 535 219 Updated Jun 22, 2026

Token efficient Claude Code full Python rebuild. AI Coding Agent in 270K LoC pure Python. Up to 200X Cost Saving!

Python 818 146 Updated Jul 29, 2026

Uni-Agent is a framework for training long-horizon agents.

Python 458 80 Updated Jul 29, 2026

与其蒸馏别人,不如蒸馏自己。欢迎加入数字永生!Inspired by colleague-skill(同事skill)。

Python 3,227 269 Updated Apr 1, 2026

从 幻觉翻译 获取基于 LaTex 源码翻译的arXiv文章

TypeScript 19 3 Updated Jul 20, 2026

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 14,995 1,560 Updated Jul 29, 2026

AgentFlow: In-the-Flow Agentic System Optimization

Python 1,981 232 Updated Feb 8, 2026

将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶

Python 20,590 2,014 Updated Jun 1, 2026

你想蒸馏的下一个员工,何必是同事。蒸馏任何人的思维方式——心智模型、决策启发式、表达DNA。Distill how anyone thinks.

Python 29,130 4,078 Updated Jul 27, 2026

A project implementing various agentic RL based on the Slime post-training framework

Python 509 35 Updated Apr 11, 2026

OpenSeeker: A search agent with open-source data and models

Python 764 59 Updated Jun 25, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,614 608 Updated May 23, 2026

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,615 568 Updated Jul 29, 2026

A modern GUI client based on Tauri, designed to run in Windows, macOS and Linux for tailored proxy experience

TypeScript 134,428 9,709 Updated Jul 29, 2026

MiroMind Deep Research Skill

22 3 Updated Mar 19, 2026

Agent harness to publish your agent chat history as Huggingface datasets.

Python 2,109 235 Updated Jun 5, 2026

bilibili video course src code

Jupyter Notebook 475 123 Updated Nov 14, 2023

OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

Python 1,115 103 Updated Jun 10, 2026

An agentic skills framework & software development methodology that works.

Shell 263,187 23,498 Updated Jul 28, 2026

slime is an LLM post-training framework for RL Scaling.

Python 7,697 1,104 Updated Jul 24, 2026

Open Visual Agentic Intelligence

2,290 318 Updated Jan 31, 2026

Scalable toolkit for efficient model reinforcement

Python 1,857 490 Updated Jul 29, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,104 393 Updated Jul 29, 2026
Python 1,298 135 Updated May 20, 2026

A curated list of reinforcement learning (RL) for agents.

111 5 Updated Jul 27, 2026

全员 AI 的公司,只有老板是人。7×24 为你做调研、发媒体、写代码、管投资。基于 Claude Code 构建。

Python 28 10 Updated Apr 30, 2026
Next