Starred repositories
Synthetic data generation, post-training, and E2B benchmark evaluation infrastructure.
The RedStone repository includes code for preparing extensive datasets used in training large language models.
Steinbeck-Lab / DECIMER.ai
Forked from OBrink/DECIMER.aiThis repository contains the code for https://decimer.ai
AI agents running research on single-GPU nanochat training automatically
Get your documents ready for gen AI
Data Efficacy for Language Model Training
The batteries-included agent harness.
Building and Scaling RL environments in the age of LLMs
An AI SKILL that provide design intelligence for building professional UI/UX multiple platforms
The agent that grows with you
将冰冷的离别化为温暖的 Skill,欢迎加入数字生命1.0!Transforming cold farewells into warm skills? It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Site-backed dataset, reconciled package files, and interactive docs for OpenAI vs Anthropic compute estimates.
Official PyTorch implementation for "Large Language Diffusion Models"
from vibe coding to agentic engineering - practice makes codex perfect
Implement a reasoning LLM in PyTorch from scratch, step by step
An agentic skills framework & software development methodology that works.
Universal scraping tool, which allows you to extract data using multiple environments
你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
Measuring and evolving with the frontier of agent work
OpenClaw-RL: Train any agent simply by talking
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
Language models scale reliably with over-training and on downstream tasks
🚀 Efficient implementations for emerging model architectures
Our code for ICLR'25 paper "DataMan: Data Manager for Pre-training Large Language Models".