Starred repositories
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
End-to-end benchmark for AI-generated GPU kernels, drawn from real production traces — turn a PyTorch reference into a DSL kernel (Triton, Gluon, FlyDSL, CuteDSL) and grade it on compilation, numer…
[Experimental] Miles-diffusion is an post-training framework for large-scale diffusion model training and production workloads, forked from and co-evolving with miles.
implementations and experimentation on mHC by deepseek - https://arxiv.org/abs/2512.24880
[ICLR 2025] Official PyTorch Implementation of Gated Delta Networks: Improving Mamba2 with Delta Rule
Standalone Tencent Hy3 support for Megatron-Lite — reference example of external model integration via register_model()
Companion code for the global workspace interpretability paper
Plugin version of oh-my-humanize, keep your favorite Coding Agent : )
Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github.io/awesome-ascend-skills/
External OMH workflow artifacts
the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing figured out twice, paper: https://arxiv.…
Scalable Agentic RL for Any Agent and Sandbox.
⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents -- All for Long running Agentized Workflows
A tutorial on modern GPU programming for machine learning systems
MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation