-
ICT, CAS
- Beijing
- rong-hash.github.io
Stars
FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? (COLM 2026)
Babel (Kimi Code edition): open-source AI-native chiplet design flow
Benchmarking LLM agents on real-world hardware bug repair tasks
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
AgentENV (AENV) is a distributed platform for running agent environments at scale.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
RuBench: repository-level agentic coding benchmark with natively authored Russian task specifications
FlashKDA: high-performance Kimi Delta Attention kernels
Kimi Code CLI β The Starting Point for Next-Gen Agents
kernelbench.com β GPU kernel engineering benchmarks for autonomous LLM coding agents. v3 archive + v-hard latest.
opensource NPU for LLM inference (this run gpt2)
A eDSL framework based on Scala and MLIR, focusing on the Hardware design.
Cycle-accurate C++ & SystemC simulator for the RISC-V GPGPU Ventus
A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.
A kernel library written in tilelang
ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning (ACL 2026)
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery π§βπ¬
[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
[ICLR 2026] Official Implementation of "FeatureBench: Benchmarking Agentic Coding for Complex Feature Development"
FrontierSWE is an ultra long-horizon coding agent benchmark that tests implementation, performance eng and ML research