-
Institute of Computing Technology, Chinese Academy of Sciences
-
09:22
(UTC -12:00) - https://kechang.xin
- https://www.zhihu.com/people/deconx
Highlights
- Pro
Lists (2)
Sort Name ascending (A-Z)
Stars
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.
DeepStack: Facilitating Co-Design Exploration of 3D DRAM-Stacked Accelerators for Distributed LLM Inference. Includes the MICRO 2026 AE artifact.
A collection of tricks and tools to speed up transformer models
Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing…
Training a general LLM-based reasoning model for agent-compiled knowledge refinement that evolves any KB with multi-turn interaction history via RL.
Can LLMs Write Correct and Efficient GPU Communication Code?
A Micro-benchmarking Tool for HPC Networks
AI 时代的伯克希尔:基于 Claude Code / Codex 的价值投资研究框架。巴菲特·芒格·段永平·李录四大师方法论 + 多Agent并行研究。| AI-era Berkshire: a value investing research framework built for Claude Code / Codex. 4 masters' methodologies + multi…
DFlash: Block Diffusion for Flash Speculative Decoding
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
a high performance library for building cache simulators
[ACL 2026] QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
Google Cloud Knowledge Catalog Tools and Samples
Debug print operator for cudagraph debugging
HRM-Text is a 1B text generation model based on the HRM architecture, strengthened by task completion and latent space reasoning.
Conveniently export torch.compile compiled products into self-contained Python files
Uni-Agent is a framework for training long-horizon agents.
Framework for evaluating and improving agents
ChampSim is an open-source trace based simulator maintained at Texas A&M University and through the support of the computer architecture community.
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign languag…
A native macOS app for syncing projects to your local filesystem
👾 Open Computer Use – Open-Source Alternative to Codex Computer Use
high-performance linear attention kernel library built on TileLang