-
University of Vienna
- Vienna, Austria
- @cai_yitao
Lists (1)
Sort Name ascending (A-Z)
Starred repositories
DeepSeek Harness: Everything is a Plugin.
Skills for AIs using the Lean programming language and theorem prover — proofs, toolchain setup, bisection, and more
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
Full stack application for creating, combining, and storing combinatorial genetic design spaces.
Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels with Hunyuan3D World Model
[ACM CSUR 2025] Understanding World or Predicting Future? A Comprehensive Survey of World Models
《动手学深度学习》:面向中文读者、能运行、可讨论。中英文版被70多个国家的500多所大学用于教学。
Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Reality
Fast and memory-efficient exact attention
🚀 Efficient implementations for emerging model architectures
Hierarchical Generation of Molecular Graphs using Structural Motifs
Junction Tree Variational Autoencoder for Molecular Graph Generation (ICML 2018)
Continuous Thought Machines, because thought takes time and reasoning is a process.
Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods
Tensors and Dynamic neural networks in Python with strong GPU acceleration
[ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoring
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
Efficient triton implementation of Native Sparse Attention.
QwQ is the reasoning model series developed by Qwen team, Alibaba Cloud.
Fully open reproduction of DeepSeek-R1
Python package built to ease deep learning on graph, on top of existing DL frameworks.
Implementation of the sparse attention pattern proposed by the Deepseek team in their "Native Sparse Attention" paper
Generative Pre-trained Graph Eulerian Transformer [ICML2025]