Highlights
- Pro
Stars
Supplementary files & data for the Combinatorial Reasoning Paper
Artifact material for [ISCA 2026] #1625 "Bringing Near Data Processing into the Low-Bit Floating-Point Era"
The official code implementation for MCI-GRU: Stock Prediction Model Based on Multi-Head Cross-Attention and Improved GRU.
LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost…
Continuous Thought Machines, because thought takes time and reasoning is a process.
A machine learning accelerator core designed for energy-efficient AI at the edge.
Some Hardware Architectures for GEMM
A lightweight UI framework based on tkinter with all UI drawn in Canvas!🎨
Code to simulate energy-based analog systems and equilibrium propagation
MixTeX multimodal LaTeX, ZhEn, and, Table OCR. It performs efficient CPU-based inference in a local offline on Windows.
Unofficial Reimplementation of VLTSeg from "Strong but Simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning"
SAURIA (Systolic-Array tensor Unit for aRtificial Intelligence Acceleration) is an open-source Convolutional Neural Network accelerator based on a GeMM systolic array engine.
DeepGEMM: clean and efficient BLAS kernel library on GPU
A digital logic designer and circuit simulator.
FlagGems is an operator library for large language models implemented in the Triton Language.
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
ROCm / triton
Forked from triton-lang/tritonDevelopment repository for the Triton language and compiler
CraftsMan: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner
Tackling the Generative Learning Trilemma with Denoising Diffusion GANs https://arxiv.org/abs/2112.07804
Source code for the ICML2019 paper "Subspace Robust Wasserstein Distances"
Official PyTorch codes of CVPR2022 Oral: Exact Feature Distribution Matching for Arbitrary Style Transfer and Domain Generalization
(NeurIPS 2024 Oral 🔥) Improved Distribution Matching Distillation for Fast Image Synthesis