Skip to content
View GCC314's full-sized avatar
C4NDM!
C4NDM!

Highlights

  • Pro

Block or report GCC314

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

AI-powered reverse engineering assistant that bridges IDA Pro with language models through MCP.

Python 11,456 1,371 Updated Aug 17, 2026

gem5 repository to study chiplet-based systems

C++ 91 19 Updated Apr 18, 2019

Artifacts for "vCXLGen: Automated Synthesis and Verification of CXL Bridges for Heterogeneous Architectures", ASPLOS'26

C 11 1 Updated Dec 22, 2025

Artifact for "C3: CXL Coherence Controllers for Heterogeneous Architectures" HPCA '26

C 6 1 Updated Dec 30, 2025

GeminiFS: A Companion File System for GPUs

C++ 86 18 Updated Aug 11, 2026
C++ 9 4 Updated Dec 13, 2024

proof of concepts and experimental code for TFHE

C++ 21 1 Updated Feb 18, 2020

Pure C++ Ver. of TFHE.

C++ 112 20 Updated Aug 17, 2026

Zama's Homomorphic Processing Unit implementation on FPGA

SystemVerilog 230 33 Updated Jul 21, 2026

example code for using DC QP for providing RDMA READ and WRITE operations to remote GPU memory

C 158 38 Updated Jul 30, 2024

TFHE-rs: A Pure Rust implementation of the TFHE Scheme for Boolean and Integer Arithmetics Over Encrypted Data.

Rust 1,644 333 Updated Aug 20, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,258 1,531 Updated Aug 19, 2026

Build userspace NVMe drivers and storage applications with CUDA support

C 444 56 Updated Dec 18, 2023

My learning notes for ML SYS.

HTML 6,900 482 Updated Aug 19, 2026

A high-performance distributed file system designed to address the challenges of AI training and inference workloads.

C++ 10,149 1,082 Updated May 7, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,705 1,186 Updated Aug 11, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,854 1,127 Updated Jul 28, 2026

llama.cpp to PyTorch Converter

Cuda 38 7 Updated Apr 8, 2024

PKU course materials on computer science and life science.

C++ 205 11 Updated Jun 15, 2026

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

C 1,408 194 Updated Jul 14, 2026

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 102,497 28,928 Updated Aug 20, 2026
C++ 19 3 Updated Dec 12, 2023

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

Python 42,966 4,935 Updated Aug 20, 2026

Let your Claude able to think

TypeScript 17,053 1,967 Updated Apr 7, 2026

[HPCA'24] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System

Python 52 7 Updated Jul 21, 2025

The official implementation of paper: SimLayerKV: A Simple Framework for Layer-Level KV Cache Reduction.

Python 55 Updated Oct 18, 2024
Jupyter Notebook 73 Updated Oct 31, 2024

Port of Facebook's LLaMA model in C/C++

C 21 4 Updated Nov 6, 2023
Next