Skip to content
View wickedfoo's full-sized avatar

Block or report wickedfoo

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Quantize transformers to any learned arbitrary 4-bit numeric format

Python 59 5 Updated Jul 2, 2026

A floating point arithmetic which works with types of any mantissa, exponent or base in modern header-only C++.

C++ 83 4 Updated Oct 22, 2024

GPU implementation of a fast generalized ANS (asymmetric numeral system) entropy encoder and decoder, with extensions for lossless compression of numerical and other data types in HPC/ML applications.

Cuda 4 Updated Mar 19, 2023

CGBN: CUDA Accelerated Multiple Precision Arithmetic (Big Num) using Cooperative Groups

Cuda 250 79 Updated Mar 31, 2026

GPU-based Distributed Point Functions (DPF) and 2-server private information retrieval (PIR).

Python 59 6 Updated Jan 27, 2023

Multithreaded Python without the GIL

Python 2,911 103 Updated May 20, 2025

Repository for nvCOMP docs and examples. nvCOMP is a library for fast lossless compression/decompression on the GPU that can be downloaded from https://developer.nvidia.com/nvcomp.

628 92 Updated Jul 13, 2026

The lightweight, fault-tolerant database built on SQLite. Designed to keep your data highly available with minimal effort.

Go 17,664 800 Updated Jul 29, 2026

Optimize floating-point expressions for accuracy

HTML 879 49 Updated Aug 3, 2026

A thin, highly portable toolkit for efficiently compiling dense loop-based computation.

C++ 147 13 Updated Jan 17, 2023

GPU implementation of a fast generalized ANS (asymmetric numeral system) entropy encoder and decoder, with extensions for lossless compression of numerical and other data types in HPC/ML applications.

Cuda 396 33 Updated Mar 18, 2026

A 16bit logarithmic fixed-point number format

Julia 9 1 Updated Dec 20, 2024

A 32-bit RISC-V Processor Designed with High-Level Synthesis

C 56 22 Updated Feb 6, 2020

The compiler is available for download. Get it!

C++ 2,572 73 Updated Nov 5, 2023

A programming language to skip the things you have already computed

JavaScript 2,020 67 Updated Sep 21, 2023

Simple C simulation of Jeff Johnson's linear-logarithmic arithmetic

C 2 1 Updated Jan 24, 2019

An exploration of log domain "alternative floating point" for hardware ML/AI accelerators.

SystemVerilog 401 39 Updated Mar 11, 2023

A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)

Python 5,636 940 Updated Jul 7, 2026

cuML - RAPIDS Machine Learning Library

Python 5,245 648 Updated Aug 4, 2026

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 102,174 28,703 Updated Aug 4, 2026

A domain specific language to express machine learning workloads.

C++ 1,767 213 Updated Apr 28, 2023

10x faster matrix and vector operations

C++ 2,514 174 Updated Oct 12, 2022

CUDA Tensor Transpose (cuTT) library

C++ 55 29 Updated Aug 10, 2017

Collective communications library with various primitives for multi-machine training.

C++ 1,440 362 Updated Jul 31, 2026

Caffe2 is a lightweight, modular, and scalable deep learning framework.

Shell 8,373 1,899 Updated Feb 7, 2023

A library for efficient similarity search and clustering of dense vectors.

C++ 40,664 4,479 Updated Aug 3, 2026

A CUDA backend for Torch7

Cuda 340 208 Updated Sep 11, 2017