Skip to content
View wickedfoo's full-sized avatar

Block or report wickedfoo

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Quantize transformers to any learned arbitrary 4-bit numeric format

Python 59 5 Updated Jul 2, 2026

A floating point arithmetic which works with types of any mantissa, exponent or base in modern header-only C++.

C++ 83 4 Updated Oct 22, 2024

GPU implementation of a fast generalized ANS (asymmetric numeral system) entropy encoder and decoder, with extensions for lossless compression of numerical and other data types in HPC/ML applications.

Cuda 4 Updated Mar 19, 2023

CGBN: CUDA Accelerated Multiple Precision Arithmetic (Big Num) using Cooperative Groups

Cuda 250 79 Updated Mar 31, 2026

GPU-based Distributed Point Functions (DPF) and 2-server private information retrieval (PIR).

Python 59 6 Updated Jan 27, 2023

Multithreaded Python without the GIL

Python 2,915 103 Updated May 20, 2025

Repository for nvCOMP docs and examples. nvCOMP is a library for fast lossless compression/decompression on the GPU that can be downloaded from https://developer.nvidia.com/nvcomp.

628 93 Updated Jul 13, 2026

The lightweight, fault-tolerant database built on SQLite. Designed to keep your data highly available with minimal effort.

Go 17,646 797 Updated Jul 25, 2026

Optimize floating-point expressions for accuracy

HTML 879 49 Updated Jul 26, 2026

A thin, highly portable toolkit for efficiently compiling dense loop-based computation.

C++ 147 13 Updated Jan 17, 2023

GPU implementation of a fast generalized ANS (asymmetric numeral system) entropy encoder and decoder, with extensions for lossless compression of numerical and other data types in HPC/ML applications.

Cuda 395 33 Updated Mar 18, 2026

A 16bit logarithmic fixed-point number format

Julia 9 1 Updated Dec 20, 2024

A 32-bit RISC-V Processor Designed with High-Level Synthesis

C 56 22 Updated Feb 6, 2020

The compiler is available for download. Get it!

C++ 2,568 73 Updated Nov 5, 2023

A programming language to skip the things you have already computed

JavaScript 2,019 67 Updated Sep 21, 2023

Simple C simulation of Jeff Johnson's linear-logarithmic arithmetic

C 2 1 Updated Jan 24, 2019

An exploration of log domain "alternative floating point" for hardware ML/AI accelerators.

SystemVerilog 401 39 Updated Mar 11, 2023

A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)

Python 5,635 941 Updated Jul 7, 2026

cuML - RAPIDS Machine Learning Library

Python 5,237 647 Updated Jul 25, 2026

Tensors and Dynamic neural networks in Python with strong GPU acceleration

Python 101,978 28,516 Updated Jul 26, 2026

A domain specific language to express machine learning workloads.

C++ 1,767 213 Updated Apr 28, 2023

10x faster matrix and vector operations

C++ 2,513 174 Updated Oct 12, 2022

CUDA Tensor Transpose (cuTT) library

C++ 55 29 Updated Aug 10, 2017

Collective communications library with various primitives for multi-machine training.

C++ 1,438 362 Updated Jul 1, 2026

Caffe2 is a lightweight, modular, and scalable deep learning framework.

Shell 8,374 1,899 Updated Feb 7, 2023

A library for efficient similarity search and clustering of dense vectors.

C++ 40,586 4,471 Updated Jul 24, 2026

A CUDA backend for Torch7

Cuda 340 208 Updated Sep 11, 2017