- San Francisco, CA
- http://rockstarresearch.com/
Stars
Fast Hadamard transform in CUDA, with a PyTorch interface
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
llama.cpp fork with additional SOTA quants and improved performance
Xanmod kernel for WSL2, built by clang with ThinLTO enabled. Build & Release are automated by Github Action.
A collection of tricks and tools to speed up transformer models
[DEPRECATED] LLR2 is a primality testing program for numbers of several specific forms.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Code for Deep RL from Human Preferences [Christiano et al]. Plus a webapp for collecting human feedback
Tutorials and implementations for "Self-normalizing networks"