Skip to content
View eqy's full-sized avatar
💭
damn that's crazy
💭
damn that's crazy
  • NVIDIA

Block or report eqy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Mirage Persistent Kernel: Compiling LLMs into a MegaKernel

Cuda 2,419 238 Updated Aug 4, 2026

Prompts for our Grok chat assistant and the `@grok` bot on X.

Jinja 4,322 489 Updated Nov 17, 2025

A place to store reusable transformer components of my own creation or found on the interwebs

Python 81 11 Updated Jul 28, 2026

A list of inputs that will beat the vast majority of Pokemon Firered games

Lua 201 9 Updated Mar 10, 2023
Common Lisp 2 2 Updated Apr 22, 2025

Staging ground for release notes for PyTorch

3 6 Updated Jul 9, 2025
Shell 2 Updated Jan 5, 2026

Stack trace visualizer

Perl 4 1 Updated Sep 2, 2020

NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the effective training time by minimizing the downtime due to fa…

Python 321 62 Updated Aug 13, 2026

Implementation for MatMul-free LM.

Python 3,085 202 Updated Dec 2, 2025

Time zone database and code

C 1,825 259 Updated Aug 11, 2026
OCaml 5 1 Updated Jun 18, 2025

World's Smallest Nintendo Wii, using a trimmed motherboard and custom stacked PCBs

HTML 795 11 Updated Oct 23, 2025
Python 456 32 Updated Apr 6, 2025

GTP engine and self-play learning in Go

C++ 4,957 744 Updated Aug 12, 2026

graph generation and analysis stuff

1 Updated Feb 26, 2024

Zero Bubble Pipeline Parallelism

Python 466 31 Updated May 7, 2025

A Pin

2 Updated Aug 1, 2023

20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.

Python 13,611 1,483 Updated Jul 20, 2026
TeX 1 Updated Aug 27, 2024

You like pytorch? You like micrograd? You love tinygrad! ❤️

Python 33,444 4,252 Updated Aug 13, 2026

The official implementation of “Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”

Python 1,003 58 Updated Jan 30, 2024
TypeScript 2 Updated May 17, 2023

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in bot…

Cuda 2 Updated Sep 29, 2022

Pax is a Jax-based machine learning framework for training large scale models. Pax allows for advanced and fully configurable experimentation and parallelization, and has demonstrated industry lead…

Python 559 72 Updated Aug 7, 2026

Composable + Tunable = Optimal

Python 2 Updated Apr 14, 2023

GPU & Accelerator process monitoring for AMD, Apple, Huawei, Intel, NVIDIA and Qualcomm

C 10,907 417 Updated May 6, 2026

A Fusion Code Generator for NVIDIA GPUs (commonly known as "nvFuser")

C++ 397 83 Updated May 31, 2026

Running large language models on a single GPU for throughput-oriented scenarios.

Python 9,358 592 Updated Oct 28, 2024
Next