Skip to content
View alecco's full-sized avatar
  • Spain

Block or report alecco

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An Optimizer for Nvidia Compilers.

Python 126 12 Updated Aug 6, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,583 7,761 Updated Aug 9, 2026

DeepGEMM: clean and efficient BLAS kernel library on GPU

Cuda 7,638 1,159 Updated Jul 20, 2026

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support

Shell 285 19 Updated Aug 9, 2026

Normalized Transformer (nGPT)

Python 211 27 Updated Nov 19, 2024

Reference implementation and examples of the CuTe Layout representation and algebra.

Python 259 24 Updated Aug 6, 2026

Quartet II Official Code

Python 80 10 Updated May 1, 2026

High-throughput tensor loading for PyTorch

Python 261 16 Updated Jul 31, 2026
Python 7 3 Updated May 7, 2026
Python 9 Updated May 26, 2026

100M tokens. Infinite compute. Lowest val loss wins.

Python 523 81 Updated Jul 3, 2026

CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Python 238 24 Updated Aug 6, 2026

Long Context Pre-Training with Lighthouse Attention

Python 65 13 Updated Jul 31, 2026

The agent that grows with you

Python 227,893 44,734 Updated Aug 9, 2026

Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.

Cuda 256 25 Updated Jun 29, 2026

Muon is an optimizer for hidden layers in neural networks

Python 2,772 128 Updated May 24, 2026

NanoGPT (124M) in 90 seconds

Python 5,655 860 Updated Aug 2, 2026

SmoothE: Differentiable E-Graph Extraction (ASPLOS'25 Best Paper)

Python 34 3 Updated Jan 15, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,836 225 Updated Aug 9, 2026

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

Python 903 242 Updated Aug 9, 2026

Node0: A collaborative event powered by Protocol Learning, our decentralized approach to AI development

Python 96 29 Updated Sep 23, 2025

Learn CUDA with PyTorch

Cuda 364 53 Updated Jun 1, 2026

Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotuning, and advanced scheduling. Run benchmarks across different …

Python 75 8 Updated Aug 9, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,607 465 Updated Jan 17, 2026

A beautiful, simple, clean, and responsive Jekyll theme for academics

HTML 15,986 13,092 Updated Aug 9, 2026

Every Code - push frontier AI to it limits. A fork of the Codex CLI with validation, automation, browser integration, multi-agents, theming, and much more. Orchestrate agents from OpenAI, Claude, G…

Rust 3,864 235 Updated Aug 8, 2026

FlashKDA: high-performance Kimi Delta Attention kernels

Cuda 1,193 112 Updated Jul 30, 2026

how few training tokens can you use to reach a target validation loss?

Python 5 Updated Apr 6, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 737 96 Updated Jul 4, 2026
Next