Skip to content
View dukebw's full-sized avatar

Highlights

  • Pro

Block or report dukebw

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Inference at the speed of light.

Rust 2,923 218 Updated Aug 9, 2026

A helm plugin that shows a diff explaining what a helm upgrade would change

Go 3,481 324 Updated Aug 1, 2026

⎈ Multi pod and container log tailing for Kubernetes -- Friendly fork of https://github.com/wercker/stern

Go 4,825 173 Updated Jul 24, 2026

FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.

Python 260 104 Updated Aug 10, 2026

Iterative agent harness improvement: run a coding agent on a hard task, generate the reusable tooling it was missing, qualify it, and replay fresh sessions with it activated. Works with Codex and C…

Python 40 3 Updated Jul 2, 2026

AIPerf is a comprehensive benchmarking tool that measures the performance of generative AI models served by your preferred inference solution.

Python 528 145 Updated Aug 10, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,916 645 Updated Jul 9, 2026

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer

Python 8,733 698 Updated Aug 10, 2026

Stacked diff support for GitHub workflows

Shell 181 17 Updated Aug 8, 2026
Python 402 50 Updated Jul 30, 2026

TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration

Python 1,720 191 Updated Mar 27, 2026

Durable, guarded goal workflows for OpenCode with persistence, safety limits, agent tools, and evidence-gated completion.

JavaScript 226 21 Updated Aug 8, 2026

Ideogram 4: Open image model at the forefront of design

Python 2,703 279 Updated Jun 30, 2026

NVIDIA FastGen: Fast Generation from Diffusion Models

Python 934 77 Updated Aug 4, 2026

Conveniently export torch.compile compiled products into self-contained Python files

Python 35 3 Updated Jun 5, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,838 227 Updated Aug 10, 2026

Ready-to-use ML training recipes to help you build and deploy models on Baseten.

Python 62 8 Updated Aug 9, 2026

The lightweight framework for building agents

Python 536 60 Updated Aug 5, 2026

AI Tensor Engine for ROCm

Python 523 460 Updated Aug 10, 2026

A modern alternative to ls

Rust 22,887 491 Updated Aug 6, 2026

A fast type checker and language server for Python

Rust 6,865 472 Updated Aug 10, 2026

Performant kernels, and other ML Systems integrations

Cuda 6 3 Updated Jul 23, 2026

From a+b to sparsemax(QK^T)V in Triton!

Jupyter Notebook 34 Updated Jun 19, 2025

Anthropic's original performance take-home, now open for you to try!

Python 4,092 930 Updated Jan 22, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,174 685 Updated Aug 10, 2026

A kernel library written in tilelang

Python 1,711 155 Updated Apr 23, 2026

CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.

Python 535 70 Updated Aug 9, 2026

Use Codex from Claude Code to review code or delegate tasks.

JavaScript 31,625 2,161 Updated Jul 8, 2026

Region-level profiling for CUDA kernels with trace, NVBit, CUPTI, NSys, and an interactive Explorer.

Python 123 11 Updated Apr 17, 2026
Next