Skip to content
View nvbkdw's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report nvbkdw

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

provision gpu on multiple clouds

Rust 1 Updated Aug 9, 2026

Distributed transactional key-value database, originally created to complement TiDB

Rust 16,800 2,317 Updated Aug 15, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,904 242 Updated Aug 15, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 4,033 684 Updated Aug 15, 2026

A pure-Python implementation of the Nvidia CuTe layout algebra intended to be approachable and easy to learn.

Python 239 19 Updated Jun 29, 2026

Rust object_store crate

Rust 310 202 Updated Aug 12, 2026

Lightweight Kubernetes

Go 33,737 2,704 Updated Aug 15, 2026

Tantivy is a full-text search engine library inspired by Apache Lucene and written in Rust

Rust 15,699 959 Updated Aug 15, 2026

The agent that grows with you

Python 231,032 45,876 Updated Aug 15, 2026

Lightweight coding agent that runs in your terminal

Rust 106,113 16,108 Updated Aug 15, 2026

Rust bindings for AppKit (macOS) and UIKit (iOS/tvOS). Experimental, but working!

Rust 2,075 82 Updated Feb 3, 2025

For developers, who are building real-time data-driven applications, Redis is the preferred, fastest, and most feature-rich cache, data structure server, and document and vector query engine.

C 76,022 24,755 Updated Aug 13, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,435 1,958 Updated Aug 9, 2026
TypeScript 1 Updated Jun 18, 2026

Storage Performance Development Kit

C 3,637 1,371 Updated Aug 15, 2026

cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign languag…

Rust 3,065 237 Updated Aug 14, 2026

A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do

1,310 164 Updated Apr 27, 2026

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.

Jupyter Notebook 3,955 282 Updated Apr 23, 2026

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

Python 819 341 Updated Aug 15, 2026

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 26,171 2,045 Updated Aug 12, 2026

AI Infra / AI Orchestration / AI Control Plane

MDX 3,718 330 Updated Aug 15, 2026

Kubernetes-native Job Queueing

Go 2,849 749 Updated Aug 15, 2026

AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI

TypeScript 90,878 11,271 Updated Aug 15, 2026

An Interactive 3D Visual Tree Map for your S3 Bucket Storage

Python 28 5 Updated Feb 13, 2025

Experimental web-based simulator for exploring metastable behaviors in distributed systems

TypeScript 85 5 Updated May 13, 2026

A heap memory profiler for Linux

C++ 4,150 241 Updated Aug 12, 2026

DeepTutor: Lifelong Personalized Tutoring. https://deeptutor.info/.

Python 35,775 4,526 Updated Aug 13, 2026

Get your documents ready for gen AI

Python 64,798 4,620 Updated Aug 15, 2026
Next