Skip to content
View doraa7's full-sized avatar

Block or report doraa7

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Manages Unified Access to Generative AI Services built on Envoy Gateway

Go 1,919 326 Updated Aug 10, 2026

Repository for open inference protocol specification

77 15 Updated May 12, 2025

The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

Python 8,777 1,009 Updated Aug 3, 2026

An inference server for your machine learning models, including support for multiple frameworks, multi-model serving and more

Python 897 239 Updated Aug 10, 2026

This NVIDIA RAG blueprint serves as a reference solution for a foundational Retrieval Augmented Generation (RAG) pipeline.

Python 729 314 Updated Aug 6, 2026

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…

Python 14,356 2,650 Updated Aug 11, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 88,742 20,529 Updated Aug 11, 2026

Port of OpenAI's Whisper model in C/C++

C++ 52,802 6,055 Updated Aug 7, 2026

LLM inference in C/C++

C++ 123,404 21,548 Updated Aug 11, 2026

Open source FPGA-based NIC and platform for in-network compute

Verilog 2,415 547 Updated Jul 5, 2024

A high-performance distributed file system designed to address the challenges of AI training and inference workloads.

C++ 10,107 1,076 Updated May 7, 2026

Mastering Computer Vision with PyTorch 2.0, published by Orange, AVA®

Jupyter Notebook 3 6 Updated Jan 18, 2025

Implement Neural Networks in Cuda from Scratch

C++ 23 3 Updated May 17, 2024

Karpenter is a Kubernetes Node Autoscaler built for flexibility, performance, and simplicity.

Go 2,081 547 Updated Aug 10, 2026

Neural network from scratch in CUDA/C++

Cuda 1 Updated Jan 17, 2025

Neural network from scratch in CUDA/C++

Cuda 95 18 Updated Sep 8, 2025

NVIDIA cuML: GPU-Accelerated Machine Learning

Python 5,250 651 Updated Aug 11, 2026

Slides and other materials from CppCon 2018

C++ 1,447 177 Updated Apr 11, 2019

BlazingSQL is a lightweight, GPU accelerated, SQL engine for Python. Built on RAPIDS cuDF.

C++ 2,012 182 Updated Sep 16, 2022
Jupyter Notebook 92 7 Updated Feb 29, 2024

DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.

Python 42,901 4,926 Updated Aug 11, 2026

Ongoing research training transformer models at scale

Python 17,387 4,349 Updated Aug 11, 2026

An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries

Python 7,451 1,119 Updated Jun 11, 2026

Open MPI jobs on Kubernetes

Makefile 120 25 Updated Apr 17, 2018

Open Fabric Interfaces

C 820 517 Updated Aug 11, 2026

This is a plugin which lets EC2 developers use libfabric as network provider while running NCCL applications.

C++ 231 103 Updated Aug 7, 2026

Open source project for data preparation for GenAI applications

HTML 951 252 Updated Jul 14, 2026

CUDA Core Compute Libraries

C++ 2,461 456 Updated Aug 11, 2026

cuVS - a library for vector search and clustering on the GPU

Cuda 832 217 Updated Aug 11, 2026

RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form building blocks for more easily writing …

Cuda 1,036 245 Updated Aug 11, 2026
Next