Skip to content
View hasanar1f's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report hasanar1f

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

Python 1,005 292 Updated Jul 24, 2026

PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)

Python 33 4 Updated Jun 10, 2026

A simple, fast and robust program-aware agentic inference system.

Python 376 35 Updated Jul 5, 2026

Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference

Go 5,039 775 Updated Jul 24, 2026

Learned LLM Prefix Cache

Python 5 2 Updated Oct 10, 2025

Google ARCore Extensions and Geospatial Creator for Unity's AR Foundation

C# 394 108 Updated Apr 22, 2026

A lightweight and fast LLM serving framework

Python 15 2 Updated Mar 5, 2026

KV cache store for distributed LLM inference

C++ 425 43 Updated Nov 13, 2025

Materials for learning SGLang

860 64 Updated Jan 5, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 30,697 7,365 Updated Jul 24, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,029 19,788 Updated Jul 24, 2026

Microsoft Azure Traces

Jupyter Notebook 1,159 184 Updated Jun 3, 2026

A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems

Python 280 15 Updated Jun 30, 2026

A framework for generating realistic LLM serving workloads

Python 165 15 Updated May 11, 2026

An efficient and scalable attention module designed to reduce memory usage and improve inference speed in large language models. Designed and implemented the Multi-Head Latent Attention (MLA) modul…

Python 25 4 Updated Jun 25, 2025

A Conversational Speech Generation Model

Python 14,696 1,484 Updated May 27, 2025
Python 63 14 Updated May 16, 2025

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++ 10,122 1,978 Updated Jul 23, 2026

A final sanity checklist to help your CS paper get accepted, not desk rejected.

1,601 145 Updated May 25, 2026

This is the official GitHub repository for our survey paper "Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models".

Python 200 6 Updated Jul 11, 2026

Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama mode…

Jupyter Notebook 18,538 2,766 Updated May 19, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 10,852 1,611 Updated Jul 24, 2026

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filli…

Python 1,221 80 Updated Apr 8, 2026

Large Language Model (LLM) Systems Paper List

2,198 118 Updated Jul 15, 2026

The official Meta Llama 3 GitHub site

Python 29,279 3,531 Updated Jan 26, 2025

This repository aims to optimize the forward pass of the Flash Attention implementation in CUDA. This is a part of a graduate course project titled “Emerging Topics in CS: High-Performance Code Gen…

Cuda 2 Updated May 3, 2025

100 days of building GPU kernels!

Cuda 616 78 Updated Apr 27, 2025

[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Python 1,023 77 Updated Jul 10, 2025

Kolmogorov Arnold Networks

Jupyter Notebook 16,326 1,565 Updated Jan 19, 2025
Next