-
Virginia Tech
- Blacksburg, VA, USA
-
19:06
(UTC -12:00) - https://hasanar1f.github.io/
- @hasanar1f
- in/hasan-arif-8b78a61a3
Highlights
- Pro
Stars
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)
A simple, fast and robust program-aware agentic inference system.
Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
Google ARCore Extensions and Geospatial Creator for Unity's AR Foundation
KV cache store for distributed LLM inference
SGLang is a high-performance serving framework for large language models and multimodal models.
A high-throughput and memory-efficient inference and serving engine for LLMs
A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems
A framework for generating realistic LLM serving workloads
An efficient and scalable attention module designed to reduce memory usage and improve inference speed in large language models. Designed and implemented the Multi-Head Latent Attention (MLA) modul…
A Conversational Speech Generation Model
CUDA Templates and Python DSLs for High-Performance Linear Algebra
A final sanity checklist to help your CS paper get accepted, not desk rejected.
This is the official GitHub repository for our survey paper "Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models".
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama mode…
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filli…
Large Language Model (LLM) Systems Paper List
This repository aims to optimize the forward pass of the Flash Attention implementation in CUDA. This is a part of a graduate course project titled “Emerging Topics in CS: High-Performance Code Gen…
[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models