Skip to content
View p208p2002's full-sized avatar

Organizations

@NCHU-NLP-Lab

Block or report p208p2002

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Checkpoint/Restore tool

C 3,956 765 Updated Aug 10, 2026

Sardeenz is a proof-of-concept application that allows you to load more than one model on a given GPU. It allows you to add more and more models onto a GPU, until it is fully utilized.

TypeScript 61 8 Updated Jun 9, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,131 130 Updated Aug 11, 2026

Algorithm powering the For You feed on X

Rust 26,975 4,590 Updated May 15, 2026

Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs

Python 1,493 209 Updated Aug 12, 2026

A small command line tool to simplify releasing software by updating all version strings in your source code by the correct increment and optionally commit and tag the changes.

Python 623 42 Updated Aug 10, 2026

A collection of examples for the ROCm software stack

C++ 311 96 Updated Aug 7, 2026

Allow torch tensor memory to be released and resumed later

Python 267 69 Updated Aug 9, 2026

My learning notes for ML SYS.

HTML 6,856 476 Updated Aug 11, 2026

Official Code Repository for the paper "Distilling LLM Agent into Small Models with Retrieval and Code Tools"

Python 257 35 Updated Oct 22, 2025

FlashInfer: Kernel Library for LLM Serving

Python 6,146 1,266 Updated Aug 12, 2026

Awesome Reasoning LLM Tutorial/Survey/Guide

Python 2,509 164 Updated Jul 27, 2026
Python 1,175 58 Updated Jan 10, 2026

Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

Cuda 11,752 1,236 Updated Aug 6, 2026

Minimal example for DeepSpeed Universal Checkpoint

Python 1 Updated Sep 20, 2024
Python 98 12 Updated Nov 6, 2024

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

Python 2,516 526 Updated Aug 11, 2026

WikiTableSet: A largest publicly available image-based table recognition dataset in three languages built from Wikipedia

Python 32 2 Updated Jun 12, 2025

Automatic GPU+CPU memory profiling, re-use and memory leaks detection using jupyter/ipython experiment containers

Jupyter Notebook 236 15 Updated Dec 15, 2023

Open GenAI Stack

Python 8,421 1,344 Updated Aug 12, 2026

Run PyTorch LLMs locally on servers, desktop and mobile

Python 3,618 246 Updated Sep 10, 2025

We collect papers about "large language models (LLM) for table-related tasks", e.g., using LLM for Table QA task. “表格+LLM”相关论文整理

633 46 Updated Apr 9, 2026

[EMNLP 2023] Enabling Large Language Models to Generate Text with Citations. Paper: https://arxiv.org/abs/2305.14627

Python 525 51 Updated Oct 9, 2024

This project has implemented the RAG function on Jetson and supports TXT and PDF document formats. It uses MLC for 4-bit quantization of the Llama2-7b model, utilizes ChromaDB as the vector databas…

Python 12 1 Updated May 16, 2024

LLM training code for Databricks foundation models

Python 4,438 591 Updated Mar 25, 2026

egui: an easy-to-use immediate mode GUI in Rust that runs on both web and native

Rust 30,033 2,101 Updated Aug 11, 2026

🐳 Web Interface for the Docker Registry HTTP API V2 written in Ruby on Rails.

Ruby 697 63 Updated Aug 11, 2026

The simplest and most complete UI for your private docker registry v2 and v3

Riot 3,513 366 Updated Aug 10, 2026

Multiple NVIDIA GPUs or Apple Silicon for Large Language Model Inference?

Jupyter Notebook 1,933 76 Updated May 13, 2024

Source code of "Reasons to Reject? Aligning Language Models with Judgments"

Python 58 5 Updated Feb 29, 2024
Next