Skip to content
View wooksu's full-sized avatar

Organizations

@nota-github

Block or report wooksu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,664 4,276 Updated Jul 25, 2026

Fully Open Framework for Democratized Multimodal Training

Python 1,151 77 Updated Jul 26, 2026

Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Python 386 20 Updated Jun 20, 2026

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 73,518 8,986 Updated Jul 24, 2026

Official implementation for "TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models" https://arxiv.org/pdf/2602.04657

Python 43 6 Updated Jun 3, 2026

omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode

TypeScript 66,605 5,430 Updated Jul 26, 2026

A simple yet powerful agent framework that delivers with open-source models

Python 4,582 471 Updated Mar 21, 2026

ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency objectives. [ICLR'26]

Python 19 1 Updated Feb 25, 2026

[ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

Python 92 1 Updated Jan 26, 2026
Python 46 1 Updated Jul 14, 2025

Nano vLLM

Python 14,642 2,355 Updated Apr 26, 2026

[NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

Python 105 5 Updated Sep 20, 2025

Official code for NeurIPS 2025 paper "GRIT: Teaching MLLMs to Think with Images"

Python 190 10 Updated Jan 16, 2026
Python 1,251 77 Updated Nov 20, 2025

Open-source unified multimodal model

Python 6,120 544 Updated May 4, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,435 64 Updated May 11, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,081 383 Updated Jul 23, 2026

State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!

Jupyter Notebook 2,329 157 Updated Apr 13, 2026

Solve Visual Understanding with Reinforced VLMs

Python 6,015 383 Updated Jul 7, 2026

Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’

Jupyter Notebook 2,263 110 Updated Oct 29, 2025

Witness the aha moment of VLM with less than $3.

Python 4,065 283 Updated May 19, 2025

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,177 19,871 Updated Jul 26, 2026

A paper list of some recent works about Token Compress for Vit and VLM

944 43 Updated Jul 20, 2026

A Compressed Stable Diffusion for Efficient Text-to-Image Generation [ECCV'24]

Python 320 20 Updated Jul 6, 2024

A 28× Compressed Wav2Lip for Efficient Talking Face Generation [ICCV'23 Demo] [MLSys'23 Workshop] [NVIDIA GTC'23]

Python 58 5 Updated Mar 8, 2024

Compressed LLMs for Efficient Text Generation [ICLR'24 Workshop]

Python 90 13 Updated Sep 13, 2024

The official NetsPresso Python package.

Jupyter Notebook 50 1 Updated Nov 20, 2025

A library for training, compressing and deploying computer vision models (including ViT) with edge devices

Python 75 11 Updated Sep 29, 2025

Repository for 2023 AI City Challenge (Track1: Multi-Camera People Tracking)

Python 38 6 Updated Oct 7, 2024

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 162,988 34,017 Updated Jul 26, 2026
Next