Skip to content
View wooksu's full-sized avatar

Organizations

@nota-github

Block or report wooksu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,885 4,364 Updated Aug 9, 2026

Fully Open Framework for Democratized Multimodal Training

Python 1,166 78 Updated Aug 10, 2026

Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Python 392 20 Updated Jun 20, 2026

Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

Python 73,946 9,049 Updated Aug 9, 2026

Official implementation for "TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models" https://arxiv.org/pdf/2602.04657

Python 43 6 Updated Jun 3, 2026

omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode

TypeScript 67,571 5,508 Updated Aug 10, 2026

A simple yet powerful agent framework that delivers with open-source models

Python 4,594 473 Updated Mar 21, 2026

ERGO (Efficient Reasoning & Guided Observation) is a large vision-language model trained with reinforcement learning on efficiency objectives. [ICLR'26]

Python 19 1 Updated Feb 25, 2026

[ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

Python 92 1 Updated Jan 26, 2026
Python 46 1 Updated Jul 14, 2025

Nano vLLM

Python 14,925 2,440 Updated Apr 26, 2026

[NeurIPS 2025] Official code for paper: Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

Python 106 5 Updated Sep 20, 2025

Official code for NeurIPS 2025 paper "GRIT: Teaching MLLMs to Think with Images"

Python 189 10 Updated Jan 16, 2026
Python 1,262 78 Updated Nov 20, 2025

Open-source unified multimodal model

Python 6,141 545 Updated May 4, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,439 68 Updated Aug 2, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,105 383 Updated Jul 30, 2026

State-of-the-art Image & Video CLIP, Multimodal Large Language Models, and More!

Jupyter Notebook 2,335 160 Updated Apr 13, 2026

Solve Visual Understanding with Reinforced VLMs

Python 6,017 385 Updated Jul 7, 2026

Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’

Jupyter Notebook 2,267 110 Updated Oct 29, 2025

Witness the aha moment of VLM with less than $3.

Python 4,062 282 Updated May 19, 2025

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 88,610 20,486 Updated Aug 10, 2026

A paper list of some recent works about Token Compress for Vit and VLM

944 46 Updated Aug 7, 2026

A Compressed Stable Diffusion for Efficient Text-to-Image Generation [ECCV'24]

Python 321 21 Updated Jul 6, 2024

A 28× Compressed Wav2Lip for Efficient Talking Face Generation [ICCV'23 Demo] [MLSys'23 Workshop] [NVIDIA GTC'23]

Python 58 5 Updated Mar 8, 2024

Compressed LLMs for Efficient Text Generation [ICLR'24 Workshop]

Python 90 13 Updated Sep 13, 2024

The official NetsPresso Python package.

Jupyter Notebook 50 1 Updated Nov 20, 2025

A library for training, compressing and deploying computer vision models (including ViT) with edge devices

Python 75 11 Updated Sep 29, 2025

Repository for 2023 AI City Challenge (Track1: Multi-Camera People Tracking)

Python 38 6 Updated Oct 7, 2024

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 163,506 34,161 Updated Aug 9, 2026
Next