Skip to content
View pengyu965's full-sized avatar
  • Department of Computer Science and Engineering, University at Buffalo
  • Buffalo, NY

Block or report pengyu965

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Awesome List for On-Policy Distillation

817 18 Updated Jul 31, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,910 4,376 Updated Aug 11, 2026

[ICLR 2026] FOCUS: Efficient Keyframe Selection for Long Video Understanding

Python 76 5 Updated Apr 23, 2026

[ACL 2026] CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

Python 5 Updated May 26, 2026

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Python 8 Updated Jul 7, 2026
Python 1 Updated May 26, 2026

TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

Python 6 Updated Jun 10, 2026

Unlock your displays on your Mac! Flexible HiDPI scaling, XDR/HDR extra brightness, virtual screens, DDC control, extra dimming, PIP/streaming, EDID override and lots more!

33,107 629 Updated Jul 23, 2026

✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

789 30 Updated Dec 8, 2025

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 77,332 6,516 Updated Aug 10, 2026
3 Updated Mar 4, 2026

WikiVideo: Article Generation from Multiple Videos

15 Updated Nov 14, 2025
Python 550 51 Updated May 10, 2026
JavaScript 5,159 319 Updated Aug 3, 2026
Python 3 1 Updated Oct 8, 2025

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,770 1,832 Updated Jan 30, 2026

An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.

Python 1,952 220 Updated Jul 25, 2026

StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and te…

Python 4,517 258 Updated Nov 7, 2025

EVA Series: Visual Representation Fantasies from BAAI

Python 2,688 187 Updated Aug 1, 2024

[ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding

Jupyter Notebook 102 10 Updated Apr 1, 2025

The official repo of "On the Perception Bottleneck of VLMs for Chart Understanding"

Jupyter Notebook 10 Updated Apr 12, 2025

[ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"

Python 901 48 Updated Aug 13, 2024

Machine Learning and Computer Vision Engineer - Technical Interview Questions

4,762 773 Updated Jan 24, 2026

The proposed simulated dataset consisting of 9,536 charts and associated data annotations in CSV format.

26 1 Updated Feb 22, 2024

The interactive graphing library for Python ✨

Python 18,726 2,835 Updated Aug 7, 2026

[NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Python 160 15 Updated Apr 22, 2025

A curated list of recent and past chart understanding work based on our IEEE TKDE survey paper: From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Mod…

242 25 Updated Dec 17, 2025

ORLM: Training Large Language Models for Optimization Modeling

Python 272 43 Updated Sep 18, 2025
Next