-
Department of Computer Science and Engineering, University at Buffalo
- Buffalo, NY
Stars
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
[ICLR 2026] FOCUS: Efficient Keyframe Selection for Long Video Understanding
[ACL 2026] CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation
Unlock your displays on your Mac! Flexible HiDPI scaling, XDR/HDR extra brightness, virtual screens, DDC control, extra dimming, PIP/streaming, EDID override and lots more!
✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
WikiVideo: Article Generation from Multiple Videos
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and te…
EVA Series: Visual Representation Fantasies from BAAI
[ICLR2025 Oral] ChartMoE: Mixture of Diversely Aligned Expert Connector for Chart Understanding
The official repo of "On the Perception Bottleneck of VLMs for Chart Understanding"
[ECCV 2024] official code for "Long-CLIP: Unlocking the Long-Text Capability of CLIP"
Machine Learning and Computer Vision Engineer - Technical Interview Questions
The proposed simulated dataset consisting of 9,536 charts and associated data annotations in CSV format.
The interactive graphing library for Python ✨
[NeurIPS 2024] CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
A curated list of recent and past chart understanding work based on our IEEE TKDE survey paper: From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Mod…
ORLM: Training Large Language Models for Optimization Modeling