-
University of Maryland
Highlights
- Pro
Lists (4)
Sort Name ascending (A-Z)
Starred repositories
A benchmark built to evaluate and improve agent capabilities for supporting legal work.
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
[ICLR 2026] "Does FLUX Already Know How to Perform Physically Plausible Image Composition?" (Official Implementation)
Evaluating repetitive behavior in Vision Language Models
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
A paper list about Token Merge, Reduce, Resample, Drop for MLLMs.
RoboChallenge Working Group Community, Committee related discussion and planning. RoboChallenge 工作组社区、委员会相关讨论和规划
RoboChallenge Inference example code
A collection of awesome video generation studies.
💃 Dance with LLM in Your Code. Minuet offers code completion as-you-type from popular LLMs including OpenAI, Gemini, Claude, Ollama, Llama.cpp, Codestral, and more.
Automation of Interactive Brokers TWS. You can download the latest release here: https://github.com/ibcalpha/ibc/releases/latest
[EMNLP 2018] PyTorch code for TVQA: Localized, Compositional Video Question Answering
Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
Reading list for research topics in multimodal machine learning
A Survey on multimodal learning research.
🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
Large Language-and-Vision Assistant for Biomedicine, built towards multimodal GPT-4 level capabilities.
A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites
Awesome-LLM-3D: a curated list of Multi-modal Large Language Model in 3D world Resources
Docker image that provides a Minecraft Server for Java Edition that automatically installs/upgrades versions, modloaders, modpacks and more at startup
collection of diffusion model papers categorized by their subareas
Layered Score Distillation for Disentangled Object Relighting
A collection of papers on neural field-based inverse rendering.
Automatically generate AirPrint Avahi service files for CUPS printers