Skip to content
View mwxely's full-sized avatar

Block or report mwxely

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Python 187 9 Updated Aug 3, 2026

Open Frontier Intelligence

8,469 660 Updated Aug 6, 2026

Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.

Swift 26,100 2,220 Updated Aug 16, 2026

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Python 23 Updated May 19, 2026

Beyond SFT-to-RL: Pre-alignment via Black-BoxOn-Policy Distillation for Multimodal RL

Python 99 2 Updated May 6, 2026

🔥[VLDB’26] Synthesizing Data Agent Trajectories via Execution-Grounded Tree Search

Python 15 Updated Jul 18, 2026

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

Python 55 2 Updated Jun 2, 2026

[Roadmap] Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

TeX 131 6 Updated Jun 9, 2026

V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

Python 34 3 Updated Apr 16, 2026

mini-claude-code: A minimal, readable Python re-implementation of Claude Code

Python 10 1 Updated Mar 31, 2026
TypeScript 10 6 Updated Aug 10, 2026

VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning.

Python 26 2 Updated Jul 20, 2026

A fork to add multimodal model training to open-r1

Python 1,600 74 Updated Feb 8, 2025

🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.

Python 3,433 210 Updated Mar 5, 2024

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Python 4,364 639 Updated Aug 6, 2026

本人的科研经验

13,730 675 Updated Jun 6, 2026

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python 5,668 574 Updated Aug 15, 2026

The open source coding agent.

TypeScript 197,930 25,491 Updated Aug 16, 2026

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 77,724 6,544 Updated Aug 16, 2026
Python 3 23 Updated Aug 12, 2025

DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation.

Python 142 3 Updated Feb 10, 2026

EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing [ICLR 2026]

Python 158 6 Updated Jul 26, 2026

Incentivizing "Thinking with Long Videos" via Native Tool Calling

Python 6 Updated Jan 5, 2026
Python 2 Updated Dec 22, 2025

Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Python 80 14 Updated Apr 29, 2026

A minimal, educational HEVC (H.265) encoder written in Python.

Python 53 1 Updated Feb 23, 2026

🔥[ICDE'26] Official repository for the paper "CARROT: A Learned Cost-Constrained Retrieval Optimization System for RAG".

Python 22 1 Updated Oct 26, 2025

Official implementation of MATPO: Multi-Agent Tool-Integrated Policy Optimization.

Python 83 38 Updated Oct 31, 2025
Next