-
Purdue University
- West Lafayette
-
18:59
(UTC -04:00) - https://dripnowhy.github.io/
Stars
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
OpenClaw-RL: Train any agent simply by talking
[ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Official implementation of Visco-Attack (EMNLP 2025 Main). An open-source one-click reproduction script is also provided.
[NIPS'25 Spotlight] Mulberry, an o1-like Reasoning and Reflection MLLM Implemented via Collective MCTS
A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explor…
[NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
[NeurIPS25 Spotlight] EMPO, A Fully Unsupervised RLVR Method
[NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"
This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
Official code for the paper, "Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning"
FeatureAlignment = Alignment + Mechanistic Interpretability
A curated list of resources dedicated to the safety of Large Vision-Language Models. This repository aligns with our survey titled A Survey of Safety on Large Vision-Language Models: Attacks, Defen…
Code for "Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate" [COLM 2025]
[ICLR 2026] Data and Code for Paper Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
[CVPR2025] T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
Align Anything: Training All-modality Model with Feedback
Robust recipes to align language models with human and AI preferences
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.