Skip to content
View DripNowhy's full-sized avatar
🐢
go
🐢
go

Block or report DripNowhy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 4 Updated May 13, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,017 109,234 Updated Aug 6, 2026

OpenClaw-RL: Train any agent simply by talking

Python 5,627 607 Updated May 23, 2026

[ICML 2026] Official implementation for paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation

Python 16 1 Updated Jun 4, 2026

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python 5,104 383 Updated Jul 30, 2026
HTML 2 Updated Mar 5, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,882 4,364 Updated Aug 9, 2026

Official implementation of Visco-Attack (EMNLP 2025 Main). An open-source one-click reproduction script is also provided.

Python 31 1 Updated Apr 11, 2026

[NIPS'25 Spotlight] Mulberry, an o1-like Reasoning and Reflection MLLM Implemented via Collective MCTS

Python 1,244 113 Updated Jan 16, 2026

A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explor…

217 5 Updated Mar 4, 2026

[NeurIPS 2025] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

Python 82 6 Updated May 31, 2025

[NeurIPS25 Spotlight] EMPO, A Fully Unsupervised RLVR Method

Python 104 4 Updated Nov 24, 2025

[NeurIPS 2025] Official Implementation of paper "Sherlock: Self-Correcting Reasoning in Vision-Language Models"

Python 31 1 Updated Jun 4, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,439 68 Updated Aug 2, 2026

Official code for the paper, "Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning"

Python 173 7 Updated Oct 23, 2025

FeatureAlignment = Alignment + Mechanistic Interpretability

Python 35 1 Updated Mar 8, 2025

A curated list of resources dedicated to the safety of Large Vision-Language Models. This repository aligns with our survey titled A Survey of Safety on Large Vision-Language Models: Attacks, Defen…

216 16 Updated May 25, 2026

Code for "Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate" [COLM 2025]

Python 182 8 Updated Jul 8, 2025

[ICLR 2026] Data and Code for Paper Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models

JavaScript 12 Updated Feb 2, 2026

[CVPR2025] T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

34 Updated Jul 10, 2025

Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks

Python 4,334 742 Updated Aug 7, 2026

Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.

Python 27,492 2,030 Updated Jan 9, 2026

An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.

Python 39,518 4,789 Updated May 1, 2026

Align Anything: Training All-modality Model with Feedback

Python 4,663 505 Updated Nov 27, 2025

AllenAI's post-training codebase

Python 3,821 573 Updated Aug 8, 2026

Robust recipes to align language models with human and AI preferences

Python 5,657 489 Updated May 26, 2026

An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

Python 9,898 996 Updated Jul 14, 2026

Implementation of the training framework proposed in Self-Rewarding Language Model, from MetaAI

Python 1,412 70 Updated Apr 11, 2024

An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.

Jupyter Notebook 2,011 314 Updated Aug 9, 2025
Next