Skip to content
View shikiw's full-sized avatar

Block or report shikiw

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,053 109,113 Updated Aug 6, 2026

(ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’

Python 60 2 Updated Jan 26, 2026

[NeurIPS 2025] Official implementation of HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance

Python 88 1 Updated Sep 18, 2025

[ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incen…

Python 1,584 27 Updated Mar 20, 2026

CYaRon: Yet Another Random Olympic-iNformatics test data generator

Python 1,649 181 Updated Mar 21, 2026

[NIPS'25 Spotlight] Mulberry, an o1-like Reasoning and Reflection MLLM Implemented via Collective MCTS

Python 1,243 113 Updated Jan 16, 2026

[ICCV 2025] MM-IFEngine: Towards Multimodal Instruction Following

Python 126 Updated Feb 13, 2026

Scalable RL solution for advanced reasoning of language models

Python 1,867 116 Updated Mar 18, 2025

MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Python 770 30 Updated Sep 7, 2025

Train transformer language models with reinforcement learning.

Python 19,079 2,910 Updated Aug 15, 2026
Python 1,179 58 Updated Jan 10, 2026

This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!

1,440 68 Updated Aug 2, 2026

Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’

Jupyter Notebook 2,268 110 Updated Oct 29, 2025

[ICML 2025] SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation

Python 316 33 Updated Nov 5, 2025

[ICCV 2025] Light-A-Video: Training-free Video Relighting via Progressive Light Fusion

Python 519 33 Updated Oct 25, 2025

[ICML 2025 Oral] An official implementation of VideoRoPE & VideoRoPE++

Python 224 5 Updated Apr 15, 2026

official code for "BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning"

Python 37 3 Updated Jan 21, 2025

📖 A curated list of resources dedicated to hallucination of multimodal large language models (MLLM).

1,036 47 Updated Sep 27, 2025

Next-Token Prediction is All You Need

Python 2,437 100 Updated Jan 12, 2026

[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Python 540 42 Updated Feb 10, 2025

open-source code for paper: Retrieval Head Mechanistically Explains Long-Context Factuality

Python 242 27 Updated Aug 2, 2024

Official implement of MIA-DPO

Python 69 4 Updated Jan 23, 2025

(CVPR 2025) PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Python 151 3 Updated Mar 6, 2025

[NeurIPS 2024] Official PyTorch implementation of LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Python 49 3 Updated Jan 14, 2025

[EMNLP 2024 Findings] ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs

Python 30 2 Updated May 22, 2025

[CVPR 2025] Official implementation of ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way

48 Updated Oct 10, 2025

Code for paper "Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models"

Python 241 4 Updated May 24, 2024

[ICCV-2025] Official implementation of Bootstrap3D: Improving Multi-view Diffusion Model with Synthetic Data

Python 94 2 Updated Jul 26, 2025

[ICCV 2025] The official code of the paper "Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate".

Python 113 2 Updated Jul 9, 2025
Next