Skip to content
View Max1Wz's full-sized avatar
  • NJU AALab
  • Nanjing, China

Block or report Max1Wz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Python 14 1 Updated Aug 11, 2026

[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Python 782 30 Updated Aug 7, 2026
23 Updated Jun 12, 2026

Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models

1,189 50 Updated Jul 14, 2026
Python 1,263 78 Updated Nov 20, 2025

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …

Python 15,117 1,586 Updated Aug 11, 2026

A curated list of models, benchmarks, tools and guides for audio editing

35 6 Updated Aug 6, 2026

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

Python 1,918 149 Updated Jul 29, 2026

Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.

Python 592 200 Updated Jul 17, 2026

MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.

Python 1,460 80 Updated Jul 24, 2026
Python 120 7 Updated Jun 5, 2026

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,771 1,832 Updated Jan 30, 2026

Interspeech 2026 audio demo website for D-SKD.

HTML 2 Updated Jun 16, 2026

JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System

Python 1,703 173 Updated Aug 11, 2026

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Python 369 4 Updated Aug 5, 2026

video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is developed by the Department of Electronic Engineering at Tsin…

Python 208 26 Updated Feb 23, 2026

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling

Python 210 6 Updated Jun 6, 2026

[NeurIPS'2025] Official repository for "LiveStar: Live Streaming Assistant for Real-World Online Video Understanding"

Python 156 7 Updated Jul 3, 2026

Structured Video Comprehension of Real-World Shorts

Python 239 7 Updated Sep 21, 2025

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)

Python 472 57 Updated Oct 29, 2025

StreamingVLM: Real-Time Understanding for Infinite Video Streams

Python 1,063 65 Updated Oct 15, 2025

A web-based annotation tool for synchronized multi-video timeline labeling and AI-assisted question generation, built for the GameplayQA benchmark.

TypeScript 20 2 Updated Jul 13, 2026

🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖

439 27 Updated Aug 5, 2026

Create beautiful slides on the web using a coding agent's frontend skills

JavaScript 27,298 2,216 Updated Jun 23, 2026

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

121 Updated May 20, 2026

Public repository for Agent Skills

Python 167,827 20,009 Updated Aug 7, 2026

Official code for "WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling"

Python 63 7 Updated Jun 27, 2026

Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with synthetic captions.

Python 135 12 Updated Mar 3, 2026
Next