-
NJU AALab
- Nanjing, China
Stars
[Official Repo] JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
Reading notes about Multimodal Large Language Models, Large Language Models, and Diffusion Models
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
A curated list of models, benchmarks, tools and guides for audio editing
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Hy3 (295B A21B), a leading reasoning and agent model in its size, with great cost efficiency.
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
JoyAI-VL-Interaction: An Open Real-time Video-Language Interaction System
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions, which is developed by the Department of Electronic Engineering at Tsin…
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
[NeurIPS'2025] Official repository for "LiveStar: Live Streaming Assistant for Real-World Online Video Understanding"
Structured Video Comprehension of Real-World Shorts
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)
StreamingVLM: Real-Time Understanding for Infinite Video Streams
A web-based annotation tool for synchronized multi-video timeline labeling and AI-assisted question generation, built for the GameplayQA benchmark.
🔥🔥🔥 [Awesome] Latest Papers, Codes & Datasets on Streaming / Online Video Understanding — Building Always-on, Real-time Video AI 🤖
Create beautiful slides on the web using a coding agent's frontend skills
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
Official code for "WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling"
Elucidated Text-To-Audio (ETTA) is a SOTA text-to-audio model with a holistic understanding of the design space and trained with synthetic captions.