-
Waseda univ.
- Tokyo
-
07:23
(UTC +09:00) - https://yutoabe.com/
- @ReactYuto
- https://qiita.com/yuAbe
- https://www.docswell.com/user/yuAbe
Highlights
- Pro
Lists (9)
Sort Name ascending (A-Z)
Stars
[CVPR 2026] Official implementation of MultiAnimate: Pose-Guided Image Animation Made Extensible
Long-form streaming TTS system for multi-speaker dialogue generation
The python library for real-time communication
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS
A framework for few-shot evaluation of language models.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-duplex dialogue systems by in- troducing a dual-channel dial…
A lightweight library for evaluating vision language models.
The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)
A suite of image and video neural tokenizers
[NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference
A fast and soft pattern search for trillion-scale corpora.
Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
emorynlp / SIGDIAL-2026
Forked from acl-org/acl-2023Repository for the SIGDIAL 2026 website
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
Saeki-M / moshi
Forked from kyutai-labs/moshiMoshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
Code for the paper Hybrid Spectrogram and Waveform Source Separation
Official Implementation of NAACL 2025 Paper: Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models