Skip to content
View abePclWaseda's full-sized avatar
💭
研究用(アルゴリズムエンジニア寄り)
💭
研究用(アルゴリズムエンジニア寄り)

Highlights

  • Pro

Block or report abePclWaseda

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[CVPR 2026] Official implementation of MultiAnimate: Pose-Guided Image Animation Made Extensible

Python 16 Updated Jun 28, 2026
Jupyter Notebook 8 Updated Jun 12, 2026

Long-form streaming TTS system for multi-speaker dialogue generation

Python 1,413 126 Updated Oct 26, 2025

The python library for real-time communication

JavaScript 4,616 434 Updated Jan 12, 2026

ARC-AGI Toolkit

Python 64 23 Updated Jun 10, 2026

Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Python 8,660 1,120 Updated Sep 14, 2024

Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS

Python 2,111 199 Updated Jul 24, 2026

A framework for few-shot evaluation of language models.

Python 13,402 3,439 Updated Jul 13, 2026

τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Python 1,658 419 Updated Jul 24, 2026

The Full-Duplex Interaction Track of the ICASSP 2026 Human-like Spoken Dialogue Systems Challenge aims to advance the evaluation of full-duplex dialogue systems by in- troducing a dual-channel dial…

Python 36 Updated Apr 27, 2026

A lightweight library for evaluating vision language models.

Python 5 Updated May 26, 2026

The repository for Springer IJCV 2025 (LR-ASD: Lightweight and Robust Network for Active Speaker Detection)

Python 132 27 Updated Mar 23, 2025

A suite of image and video neural tokenizers

Jupyter Notebook 1,731 90 Updated Feb 11, 2025

[NeurIPS 2025] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Python 2,975 492 Updated May 22, 2026

Run agents like Hermes, LangChain Deep Agents, and OpenClaw more securely inside NVIDIA OpenShell with managed inference

TypeScript 21,910 2,974 Updated Jul 24, 2026

A fast and soft pattern search for trillion-scale corpora.

Python 237 11 Updated Feb 28, 2026

Versatile Evaluation of Speech and Audio

Python 424 50 Updated Jul 21, 2026

Kyutai with an "eye"

Python 254 34 Updated Mar 26, 2025
Jupyter Notebook 509 89 Updated Jul 3, 2026

Transformer with Local Modeling by Convolution for Speech Separation and Enhancement

Python 133 9 Updated Aug 8, 2025

MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols

JavaScript 20 1 Updated Nov 19, 2025

Repository for the SIGDIAL 2026 website

JavaScript 1 Updated Jul 20, 2026

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.

Jupyter Notebook 3,904 278 Updated Apr 23, 2026

Benchmarking scripts for Gaia

Python 17 1 Updated Apr 10, 2025

Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.

Python 1 Updated Jan 14, 2026

Code for the paper Hybrid Spectrogram and Waveform Source Separation

Python 10,343 1,546 Updated Apr 24, 2024

Official Implementation of NAACL 2025 Paper: Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models

18 1 Updated Apr 30, 2025
Next