Stars
High performance self-hosted photo and video management solution.
Chinese Mandarin Grapheme-to-Phoneme Converter. 中文轉注音或拼音 (INTERSPEECH 2022)
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
Implementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in Pytorch
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
搞定C++:punch:。C++ Primer 中文版第5版学习仓库,包括笔记和课后练习答案。
Solutions to Exercises in C++ Primer 5th Edition
Clone a voice in 5 seconds to generate arbitrary speech in real-time
⚡️ A minimal portfolio template for Developers
Self-Supervised Speech Pre-training and Representation Learning Toolkit
The PyTorch-based audio source separation toolkit for researchers
A face recognition solution on mobile device.
A MNIST-like fashion product database. Benchmark 👇
Code for Switchable Normalization from "Differentiable Learning-to-Normalize via Switchable Normalization", https://arxiv.org/abs/1806.10779
A Simple Tool to Evaluate Your Models on Megaface Benchmark Implemented in Python and Mxnet
MMdnn is a set of tools to help users inter-operate among different deep learning frameworks. E.g. model conversion and visualization. Convert models between Caffe, Keras, MXNet, Tensorflow, CNTK, …
Command-line program to download videos from YouTube.com and other video sites
Supporting code for my article on video streaming with Flask.
Simple Online Realtime Tracking with a Deep Association Metric
Models and examples built with TensorFlow
The world's simplest facial recognition api for Python and the command line
Deformable Convolutional Networks
State-of-the-art 2D and 3D Face Analysis Project