| Uniaudio: An audio foundation model toward universal audio generation D Yang, J Tian, X Tan, R Huang, S Liu, X Chang, J Shi, S Zhao, J Bian, ... arXiv preprint arXiv:2310.00704, 2023 | 320* | 2023 |
| Kimi-audio technical report D Ding, Z Ju, Y Leng, S Liu, T Liu, Z Shang, K Shen, W Song, X Tan, ... arXiv preprint arXiv:2504.18425, 2025 | 296* | 2025 |
| Hifi-codec: Group-residual vector quantization for high fidelity audio codec D Yang, S Liu, R Huang, J Tian, C Weng, Y Zou arXiv preprint arXiv:2305.02765, 2023 | 262 | 2023 |
| Longcat-flash-omni technical report MLC Team, B Wang, B Xiao, B Zhang, B Rong, B Chen, C Wan, C Zhang, ... arXiv preprint arXiv:2511.00279, 2025 | 209* | 2025 |
| Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt D Yang, S Liu, R Huang, C Weng, H Meng IEEE/ACM Transactions on Audio, Speech, and Language Processing 32, 2913-2925, 2024 | 187 | 2024 |
| Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens X Wang, M Jiang, Z Ma, Z Zhang, S Liu, L Li, Z Liang, Q Zheng, R Wang, ... arXiv preprint arXiv:2503.01710, 2025 | 180 | 2025 |
| Any-to-Many Voice Conversion with Location-Relative Sequence-to-Sequence Modeling S Liu, Y Cao, D Wang, X Wu, X Liu, H Meng IEEE/ACM Transactions on Audio Speech and Language Processing, 2020 | 166 | 2020 |
| Speech emotion recognition using capsule networks X Wu, S Liu, Y Cao, X Li, J Yu, D Dai, X Ma, S Hu, Z Wu, X Liu, H Meng ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and …, 2019 | 160 | 2019 |
| DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs S Liu, D Su, D Yu ICML 2022 Workshop on Machine Learning for Audio Synthesis, 2022 | 105 | 2022 |
| Diffsvc: A diffusion probabilistic model for singing voice conversion S Liu, Y Cao, D Su, H Meng 2021 IEEE automatic speech recognition and understanding workshop (ASRU …, 2021 | 104 | 2021 |
| The singing voice conversion challenge 2023 WC Huang, LP Violeta, S Liu, J Shi, T Toda 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 1-8, 2023 | 103 | 2023 |
| Adversarial attacks on spoofing countermeasures of automatic speaker verification S Liu, H Wu, H Lee, H Meng 2019 IEEE automatic speech recognition and understanding workshop (ASRU …, 2019 | 93 | 2019 |
| Defense against adversarial attacks on spoofing countermeasures of ASV H Wu, S Liu, H Meng, H Lee ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and …, 2020 | 87 | 2020 |
| Voice conversion across arbitrary speakers based on a single target-speaker utterance S Liu, J Zhong, L Sun, X Wu, X Liu, H Meng Proc. Interspeech 2018, 496-500, 2018 | 72 | 2018 |
| End-to-end voice conversion via cross-modal knowledge distillation for dysarthric speech reconstruction D Wang, J Yu, X Wu, S Liu, L Sun, X Liu, H Meng ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and …, 2020 | 67 | 2020 |
| End-to-end accent conversion without using native utterances S Liu, D Wang, Y Cao, L Sun, X Wu, S Kang, Z Wu, X Liu, D Su, D Yu, ... ICASSP 2020, 2020 | 67 | 2020 |
| Fastsvc: Fast cross-domain singing voice conversion with feature-wise linear modulation S Liu, Y Cao, N Hu, D Su, H Meng 2021 ieee international conference on multimedia and expo (icme), 1-6, 2021 | 64 | 2021 |
| Simplespeech 2: Towards simple and efficient text-to-speech with flow-based scalar latent transformer diffusion models D Yang, R Huang, Y Wang, H Guo, D Chong, S Liu, X Wu, H Meng IEEE Transactions on Audio, Speech and Language Processing 33, 2634-2646, 2025 | 59 | 2025 |
| End-to-end code-switched tts with mix of monolingual recordings Y Cao, X Wu, S Liu, J Yu, X Li, Z Wu, X Liu, H Meng ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and …, 2019 | 56 | 2019 |
| Speech emotion recognition using sequential capsule networks X Wu, Y Cao, H Lu, S Liu, D Wang, Z Wu, X Liu, H Meng IEEE/ACM Transactions on Audio, Speech, and Language Processing 29, 3280-3291, 2021 | 49 | 2021 |