Speech and Singing Voice Generation
Speech synthesis, singing voice synthesis, voice conversion, singing technique control, and zero-shot vocal modeling.
I build expressive AI systems for voice, music, and human motion.
Ph.D. student at the National University of Singapore, exploring controllable generation systems that preserve expression, style, and human intent across sound and movement.
I study generative AI for voice, music, and human motion, with an emphasis on controllable systems that support expressive creation.
I work where sound, intelligence, and human expression meet.
I am a Ph.D. student at the School of Computing, National University of Singapore, advised by Prof. Ye Wang in the Sound and Music Computing Lab.
My research covers speech and singing voice synthesis, neural audio codecs, music generation, talking head generation, co-speech gesture generation, and affective multimodal learning.
I am especially interested in models and interfaces that make generated media easier to guide, evaluate, and use in creative or communicative settings.
Before and alongside my research, I have spent many years studying piano and singing. This musical background shapes the questions I care about: how generated sound can remain controllable, expressive, and useful for real creative workflows.
For selected performances and music projects, check my music page.
Three threads connect most of my work: expressive voice generation, music and audio intelligence, and multimodal human signal modeling.
Speech synthesis, singing voice synthesis, voice conversion, singing technique control, and zero-shot vocal modeling.
Neural audio codecs, controllable audio generation, piano transcription/rendering, music style transfer, and creative ML systems.
Talking head generation, co-speech gesture generation, multimodal emotion analysis, and human-centered generation.
A compact timeline for papers, internships, teaching, and milestones.
Research across voice, music, and human-centered multimodal systems. Browse by year, filter by field, or search the collection.
* Equal contribution · † Corresponding author
arXiv preprint, 2026.
EDICT combines reference-based timbre editing with segment-specific expressive instructions, using an edited acoustic reference and KV-cache reconstruction to maintain voice consistency and acoustic continuity.
@misc{zhao2026editwhospeaks,
title = {Edit Who Speaks, Control How They Speak: Global Timbre Editing and Local Instruction Control for {TTS}},
author = {Junchuan Zhao and Chenglin Xu and Wei Zeng and Haoyang Li and Yiwen Guo and Ye Wang},
year = {2026},
eprint = {2610.11437},
archivePrefix = {arXiv},
primaryClass = {cs.SD},
doi = {10.48550/arXiv.2610.11437},
url = {https://arxiv.org/abs/2610.11437}
}
arXiv preprint, 2026.
A study of preference signals, pair construction, and DPO objectives for non-verbal vocalization synthesis, introducing a non-verbal-aware character error rate to balance vocalization realization and lexical accuracy.
@misc{li2026preferenceoptimization,
title = {Preference Optimization for Non-Verbal Vocalization Synthesis},
author = {Haoyang Li and Chenglin Xu and Junchuan Zhao and Yuang Cao and Liumeng Xue and Yiwen Guo and Eng Siong Chng},
year = {2026},
eprint = {2608.24163},
archivePrefix = {arXiv},
primaryClass = {eess.AS},
doi = {10.48550/arXiv.2608.24163},
url = {https://arxiv.org/abs/2608.24163}
}
arXiv preprint, 2026.
arXiv preprint, 2026.
arXiv preprint, 2026.
Interspeech, 2026.
arXiv preprint, 2026.
ACL Main, 2026.
ICASSP, 2026.
ICASSP, 2026.
IEEE TASLP, 2026.
ICLR, 2026.
IEEE/ACM TASLP, 2024.
No publications match this search.
Teaching, academic service, honours, and contact links collected in one place.
For research discussions, collaboration, code, CV, and professional profiles.