IEEE / ACM Transactions on Audio, Speech, and Language Processing · 2024

SinTechSVS — A Singing Technique Controllable Singing Voice Synthesis System

Junchuan Zhao · Low Qi Hong Chetwin · Ye Wang
National University of Singapore, School of Computing
Eight singing techniques surrounding the SVS core A radial diagram showing four pitch techniques — scooping, bend, melisma, drop — on the left, and four timbre techniques — vocal fry, falsetto, breathy, belting — on the right, all connected to a central SVS node. Each technique is clickable and jumps to its section. SinTechSVS core system Scooping Bend Melisma Drop Vocal Fry Falsetto Breathy Belting
Pitch techniques Timbre techniques

Click a technique to jump to its samples.

01 — Abstract

Closing the gap between synthesis and human expressivity

The precise control of singing techniques is of utmost importance in achieving emotionally expressive vocal performances. To bridge the gap between current Singing Voice Synthesis (SVS) systems and human singers, our paper focuses on developing an SVS system that allows for control over singing techniques.

We introduce SinTechSVS, a singing technique controllable SVS system composed of a singing technique annotator, a singing technique controllable synthesizer, and a singing technique recommender. Our approach leverages transfer learning for efficient singing technique annotation and adapts the DiffSinger framework with additional style encoders and an attention-based singing technique local score (STLS) module to enhance singing technique controllability. We also propose a Seq2Seq singing technique recommender for the new task of Singing Technique Recommendation (STR).

Experimental results demonstrate that SinTechSVS significantly improves the quality and expressiveness of synthesized vocal performances, with comparable general synthesis capabilities to state-of-the-art SVS systems and enhanced control over singing techniques, as evidenced by objective and subjective evaluations. To the best of our knowledge, SinTechSVS is the first SVS capable of controlling singing techniques.

02 — System

Overall architecture

SinTechSVS consists of three key components: a singing technique annotator (STA), a singing voice synthesizer conditioned on singing techniques (SVS), and a singing technique recommender (STR). The "OR" symbol denotes that the SVS input is either a user-specified technique sequence or the predicted sequence from the STR.

Overall architecture diagram of SinTechSVS showing the singing technique annotator, the technique-conditioned synthesizer, and the technique recommender.
Training and inference pipeline diagram for SinTechSVS across three stages.
The training process consists of three steps, each laying the foundation for the next. Modules with full shadows remain unfixed during the step; modules with half shadows are first fixed, then unfixed. (Top-left) Training of STA. (Bottom-left) Training of SVS. (Bottom-right) Inference of SinTechSVS.
03 — Ground Truth

Singing technique annotations

Samples of singing techniques manually annotated on the Opencpop dataset. In sentence-level samples, bolded syllables are sung in the specified technique.

Scooping approaching the target pitch from below
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
具(ju) 象(xiang)
Bend pitch deviates and returns within a note
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
没(mei) 限(xian) 期(qi)
Drop pitch falls away at the end of a note
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
再(zai) 给(gei) 我(wo) 两(liang) 分(fen) 钟(zhong)
Melisma multiple pitches sung across one syllable
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
双(shuang) 眼(yan)
Vocal Fry low, creaky glottal vibration
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
我(wo) 们(men) 都(dou) 需(xu) 要(yao) 勇(yong) 气(qi)
Falsetto light, breathy upper-register voice
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
没(mei) 有(wo) 你(you) 根(gen) 本(ben) 不(bu) 想(xiang) 逃(tao)
Breathy audible airflow blended into the tone
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
我(wo) 不(bu) 会(hui) 发(fa) 现(xian) 我(wo) 难(nan) 受(shou)
Belting powerful chest-register projection
Word-level Sample 1
Word-level Sample 2
Sentence-level Sample
一(yi) 辈(bei) 子(zi) 暖(nuan) 暖(nuan) 的(de) 好(hao)
04 — Resources

Data acquirement and annotation statistics

Download the annotation file

The manual singing technique annotation file for the Opencpop dataset is available for direct download from the SinTechSVS annotation dataset on Hugging Face. The annotation is for research purposes only.

For the Opencpop dataset itself, please strictly follow the instructions at wenet.org.cn/opencpop — we have no right to grant access to it directly.

Usage requires

  • Research purposes only
  • Agreement to the license

Distribution of manually annotated Opencpop labels

Bar chart showing the distribution of manually annotated singing technique labels in the Opencpop dataset.

"Whisper" and "hiccup" are removed due to the small amount of available labels.

Distribution of duration per technique

Violin plot showing the duration distribution of each singing technique.

Mel-spectrograms of the timbre and pitch singing techniques

Mel-spectrograms of timbral and pitch singing techniques.
05 — Results

Synthesis with singing technique control

Synthesized samples conditioned on singing techniques. In the word-level lyric sequence, bolded syllables are sung in the specified technique. Regular / Straight denotes synthesis without any technique conditioning, serving as a reference for comparison.

Scooping
冰 刀 的 圈 bing dao hua de quan
Regular / Straight
SinTechSVS
Bend
又 无 可 you wu ke nai he
Regular / Straight
SinTechSVS
Drop
小 火 车 摆 的 旋 律 xiao huo che bai dong de xuan lv
Regular / Straight
SinTechSVS
Melisma
你 在 世 俗 里 的 名 字 被 人 用 ni zai shi su li de ming zi bei ren yong le
Regular / Straight
SinTechSVS
Vocal Fry
喔 喔 wo wo
Regular / Straight
SinTechSVS
Falsetto
我 恨 你 wo hen ni
Regular / Straight
SinTechSVS
Breathy
很 少 人 看 诗 hen shao ren kan shi
Regular / Straight
SinTechSVS
Belting
你 好 吗 ni hao ma
Regular / Straight
SinTechSVS
06 — Results

Synthesis with singing technique recommendation

Synthesized samples conditioned on singing techniques recommended directly from the music score. STan denotes SinTechSVS using annotated (ground-truth) technique labels; SinTechSVS uses the STR module to predict techniques for the input.

Ground Truth STan SinTechSVS
07 — Results

Recommendation on unseen music scores

Technique recommendations and corresponding synthesis on previously unseen score samples. Pitch abbreviations: STR straight · SCO scooping · BEND bend · DROP drop · MEL melisma. Timbre abbreviations: REG regular · FRY vocal fry · FAL falsetto · BRE breathy · BEL belting. Word-level pitch, lyric, slur, and technique sequences are separated by "|".

Unseen Sample 1

Lyrics但(dan) | 我(wo) | 早(zao) | 已(yi) | 学(xue) | 会(hui) | 一(yi) | 个(ge) | 人(ren) | 想(xiang) | 你(ni)
PitchA3 | F#4/Gb4 E4 D4 | F#4/Gb4 | B4 | C#5/Db5 | B4 | A4 | F#4/Gb4 | E4 | D4 | F#4/Gb4 E4 D4
Slur0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1
Pitch Tech.STRMELREGREGSCOREGREGDROPREGREGMEL
Timbre Tech.REGBELBELFALFALFALFALBREBREREGBRE
Synthesized Audio

Unseen Sample 2

Lyrics看(kan) | 着(zhe) | 我(wo) | 的(de) | 脚(jiao) | 印(yin) | 一(yi) | 个(ge) | 人(ren) | 一(yi) | 步(bu) | 步(bu) | 好(hao) | 寂(ji) | 寞(mo)
PitchF4 | G4 | E4 | D4 | C4 | D4 C4 A4 | E4 | F4 | C5 | E4 | F4 | C5 | E4 | F4 | F4
Slur0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1
Pitch Tech.STRSTRSCOSTRREGMELSTRSTRSCOSTRSTRSCOSTRSTRSCO
Timbre Tech.BELBELBELBELBELREGREGREGFALREGREGFALREGREGBEL
Synthesized Audio

Unseen Sample 3

Lyrics有(you) | 什(shen) | 么(me) | 方(fang) | 法(fa) | 让(rang) | 自(zi) | 己(ji) | 真(zhen) | 的(de) | 忘(wang) | 记(ji) | ha | ha
PitchG4 | A4 | A#4/Bb4 | A4 F4 | C4 | A#4/Bb4 | A4 | F4 | A#4/Bb4 | A4 | A#4/Bb4 | C5 | A#4/Bb4 A4 | A4 G4 F4
Slur0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 1
Pitch Tech.STRSTRSCOBENDBENDSTRSTRSTRBENDSTRSTRSCODROPMEL
Timbre Tech.REGREGREGBREBREBELBELBELBREBREBREFALBREBRE
Synthesized Audio
08 — Reference

Citation

If you use SinTechSVS in your research, please cite the IEEE/ACM Transactions on Audio, Speech, and Language Processing journal paper.

@article{zhao2024sintechsvs,
  title={Sintechsvs: A singing technique controllable singing voice synthesis system},
  author={Zhao, Junchuan and Chetwin, Low Qi Hong and Wang, Ye},
  journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
  volume={32},
  pages={2641--2653},
  year={2024},
  publisher={IEEE}
}
09 — Terms

Contact and license

If you have any questions about the paper, please contact the first author Junchuan at junchuan@comp.nus.edu.sg.

  • The singing technique annotation for Opencpop is available to download for non-commercial purposes under the Opencpop license.
  • This annotation may not be sold, leased, published, or distributed to any third party without written permission from the administrator.
  • The National University of Singapore is not responsible for errors in the annotation's content or any damages resulting from its use. The administrator may update these conditions of use at any time.