论文信息 - Beijing Opera Synthesis Based on Straight Algorithm and Deep Learning

Beijing Opera Synthesis Based on Straight Algorithm and Deep Learning

Speech synthesis is an important research content in the field of human-computer interaction and has a wide range of applications. As one of its branches, singing synthesis plays an important role. Beijing Opera is a famous traditional Chinese opera, and it is called Chinese quintessence. The singing of Beijing Opera carries some features of speech but it has its own unique pronunciation rules and rhythms which differ from ordinary speech and singing. In this paper, we propose three models for the synthesis of Beijing Opera. Firstly, the speech signals of the source speaker and the target speaker are extracted by using the straight algorithm. And then through the training of GMM, we complete the voice control model to input the voice to be converted and output the voice after the voice conversion. Finally, by modeling the fundamental frequency, duration, and frequency separately, a melodic control model is constructed using GAN to realize the synthesis of the Beijing Opera fragment. We connect the fragments and superimpose the background music to achieve the synthesis of Beijing Opera. The experimental results show that the synthesized Beijing Opera has some audibility and can basically complete the composition of Beijing Opera. We also extend our models to human-AI cooperative music generation: given a target voice of human, we can generate a Beijing Opera which is sung by a new target voice.

Wei Zhao | Cong Jin | Xueting Wang

[1] Andrew W. Senior,et al. Fast and accurate recurrent neural network acoustic models for speech recognition , 2015, INTERSPEECH.

[2] Soumith Chintala,et al. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks , 2015, ICLR.

[3] Yoshua Bengio,et al. Generative Adversarial Nets , 2014, NIPS.

[4] Thomas Brox,et al. Synthesizing the preferred inputs for neurons in neural networks via deep generator networks , 2016, NIPS.

[5] Chung-Hsien Wu,et al. HMM-based Mandarin Singing Voice Synthesis Using Tailored Synthesis Units and Question Sets , 2013, Int. J. Comput. Linguistics Chin. Lang. Process..

[6] Hideki Kawahara,et al. YIN, a fundamental frequency estimator for speech and music. , 2002, The Journal of the Acoustical Society of America.

[7] D. Schwarz,et al. Corpus-Based Concatenative Synthesis , 2007, IEEE Signal Processing Magazine.

[8] Mark A. Clements,et al. A singing voice synthesis system based on sinusoidal modeling , 1997, 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing.

[9] J. Bonada,et al. Synthesis of the Singing Voice by Performance Sampling and Spectral Models , 2007, IEEE Signal Processing Magazine.

[10] Hung-Yan Gu,et al. Mandarin singing voice synthesis using ANN vibrato parameter models , 2008, ICMLC 2008.