Abstract

Voice conversion (VC) is an emerging technology in speech processing that aims to modify an utterance so that it sounds as though it were spoken by another speaker while preserving its linguistic content. High-quality voice conversion has broad applications, including speech synthesis, assistive communication, and entertainment applications such as multilingual dubbing. However, current embedding-guided voice conversion (EGVC) frameworks often struggle with generalization and naturalness under regional data-scarcity conditions. This study explores these limitations by evaluating an EGNVC framework adapted for low-resource regional-language pairs. The proposed framework incorporates the Harvest pitch-extraction algorithm alongside pretrained speaker representations to guide cross-gender pitch transitions while attempting to preserve speaker-identity profiles. Experimental results show that, although the framework successfully shifts macro-level pitch contours across genders, spectral alterations lead to substantial acoustic distortion and reduced intelligibility. Specifically, Kannada speech conversion achieved a localized objective intelligibility score of STOI = 0.12, whereas Malayalam transformations exhibited substantial spectral variation, with an MCD of 169.75, highlighting significant language-specific barriers to regional voice conversion. Kannada achieved higher intelligibility (STOI = 0.12) than Malayalam, whereas Malayalam required greater spectral modification (MCD = 169.75), indicating language-specific challenges in voice conversion. These baseline metrics delineate the empirical limitations of current embedding-guided architectures for Dravidian languages and indicate that substantial advances in spectral mapping are required before such systems can be integrated into real-time assistive or localized voice-synthesis applications.

Keywords

Voice Conversion, Speaker Embeddings, Neural Voice Transformation, Indian Regional Languages, Generative Adversarial Networks (GANs), Pitch Extraction,

Downloads

Download data is not yet available.

References

  1. T. Walczyna, Z. Piotrowski, Overview of voice conversion methods based on deep learning. Applied sciences, 13(5), (2023) 3100. https://doi.org/10.3390/app13053100
  2. B. Li, J. Chen, Y. Xu, W. Li, Z. Liu, DRAW: Dual-Decoder-Based Robust Audio Watermarking Against Desynchronization and Replay Attacks. IEEE Transactions on Information Forensics and Security, 19, (2024) 6529 – 6544. https://doi.org/10.1109/TIFS.2024.3416047
  3. A. Triantafyllopoulos, B.W. Schuller, G. İymen, M. Sezgin, X. He, Z. Yang, P. Tzirakis, S. Liu, S. Mertes, E André, An overview of affective speech synthesis and conversion in the deep learning era. Proceedings of the IEEE, 111(10), (2023) 1355-1381. https://doi.org/10.1109/JPROC.2023.3250266
  4. Z. Yang, X. Jing, A. Triantafyllopoulos, M. Song, I. Aslan, B.W. Schuller, (2022) An overview & analysis of sequence-to-sequence emotional voice conversion. arXiv preprint arXiv:2203. 15873. https://doi.org/10.48550/arXiv.2203.15873
  5. W.C. Huan, L.P. Violeta, S. Liu, J. Shi, T. Toda, (2023) The singing voice conversion challenge 2023. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, Taipei, Taiwan. https://doi.org/10.1109/ASRU57964.2023.10389671
  6. M. I. Hussan, D. Saidulu, P.T. Anitha, A. Manikandan, P. Naresh, Object detection and recognition in real time using deep learning for visually impaired people. International Journal of Electrical and Electronics Research, 10(2), (2022) 80-86. https://doi.org/10.37391/ijeer.100205
  7. A. Mehrish, N. Majumder, R. Bharadwaj, R. Mihalcea, S. Poria, A review of deep learning techniques for speech processing. Information Fusion, 99, (2023) 101869. https://doi.org/10.1016/j.inffus.2023.101869
  8. J.H. Belz, L.M. Weilke, A. Winter, P. Hallgarten, E. Rukzio, T. Grosse-Puppendahl, Story-Driven: Exploring the Impact of Providing Real-time Context Information on Automated Storytelling. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 73, (2024) 1-15. https://doi.org/10.1145/3654777.3676372
  9. S. Aggarwal, S. Uttam, S. Garg, S. Garg, K. Jain, S. Aggarwal, Advancements in end-to-end audio style transformation: a differentiable approach for voice conversion and musical style transfer. AI, 6(1), (2025) 16. https://doi.org/10.3390/ai6010016
  10. C.W. Bang, C. Chun, Effective zero-shot multi-speaker text-to-speech technique using information perturbation and a speaker encoder. Sensors, 23(23), (2023) 9591. https://doi.org/10.3390/s23239591
  11. W. Slam, Y. Li, N. Urouvas, Frontier research on low-resource speech recognition technology. Sensors, 23(22), (2023) 9096. https://doi.org/10.3390/s23229096
  12. X. Xu, L. Shi, X. Chen, P. Lin, J. Lian, J. Chen, Z. Zhang, E.R. Hancock, Any-to-any voice conversion with multi-layer speaker adaptation and content supervision. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31, (2023) 3431-3445. https://doi.org/10.1109/TASLP.2023.3306716
  13. A. Angra, H. Muralikrishna, D.A. Dinesh, V. Thenkanidiyoor, Exploring aggregated wav2vec 2.0 features and dual-stream TDNN for efficient spoken dialect identification. IEEE Access, 13, (2024) 3115-3129. https://doi.org/10.1109/ACCESS.2024.3523951
  14. K. Ezzine, J.D. Martino, M. Frikha, Any-to-one non-parallel voice conversion system using an autoregressive conversion model and lpcnet vocoder. Applied Sciences, 13(21), (2023) 11988. https://doi.org/10.3390/app132111988
  15. W.S. Hsu, G.T. Lin, W.H. Wang, Enhancing dysarthric voice conversion with fuzzy expectation maximization in diffusion models for phoneme prediction. Diagnostics, 14(23), (2024) 2693. https://doi.org/10.3390/diagnostics14232693
  16. D. Yook, G. Han, H.P. Chang, I.C. Yoo, Cyclediffusion: Voice conversion using cycle-consistent diffusion models. Applied Sciences, 14(20), (2024) 9595. https://doi.org/10.3390/app14209595
  17. K. Zhou, B. Sisman, R. Rana, B.W. Schuller, H. Li, Emotion intensity and its control for emotional voice conversion. IEEE Transactions on Affective Computing, 14(1), (2022) 31-48. https://doi.org/10.1109/TAFFC.2022.3175578
  18. D.S. Asudani, N.K. Nagwani, P. Singh, Impact of word embedding models on text analytics in deep learning environment: a review. Artificial intelligence review, 56(9). (2023): 10345-10425. https://doi.org/10.1007/s10462-023-10419-1
  19. M. Matassoni, S. Fong, A. Brutti, Speaker anonymization: Disentangling speaker features from pre-trained speech embeddings for voice conversion. Applied Sciences, 14(9), (2024) 3876. https://doi.org/10.3390/app14093876
  20. M. Amarjouf, E.H. Ibn Elhaj, M. Chami, K. Ezzine, J. Di Martino, Assessment of Self-Supervised Denoising Methods for Esophageal Speech Enhancement. Applied Sciences, 14(15), (2024) 6682. https://doi.org/10.3390/app14156682
  21. M. Zhang, Y. Zhou, Y. Ren, C. Zhang, X. Yin, H. Li, Refxvc: Cross-lingual voice conversion with enhanced reference leveraging. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32, (2024) 4146-4156. https://doi.org/10.1109/TASLP.2024.3439996
  22. K. Guo, Y. Xu, S. Zhang. IAFF-VC: Any-to-Any Voice Conversion Using Attentional Feature Fusion. Circuits, Systems, and Signal Processing, 44 (2025) 8489-8509. https://doi.org/10.1007/s00034-025-03198-3
  23. R. Yamamoto, R. Yoneyama, L.P. Violeta, W.C. Huang, T. Toda, (2023) A comparative study of voice conversion models with large-scale speech and singing data: The T13 systems for the singing voice conversion challenge 2023. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), IEEE, Taipei, Taiwan. https://doi.org/10.1109/ASRU57964.2023.10389779
  24. B. Sisman, J. Yamagishi, S. King, H. Li, An overview of voice conversion and its challenges: From statistical modeling to deep learning. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29, (2020) 132-157. https://doi.org/10.1109/TASLP.2020.3038524
  25. S. Chen, Y. Gu, J. Cui, J. Zhang, R. Chen, L. Dai, Lcm-svc: Latent diffusion model based singing voice conversion with inference acceleration via latent consistency distillation. In 2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP), IEEE, Beijing, China. https://doi.org/10.1109/ISCSLP63861.2024.10800358
  26. W. Lu, X. Zhao, N. Guo, Y. Li, J. Wei, J. Tao, J. Dang, One-shot emotional voice conversion based on feature separation. Speech Communication, 143, (2022) 1-9. https://doi.org/10.1016/j.specom.2022.07.001
  27. H. Kameoka, T. Kaneko, K. Tanaka, N. Hojo, S. Seki, (2020). Voicegrad: Non-parallel any-to-many voice conversion with annealed langevin dynamics. Arxiv Preprint Arxiv: 2010.02977. https://doi.org/10.48550/arXiv.2010.02977
  28. H. Guo, C. Liu, C.T. Ishi, H. Ishiguro, (2023) Using joint training speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) IEEE, Taipei, Taiwan. https://doi.org/10.1109/ASRU57964.2023.10389651
  29. K. Qian, Y. Zhang, S. Chang, X. Yang, M. Hasegawa-Johnson, AutoVC: Zero-shot voice style transfer with only autoencoder loss. In Proceedings of the 36th International Conference on Machine Learning, PMLR, 97, (2019) 5210–5219.
  30. Y.A. Li, A. Zare, N, Mesgarani, (2021) StarGANv2-VC: A diverse, unsupervised, non-parallel framework for natural-sounding voice conversion. Arxiv Preprint Arxiv: 2107.10394. https://doi.org/10.48550/arXiv.2107.10394
  31. A. Das, S. Ghosh, T. Polzehl, I. Siegert, S. Stober, (2023). StarGAN-VC++: Towards emotion preserving voice conversion using deep embeddings. Arxiv Preprint Arxiv: 2309.07592. https://doi.org/10.48550/arXiv.2309.07592
  32. RVC-Project, (2023). Retrieval-based-Voice-Conversion-WebUI. GitHub repository. https://github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI