Synthetic Voices in Education: Exploring the Impact of GenAI-based Text-to-Speech Lectures on Student Engagement and Comprehension

2026.09.25

Fiscal Year
FY 2025
April 2025 – March 2026

Principal Investigator
Masako Hayashi
Associate Professor, Institute for Excellence in Higher Education

Co-Investigators
Marvin Eder
Graduate School of Information Sciences

Research Keywords
Generative AI ; Synthetic speech ; Multilingual education ; Flipped classroom ; XR/metaverse-based international collaborative learning

1. Research Overview

This study examined the educational potential of generative AI-based speech synthesis for creating multilingual flipped-classroom materials more efficiently in international collaborative learning using XR and the metaverse.

In international collaborative courses, the flipped-classroom format—in which students watch lecture videos before class and focus on collaborative work, discussion, and presentations during class—is effective. However, producing lecture videos in multiple languages, such as Japanese and English, places a heavy burden on instructors. In addition, while standard text-to-speech (TTS) voices are efficient, it is necessary to examine how learners perceive them in terms of naturalness and human-likeness.

In this study, we therefore built a video production pipeline that automatically generates narrated lecture videos in Japanese and English from PowerPoint slides and lecture scripts. We also created a synthetic voice trained on the instructor’s own voice and compared natural speech, the trained synthetic voice, and standard TTS voices.

Thirty-two students watched lecture videos with identical content that differed only in voice, and rated them in terms of naturalness, human-likeness, ease of understanding, and willingness to continue listening. The results showed that standard TTS voices received relatively low ratings, whereas the trained synthetic voice received ratings close to those of natural speech. This suggests that synthetic voices are promising not as a simple replacement for natural speech, but as an assistive technology that supports the creation of multilingual teaching materials.

This study is foundational research that connects generative AI, speech information processing, multilingual education, and XR/metaverse-based international collaborative learning, contributing to more efficient material development and improved accessibility in international education.

2. Significance of the Research and Future Prospects

The significance of this study lies in demonstrating the potential to reduce instructors’ burden in producing flipped-classroom videos in multiple languages and to provide more accessible materials for the diverse students participating in international collaborative learning.

In international collaborative learning using XR and the metaverse, it is important to devote class time not to one-way explanations but to dialogue, collaborative creation, presentations, and reflection among students. This requires a mechanism for efficiently preparing lecture videos for pre-class study. The video generation pipeline developed in this study can automatically produce Japanese and English lecture videos from PowerPoint slides and lecture scripts, making it applicable to future support for multilingual education.

At the same time, synthetic voices trained on an instructor’s own voice require careful ethical consideration, including the individual’s consent, scope of use, storage and management, restrictions on third-party use, and clear disclosure that the voice is AI-generated. Going forward, rather than only expanding research on synthetic speech itself, we will examine the role of voice interfaces as one component of a broader AI-based educational infrastructure supporting international collaborative learning, such as AI facilitators and AI orchestrators.

In the future, we aim to integrate generative AI, XR and the metaverse, learning analytics, and multilingual support to develop a safe and practical educational support system for collaborative learning in international education.

3. Conclusion

This study examined the educational use of generative AI-based speech synthesis to support the creation of multilingual flipped-classroom materials for international collaborative learning using XR and the metaverse. We built a system that automatically generates Japanese and English lecture videos from PowerPoint slides and lecture scripts, and compared natural speech, a trained synthetic voice, and standard TTS voices. The trained synthetic voice received ratings close to those of natural speech, indicating its potential to support the creation of multilingual teaching materials.