Further Utilization of Classical Japanese Book Information through Character Image Processing Technology
2026.09.23
Fiscal Year
FY 2025
April 2025 – March 2026
Principal Investigator
Shinichiro Omachi
Professor, Graduate School of Engineering
Co-Investigators
Yoshiyasu Sakura
Professor, Graduate School of Arts and Letters
Research Keywords
Classical Japanese books ; Kuzushiji (cursive Japanese characters)
1. Research Overview
Image data of classical Japanese books (kotenseki) consist of handwritten characters, so transcribing them into printed text requires enormous effort. Using conventional methods, converting the vast collection of roughly 300,000 classical books into searchable text data is effectively impossible in terms of time, labor, and funding. This study aims to establish a highly accurate search technology for this vast body of classical book images and, at the same time, to rethink research methods in the humanities, thereby fostering new creativity in humanities research.
Within this project, highly accurate search technology for classical books is the key technology that determines the success of the research. Many classical books are written in kuzushiji, a cursive style in which the dots and strokes of characters are abbreviated or connected, making them difficult to decipher without specialized training. Although the accuracy of machine-learning-based optical character recognition (OCR) has improved, conventional technologies merely convert images into text and cannot explain why a character is read in a particular way. This study therefore focused on the latest AI capable of understanding images and text simultaneously—multimodal large language models. By additionally training such a model specifically for kuzushiji, we aimed both to substantially improve recognition accuracy and to clarify the basis of its judgments. In our experiments, the model achieved high accuracy in single-character recognition and substantially reduced the error rate in recognizing entire vertical lines of text. We also found that a model architecture that preserves the aspect ratio of images is particularly important for line-level recognition.
In parallel, as an effort to enhance creativity in the humanities, we investigated and examined the correlation between the characteristics and significance of classical works themselves and classical book information—centered on the distinction among writing systems, namely kanji, hiragana, and katakana. In studying mainly war tales (gunki monogatari), which include diverse textual variants that differ greatly in their orthography, a new and important issue emerged: in some works, differences in orthography play a decisive role in the differences in expressive content between variant texts, whereas in others their role is largely limited.
2. Significance of the Research and Future Prospects
By making it possible to recognize kuzushiji using multimodal large language models, the results of this study are considered an important step toward understanding the content of classical books and automatically generating modern Japanese translations of them.
3. Conclusion
Japan holds image data of approximately 300,000 classical books, most of which are written in kuzushiji, and converting them into text requires vast amounts of time, labor, and cost. In this study, we sought to improve kuzushiji recognition accuracy by additionally training the latest AI capable of understanding images and text simultaneously (a multimodal large language model) specifically for kuzushiji. In experiments, the model achieved high accuracy in single-character recognition and substantially improved the error rate for entire lines. These results contribute to establishing highly accurate search technology for classical books and open new possibilities for humanities research.