PROJECT

THEME

2025年度:文字画像処理技術による古典籍情報の有効化

代表者
大町 真一郎 (工学研究科)
分担者
  • 佐倉 由泰 (文学研究科)

INTRODUCTION

AI-Generated Comment on the Study



古典籍の多くは「くずし字」で書かれており、その翻刻には高度な専門知識と膨大な労力が必要とされます。約30万点に及ぶ古典籍の画像データを検索可能なテキストへと変換することは、人文学研究における長年の課題となってきました。

本研究では、画像とテキストを同時に理解できるマルチモーダル大規模言語モデルをくずし字に特化して追加学習させることで、一文字認識および行単位の認識の双方において認識精度を大きく向上させています。また、画像の縦横比を保持するモデル構造の重要性を明らかにした点や、軍記物語を対象として表記の違いと表現内容との関係を考察した点は、情報科学と人文学を架橋する本研究ならではの成果といえます。

本研究で得られた知見は、古典籍の高精度な検索や内容理解、さらには現代語訳の自動生成へとつながる基盤となるものです。今後、日本の文化史やリテラシー史を新たな視点から捉え直す研究への発展が期待されます。




Prompt Used for AI-Generated Comment
Please read the following research summary and write a short comment on the study.
The comment should be suitable for inclusion in an academic research report. It should evaluate the significance, originality, interdisciplinary value, and future potential of the study. Please keep the tone formal, balanced, and positive, but avoid exaggerated claims. The comment should be understandable to a general academic audience and should not be too technical.
Preferred length: 100–150 words.
(Claude)



Research Details



Link: https://www.aisogochi.tohoku.ac.jp/archives/1200