38.3 Million AI Training Data Sets from Magazines, Textbooks, and Academic Papers Made Public

by Yoon Juhye Posted : September 17, 2026, 09:12Updated : September 17, 2026, 09:12

More than 38 million data sets, including books, tables, illustrations, advertisements, and photographs, will be made available to the public.


The National Library of Korea announced on September 17 that it will publicly release artificial intelligence (AI) training data built from national literature for the first time.


The release includes 3,974 text data sets (Hidden Text PDF, XML, TXT, JSON), 21,485 image data sets such as tables, illustrations, advertisements, and photographs, as well as 38.3 million character datasets, totaling 38.33 million items.


Along with this data release, the existing project platform, 'shared library,' has been revamped to optimize it for AI training data usage. This upgrade focuses on enhancing convenience for various purposes, including AI training, academic research, service and app development, and content creation, making it easily accessible to everyone.


The open data consists of digitized original texts from the National Library of Korea, including modern magazines from the 1930s and 1940s, government publications and textbooks from the 1940s to 1960s, and open access academic papers that have been authorized for AI training use. Anyone can freely browse and download the materials from the 'shared library.'


The text data is provided in a refined, high-quality format that is easy for AI to recognize, allowing for immediate use in research and service development. The character datasets can be directly utilized for AI training, which is expected to benefit small and venture companies struggling to secure large-scale data for AI-based optical character recognition (OCR) and natural language processing.


The revamped 'shared library' has also strengthened its features to support AI and digital humanities research and expand the use of public data. By organizing and processing the vast knowledge resources accumulated digitally into a format that AI can learn from, large-scale data-driven humanities research, which was previously challenging, is now possible. Researchers are expected to conduct digital humanities studies more easily by analyzing various materials such as literature, records, and images with AI to derive new meanings.


Moreover, the significance of this initiative lies in the large-scale opening of national literature that can be directly utilized for AI training, alleviating some of the burdens of costs and copyright issues that researchers and small and venture companies have faced in securing training data.


The platform's user environment has also been improved to be more user-friendly. It features an intuitive layout and a navigation structure based on material and image types, allowing anyone to easily access the data without a separate registration process, making it easy for both professional researchers and the general public to search and utilize the materials.


The National Library of Korea plans to continue expanding data construction and opening, including the additional release of text data from over 300 Korean novels by the end of this year, and to enhance AI utilization services, developing the 'shared library' into a continuously growing data platform.


Lee Hyun-joo, head of the Digital Information Planning Division at the National Library of Korea, stated, “We will steadily expand the scope of data openness to create a library in the AI era where all citizens can benefit from public knowledge resources.”





* This article has been translated by AI.