The OECD.AI Policy Navigator

Our policy navigator is a living repository from more than 80 jurisdictions and organisations. Use the filters to browse initiatives and find what you are looking for.

KB Lab


Added by:   OECD analyst
Added on:   20 Jul 2026
Updated by:   OECD analyst
Updated on:   20 Jul 2026

KB-labb (The KBLab) at the National Library of Sweden is a data-driven infrastructure for research. It provides access to the library’s massive digital collections—including newspapers, books, and radio/TV—for large-scale analysis.

Name in original language

KB-labb

Initiative overview

KBLab is a national infrastructure for data-driven research, established by the Royal Library (KB). The initiative aims to bridge the gap between large-scale digitized cultural heritage and traditional humanities methods. As collections reach petabyte levels – encompassing everything from historical newspapers to social media and audiovisual broadcasts – manual analysis becomes impossible. KBLab makes this "big data" accessible through computational tools and high-performance environments, transforming the library from a place for reading to an arena for active computing.

The main goal is to lower the thresholds for KB's digital vault. Through a "co-design" environment, the lab facilitates the development of large-scale AI models, especially language (LLM) and vision models trained on Swedish material. These models function as a "digital public good" that enables complex analyses, such as entity recognition and structural mapping in millions of documents. The aim is to create an ecosystem where computer science and humanities meet to generate new knowledge from cultural heritage. From being a pilot project, KBLab has become a central pillar of KB's digital strategy. The operations have shifted from pure data deliveries to providing an HPC cluster (High-Performance Computing) where researchers can apply their code directly to the data. The lab has contributed greatly to the open-source community through models such as "KB-BERT", which have become the industry standard for processing Swedish text.

KBLab has digitized and released massive datasets (newspapers, books, radio), leading to the creation of KB-BERT and KBLab’s Whisper, which are now industry standards for Swedish NLP. These models have been downloaded hundreds of thousands of times, enabling private and public sectors to automate text analysis and transcription with high precision. Quantitative metrics for impact include model download counts (Hugging Face), GPU-cluster utilization, and the volume of petabytes processed. Qualitative tracking involves monitoring peer-reviewed citations and the integration of KBLab. 

Other relevant details

KBLab’s primary challenge is the tension between open science and legal constraints (GDPR/copyright). We addressed this via a "lab-as-a-service" model—bringing researchers to the data rather than distributing restricted files. Another challenge is that developing AI models requires heavy GPU infrastructure and a cultural shift. We responded by forming interdisciplinary teams where data scientists and scholars co-create, ensuring technical outputs are research-ready and relevant. Ultimately, success required sustained funding, HPC access, and a shift in viewing the library as a "data provider" rather than just a "bookshelf." Going forward, the focus is on institutionalizing data-driven methods throughout the Swedish cultural heritage sector. This includes handling complex multimodal data, such as synchronized audio and text from broadcast archives. By integrating workflows into KB's core services and collaborating internationally, KBLab strives to be a role model for how national libraries can function in the AI ​​era through the concept of "Collections as Data".

About the policy initiative