The OECD.AI Policy Navigator

Our policy navigator is a living repository from more than 80 jurisdictions and organisations. Use the filters to browse initiatives and find what you are looking for.

DigiTeks - Application of AI for understanding the Serbian language and digitalization in the field of legislation


Added by:   OECD analyst
Added on:   30 Sep 2026
Updated by:   OECD analyst
Updated on:   30 Sep 2026

The initiative is a flexible module for automatic text correction, designed to improve efficiency, quality, and reliability in writing and communication. While targeted at “JP Službeni Glasnik”, it potentially benefits institutions, professionals, and anyone working with Serbian-language content. Developed to address the need for accurate OCR and text analysis, it opens opportunities for integration with existing tools and broader language applications.

Name in original language

DigiTeks - Primena veštačke inteligencije na razumevanje srpskog jezika i digitalizaciju u oblasti zakonodavstva

Initiative overview

The initiative addresses persistent challenges in digitising and correcting Serbian legal and administrative documents by integrating two complementary modules. The first module performs information extraction using Google Tesseract and Apache Tika, adapted to process both scanned and text-based documents. Its output is structured JSON enriched with confidence scores, allowing more dependable downstream analysis.

The second module focuses on correcting misread or inconsistent text by applying the Jerteh 355 language model, which is fine-tuned on legislative materials and enhanced through whole-word masking. By combining OCR confidence metrics, contextual probabilities, and character-level similarity, it generates accurate HTML reconstructions suitable for publication or further processing.

The objective is to deliver a modern and adaptable system that raises efficiency, quality, and reliability in text digitisation. In addition to improving document accuracy, the system supports advanced linguistic and analytical tasks such as lemmatisation, indexing, and semantic search. This enables institutions, professionals, and researchers to work with Serbian legal and administrative texts more effectively and to access higher-quality digital resources.

Looking ahead, the initiative is designed for evolution and wider applicability. Future development includes integration with existing text-processing tools, straightforward replacement of models as new versions emerge, and expansion to additional institutions and document types. The underlying architecture emphasises adaptability, allowing the system to gradually develop into a broader platform for Serbian-language technologies. It can support complementary functions such as entity recognition, terminology extraction, and style correction.

Ultimately, the initiative establishes a foundation that can contribute to long-term digital transformation in legislative and administrative environments.

Results, outcomes and impacts

The implementation of the solution has been successfully completed. The executable files, as well as the full code of the final version of the software (key event 3), have been published on the public GitHub repository (procesaur/digiteks), so that the software is available to the public for download and use.

The developed software is also available on the Internet and as a web application that can be tested. The initial version of the new language model, which was trained on documents obtained from a public entity (JP Službeni Glasnik), as well as on additional prepared resources (key event 2), has been published for testing on the Hugging Face platform.

Other relevant details

Challenges and lessons learned: The initiative faced initial challenges with OCR accuracy for Serbian Cyrillic, poor scan quality, and complex legal formatting. Multi-GPU execution on Windows also caused instability due to unsupported IPC protocols. These were addressed by enhancing preprocessing with Poppler, fine-tuning the Jerteh 355 model on legislative texts, and disabling IPC to stabilise execution. Lessons learned include the importance of domain adaptation, preprocessing quality, and flexible model design. Success requires reliable infrastructure, high-quality training data, institutional support, and iterative refinement.

About the policy initiative


Category:

  • AI policy initiatives, programmes and projects

Initiative type:

  • AI use cases/projects in the public sector

Status:

  • Inactive – initiative complete

Start Year:

  • 2024

End Year:

  • 2025

Target Sectors:


OECD AI Principles:

—

Other relevant urls:

—