Name in original language
DigiTeks - Primena veštačke inteligencije na razumevanje srpskog jezika i digitalizaciju u oblasti zakonodavstva
Initiative overview
The initiative addresses persistent challenges in digitising and correcting Serbian legal and administrative documents by integrating two complementary modules. The first module performs information extraction using Google Tesseract and Apache Tika, adapted to process both scanned and text-based documents. Its output is structured JSON enriched with confidence scores, allowing more dependable downstream analysis.
The second module focuses on correcting misread or inconsistent text by applying the Jerteh 355 language model, which is fine-tuned on legislative materials and enhanced through whole-word masking. By combining OCR confidence metrics, contextual probabilities, and character-level similarity, it generates accurate HTML reconstructions suitable for publication or further processing.
The objective is to deliver a modern and adaptable system that raises efficiency, quality, and reliability in text digitisation. In addition to improving document accuracy, the system supports advanced linguistic and analytical tasks such as lemmatisation, indexing, and semantic search. This enables institutions, professionals, and researchers to work with Serbian legal and administrative texts more effectively and to access higher-quality digital resources.
Looking ahead, the initiative is designed for evolution and wider applicability. Future development includes integration with existing text-processing tools, straightforward replacement of models as new versions emerge, and expansion to additional institutions and document types. The underlying architecture emphasises adaptability, allowing the system to gradually develop into a broader platform for Serbian-language technologies. It can support complementary functions such as entity recognition, terminology extraction, and style correction.
Ultimately, the initiative establishes a foundation that can contribute to long-term digital transformation in legislative and administrative environments.
Results, outcomes and impacts
The implementation of the solution has been successfully completed. The executable files, as well as the full code of the final version of the software (key event 3), have been published on the public GitHub repository (procesaur/digiteks), so that the software is available to the public for download and use.
The developed software is also available on the Internet and as a web application that can be tested. The initial version of the new language model, which was trained on documents obtained from a public entity (JP Službeni Glasnik), as well as on additional prepared resources (key event 2), has been published for testing on the Hugging Face platform.





























