The OECD.AI Policy Navigator

Our policy navigator is a living repository from more than 80 jurisdictions and organisations. Use the filters to browse initiatives and find what you are looking for.

Khmer Intelligent Document AI Agent (KIDAI-A)


Added by:   OECD analyst
Added on:   03 Aug 2026
Updated by:   OECD analyst
Updated on:   03 Aug 2026

KIDAI-A is Cambodia’s first sovereign AI agent purpose-built for the Khmer language. It processes Khmer and English text, speech, and image documents, enabling government agencies, businesses, and citizens to query, summarise, and extract information from official documents in their native language. KIDAI-A closes the critical gap left by global AI tools that perform poorly on Khmer, a low-resource language, advancing digital sovereignty and public service efficiency.

Initiative overview

Documents are the backbone of governance, commerce, and social life. Yet virtually all leading Document AI systems, designed for high-resource languages like English, perform poorly on Khmer, Cambodia’s official script used by over 17 million people. This creates a compounding digital divide: government officials cannot efficiently query Khmer-language policy archives; businesses struggle to automate Khmer invoice processing; and citizens face barriers accessing public information in their own language. No commercially available AI solution adequately addresses this gap.

Government Use Cases in Action:
1. Admin Assistant: Summarises official letters, retrieves staff records, and answers queries about internal reports in real time.
2. Policy Assistant: Enables officials to query national AI strategies (e.g., NAIS), MEF IT frameworks, and regulatory documents conversationally.
3. Data Entry Automation: Extracts structured data from invoices, sale reports, and forms, dramatically reducing manual processing time.

KIDAI-A locally hosts open-source LLMs (Qwen, Gemma, GPT-OSS variants) to ensure data sovereignty, no sensitive government data leaves Cambodia’s infrastructure. The platform is offered through Subscription (PaaS), API services (OCR, STT, MT, RAG), Knowledge Hub digitisation services, and vertical solutions. Target customers include government agencies, education and research institutes, the financial sector, enterprises, SMEs, and legal services. The team is pursuing co-development partnerships with government ministries and plans to expand KIDAI-A into a national-grade AI infrastructure for Cambodia’s digital government ecosystem. KIDAI-A represents more than a product, it is a national movement. Cambodia, like many lower-middle-income nations, faces a structural risk: dependence on foreign AI tools that neither understand Khmer culture, nor protect Cambodian data. KIDAI-A is Cambodia’s answer to this challenge, a fully sovereign AI stack designed, trained, and deployed by Cambodians, for Cambodians.

Results, outcomes and impacts: KIDAI-A is actively under development with the following milestones achieved and projected:
1. Proprietary Khmer speech recognition model trained on 500+ hours of audio data, achieving high transcription accuracy for Cambodian speech.
2. In-house OCR and document layout analysis models functional for Khmer/English document extraction.
3. Working prototype deployed locally — vector store loaded with 406+ government documents, demonstrating real-time multilingual Q&A.
4. Projected impact: reduce document processing time for government staff by 60-80%; enable digital access to Khmer-language policy archives; support Cambodia’s national digitisation agenda.

Other relevant details

Challenges and lessons learned: 1. Data scarcity: Khmer is a low-resource language with limited publicly available annotated datasets for OCR, ASR, and NLP training. The team addressed this by building proprietary datasets through manual annotation and semi-supervised learning. 2. Script complexity: Khmer’s abugida script with complex ligatures and stacked consonants poses unique challenges for layout analysis and OCR, requiring custom model architectures beyond standard approaches. 3. Compute constraints: Training large-scale Khmer models demands significant GPU resources. Lessons Learned & Success Conditions 1. Sovereign-first design is essential: Using locally-hosted LLMs and on-premise vector stores from the outset eliminates data sovereignty risks and builds government trust. 2. Domain expertise + AI: Deep Khmer linguistic expertise embedded within the AI team is irreplaceable. 3. Government partnership is the critical path to scale.

About the policy initiative


Category:

  • AI policy initiatives, programmes and projects

Initiative type:

  • AI use cases/projects in the public sector

Status:

  • Proposed or under development

Start Year:

  • 2025

Other relevant urls: