Initiative overview
The Virtual Assistant is an attempt to help the citizen navigate the multitude of available public services ranging from reporting an environmental violations, to registering the birth of a child. Every year many services are introduced, changed, or retired, making it difficult for the citizen to keep track of them. The ability to ask questions in natural language makes the services more approachable, and by extension helps every citizen access everything the modern state has to offer. As a secondary objective, the initiative strives to build trusted public AI – both in the terms of the information provided by the Virtual Assistant, and by legally and technically enshrining the safety of the citizen data.
The virtual assistant differs from the commercial offerings by the strict validation and verification of the knowledge sources and close integration with government databases. The answers it provides are limited to the official documentation – it is also designed not to give opinions or judgement-based advice. This ensures the assistant can be trusted to help citizens, while at the same time saving the budgetary cost of citizens asking general-knowledge questions. Furthermore, the assistant is embedded in mObywatel - the official Polish state application that lets people keep digital versions of important documents and use many public services safely on their mobiles. Therefore, if the user asks about a service or a document available in the app, the assistant can link directly to that service, ensuring seamless user experience.
Finally, the Virtual Assistant is designed from the ground up to be privacy-preserving. It does not link the conversations to a user preventing any profiling. Every message goes through a multi-step process designed to purge any personally identifiable information. The LLM infrastructure is fully self-managed, to prevent external API providers from accessing the conversation data.
The Virtual Assistant has been used 600 000 times in the two months from its introduction. The performance of the assistant is measured in real time, based on the user feedback (surveys after a conversation and the ability to up and downvote specific answers). Specific answers receive 62% positive votes and 38% negative, while the whole conversations have an equal split of 44% good or very good and 44% of bad or very bad. The specific rated conversations and messages are then assessed both manually and automatically by analytical and data science teams to guide the development of the system. Moreover, the project is assessed using social science frameworks, i.e: moderated workshops and unmoderated online forms.



























