Initiative overview
The trial indicated benefits from adopting generative AI, as well as challenges and concerns requiring ongoing monitoring. Overall, Copilot use was moderate and focused on a few use cases. However, most trial participants across classifications and job families were optimistic about Copilot and wished to keep using it.
Only a third of trial participants across classifications and job families used Copilot daily. Copilot was mainly used to summarise information and re-write content. Copilot in Microsoft Word and Teams were viewed favourably and used most frequently. Access barriers prevented Copilot use in Outlook.
In terms of perceived improvements to efficiency and quality, participants estimated time savings of up to an hour per day when summarising information, preparing a first draft of a document and searching for information. The highest efficiency gains were perceived by staff in entry level to intermediate roles (APS levels 3-6), first-line supervisory roles (Executive Level (EL) 1) and ICT roles.
The majority of managers (64%) perceived uplifts in efficiency and quality in their teams. More than a third (40%) of trial participants reported being able to reallocate their time to higher-value activities such as staff engagement and strategic planning. The trial also identified potential for Copilot to improve inclusivity and accessibility in the workplace and in government communication.
The trial found adoption requires a concerted effort to address barriers, including key integration, data security and information management considerations agencies must consider before adopting Copilot. These include scalability and performance of the GPT integration and understanding of the context of the large language model.
Training in prompt engineering and use cases tailored to agency needs is required to build capability and confidence in Copilot. Clear communication and policies are required to address uncertainty regarding the security of Copilot, accountabilities and expectations. Adaptive planning is also needed to reflect the rolling feature release cycle of Copilot alongside governance structures that reflect agencies’ risk appetite and clear roles and responsibilities across government to provide advice on generative AI use.
Broader concerns on AI that require active monitoring include the potential public sector workforce impacts of generative AI, particularly on entry-level jobs and women. Large language model (LLM) outputs may be biased towards ‘Western’ norms and may not appropriately use cultural data and information, raising particular concerns for First Nations data sovereignty. Trial participants also raised broader concerns regarding vendor lock-in and competition as well as the environmental impacts of growing generative AI use.




























