The OECD.AI Policy Navigator

Our policy navigator is a living repository from more than 80 jurisdictions and organisations. Use the filters to browse initiatives and find what you are looking for.

Microsoft 365 Copilot Trial


Added by:   OECD analyst
Added on:   14 Aug 2026
Updated by:   OECD analyst
Updated on:   14 Aug 2026

In 2024, the DTA led the 6-month Australian Government trial of Microsoft 365 Copilot – a generative AI assistant embedded in the Microsoft 365 application suite. The trial involved the distribution of over 5,700 Copilot licenses across 56 federal government agencies. Broadly, the trial and evaluation tested whether the anticipated benefits of generative AI capabilities translated into real‑world adoption by workers.

Initiative overview

The trial indicated benefits from adopting generative AI, as well as challenges and concerns requiring ongoing monitoring. Overall, Copilot use was moderate and focused on a few use cases. However, most trial participants across classifications and job families were optimistic about Copilot and wished to keep using it. 

Only a third of trial participants across classifications and job families used Copilot daily. Copilot was mainly used to summarise information and re-write content. Copilot in Microsoft Word and Teams were viewed favourably and used most frequently. Access barriers prevented Copilot use in Outlook. 

In terms of perceived improvements to efficiency and quality, participants estimated time savings of up to an hour per day when summarising information, preparing a first draft of a document and searching for information. The highest efficiency gains were perceived by staff in entry level to intermediate roles (APS levels 3-6), first-line supervisory roles (Executive Level (EL) 1) and ICT roles. 

The majority of managers (64%) perceived uplifts in efficiency and quality in their teams. More than a third (40%) of trial participants reported being able to reallocate their time to higher-value activities such as staff engagement and strategic planning. The trial also identified potential for Copilot to improve inclusivity and accessibility in the workplace and in government communication. 

The trial found adoption requires a concerted effort to address barriers, including key integration, data security and information management considerations agencies must consider before adopting Copilot. These include scalability and performance of the GPT integration and understanding of the context of the large language model. 

Training in prompt engineering and use cases tailored to agency needs is required to build capability and confidence in Copilot. Clear communication and policies are required to address uncertainty regarding the security of Copilot, accountabilities and expectations. Adaptive planning is also needed to reflect the rolling feature release cycle of Copilot alongside governance structures that reflect agencies’ risk appetite and clear roles and responsibilities across government to provide advice on generative AI use. 

Broader concerns on AI that require active monitoring include the potential public sector workforce impacts of generative AI, particularly on entry-level jobs and women. Large language model (LLM) outputs may be biased towards ‘Western’ norms and may not appropriately use cultural data and information, raising particular concerns for First Nations data sovereignty. Trial participants also raised broader concerns regarding vendor lock-in and competition as well as the environmental impacts of growing generative AI use.

Other relevant details

Results, outcomes and impacts: Over 2,000 trial participants from over 50 agencies contributed to the evaluation. The evaluation adopted a mixed-methods approach, involving surveys (pre-use, mid-trial, post-use), consultations (interviews, focus groups) and document/data review (research papers, trial issues register, internal agency evaluations). Key survey results: Majority of respondents agreed Copilot improved task completion speed (69%) and work quality (61%). Majority of managers surveyed (65%) noticed quality and efficiency improvements in their team. Some participants identified novel use cases which were highly specific to their roles, such as coding for task automation and converting technical documentation into plain language for non-technical audiences. Challenges and lessons learned: The positive sentiment towards Copilot was not uniformly observed across all MS products or activities. In particular, MS Excel and Outlook Copilot functionality did not meet expectations. The quality of Copilot’s output limited the scale of productivity benefits. Overall, Copilot’s improvements to work quality were more subdued than improvements to work efficiency. While the majority of trial participants considered Copilot effective at developing first drafts of documents and lifting overall quality, editing was almost always needed to tailor content for the audience or context, reducing total efficiency gains. The evaluation identified several barriers to generative AI adoption requiring concerted effort to overcome. These included technical integration challenges with non-Microsoft 365 applications, staff capability, uncertainty around recordkeeping and other public administration obligations and negative stigmas and ethical concerns associated with generative AI.

About the policy initiative


Category:

  • AI policy initiatives, programmes and projects

Initiative type:

  • AI use cases/projects in the public sector

Status:

  • Inactive – initiative complete

Start Year:

  • 2023

End Year:

  • 2024

OECD AI Principles: