Background information
Large language models (LLMs) can generate a broad range of content, including text, code, images, audio, and other media. They are being applied across various contexts, ranging from cognitive tasks to manufacturing.
Artificial Analysis performs intelligence, quality, performance and price benchmarking on AI models, inference API endpoints and systems. See their methodology page for more details about terminology and benchmarking approaches.
Coverage
The database includes publicly available LLMs and does not account for strictly on-premise or non-disclosed options. While Artificial Analysis does not include metrics for all AI models offered by cloud providers, they attempt to cover all new models that are relevant to AI developers and practitioners.
Update Frequency
Data is updated on a continual basis at Artificial Analysis, and will be updated on a monthly basis on OECD.AI.
Dimensions
The “developer” is the firm that developed the model. A model’s country assignment is determined by the location of the headquarters of the developer.
The dataset also includes additional information when available, in particular whether the model has been released in open weight with a permissive license (allowing commercial use).
Models are further categorised based on the modality of the input and output. Modalities can be text, audio, image, or video. These are grouped into text or multi-modal, where the latter represents an input or output of at least two modalities. The vast majority of multi-modal models are those with text and an additional input or output.
Quality
Quality is measured using the Artificial Analysis Intelligence Index: A composite benchmark aggregating nine challenging evaluations to provide a holistic measure of AI capabilities across mathematics, science, coding, and reasoning. See the examples and frequently asked questions for additional details.
Quality is displayed in two forms:
- The average intelligence index score of all models, by model release date.
- The highest intelligence index score achieved by a model developed in each country, by date. Previously released models are carried forward, so the score at any given date represents the best-performing model released by a developer in that country up to that date.
A frontier model, therefore, is the model in the specified geographic area that performs the best according to the Intelligence Index.


























