This is the Trace Id: 3b353670d1367de3cc4957abc0beb2d7

Data for AI model development

Explore our approach to data for training generative AI models.
Two people sitting at a table reviewing content on a tablet, with a coffee cup nearby.
Trust and transparency are core to how we approach data for training AI models here at Microsoft.

The purpose of this page is to share our approach, outlining the categories of data we use for training and what we are doing to give web publishers and content creators control over how their content is used to train these models.

What does it mean to train an AI model?

AI models make predictions based on data. During the training process, an AI model identifies patterns, correlations, and concepts across vast amounts of data and then uses what it has learned to generate answers to questions or requests. The model can be tuned using different categories of training data to optimize generation capabilities for specific classes of information such as natural language, software code, or images. It can also be grounded with an organization’s own data—for example, in an industry vertical such as agriculture, health, or retail.

Generative AI models are not designed to store, access, or reproduce the original training data. Instead, these models are intended to generate new expressive works and content.

Two people standing indoors, reviewing a dashboard on a large monitor with charts and metrics displayed.

Discover how Microsoft trains generative AI models on varied data

Large and varied datasets are integral to creating today’s AI systems. It’s critical that a broad diversity of data is used to develop models that are more accurate, more representative of different cultures and languages, and that learn and apply the full spectrum of human experience and ingenuity.
As identified in our model cards and other publicly available documentation, Microsoft sources the following categories of data for training our generative AI models responsibly:
  • We train on select publicly available data, from a mix of industry-standard machine learning datasets and web crawls. We exclude sources with paywalls and content that violates our policies, and that have opted out of training using web controls that we published. Web-crawled data is safety-filtered and processed to remove illegal content, including child sexual abuse material (CSAM), in accordance with applicable laws and Microsoft’s Responsible AI policies.

    We respect standards such as robots.txt tags that web publishers use to opt out of web crawling. We don’t use data from domains listed in the Office of the United States Trade Representative (USTR) Notorious Markets for Counterfeiting and Piracy list.
  • This is training data that we access through negotiated arrangements such as with publishers and copyright owners.
  • This data is collected from select Microsoft consumer services and protected in accordance with our Microsoft Privacy Statement. Our Microsoft Copilot blog and the Copilot FAQ explain how we use consumer data from Copilot, Bing, and Microsoft Start (MSN) to help train our models in Copilot. We make it easy for users to opt out and put in place safeguards such as removing identifiers that may identify you to protect user privacy. We do not use our enterprise customers’ data without their permission.
  • We sometimes train on synthetic datasets. These we create by prompting large language models (LLMs) to generate specifically requested output that augments scarce or limited real world data. The results are then reviewed and filtered, including by human reviewers, to ensure they are of sufficient quality to be part of the training dataset.
  • We include human feedback from AI trainers, red teamers, employees, and external annotation partners, in our training process. This includes, for example, human feedback that reinforces quality output to a user’s prompt, improving the end-user experience. Red teamers also test the model against harmful or unsafe requests, and their feedback is used to improve the model’s safety behaviors.
Along with our own Microsoft models, we incorporate various other models, including from OpenAI, into our products and services. Learn more about how OpenAI trains models.

For beneficial and inclusive AI, some web publishers and content owners want more say in how their content is used. Web publisher controls are helpful but not the complete solution. We are working alongside others in industry, publishing, standard bodies, and the content community to identify a common approach we all can advance and adopt.
Person standing indoors using a laptop in a bright office setting.

Responsible AI at Microsoft

Explore the tools, practice, and policies we’ve created to uphold our responsible AI principles.

More resources

Three people reviewing content together on a laptop in a living room setting.

Responsible AI Transparency Report

Learn how we build AI responsibly, support our customers, evolve, and grow.
Two people seated at a table using laptops during a discussion.

Transparency and Control in Consumer Data Use

Control how Copilot learns from your data.
Person holding a tablet displaying bar charts and analytics.

Advancing open data

Realize the benefits of more open and accessible data.

Follow Microsoft

English (United States) Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads