Training DataTraining data is the collection of examples, records, documents, images, audio, code, labels, and other inputs used during training to update an AI model’s parameters. For large language models (LLMs), it can include web pages, books, code repositories, academic text, conversations, licensed archives, instruction-response pairs, and human preference ratings.
Learn more:
Training Data: How Data Shapes an AI Model