Contents What training does not include

Where the ability comes from

What training does not include

No major vendor discloses what its current model was trained on. That absence is the fact to start from, and it is also what makes the boundary on this page real. OpenAI withheld GPT-4's dataset construction, citing competition and safety1. Stanford's 2025 transparency index records Anthropic marking its dataset list, licensed sources, language composition and domain composition as proprietary and not disclosed2. OpenAI's own legal disclosure names only broad categories: publicly available data, data accessed through third-party partners, and information users, trainers and researchers provide or generate, plus synthetic data, with no datasets or proportions given3.

So nobody outside these companies can determine the exact training mix from what has been disclosed. What is documented, and far more useful, is what is outside it.

Outside the training text What the documentation points to instead
Your own files a File Search tool, which has to be added
Events after the cutoff a Web Search tool, which has to be added
Your own past chats with the app user prompts and outputs were excluded from training
The exact contents of the mix withheld as proprietary

OpenAI states that bridging the gap between a model's training cutoff and current events requires adding Web Search for the open internet, or File Search to retrieve from your own files and databases, which means the base model has neither by default4. Anthropic states that the Claude 3 models were "not trained on any user prompt or output data submitted to us by users", free, paid and API alike5. Your documents and your earlier conversations are not in there, and that is documented rather than inferred.

Recency is the same kind of boundary. OpenAI wrote that GPT-4 generally lacked knowledge of events after September 2021, the month most of its pretraining data stopped1. That is one 2023 model and not a current figure, and current models publish their own dates.

The boundary: this is not a claim that nothing about you could be in there. OpenAI acknowledges that despite steps to reduce it, some training data may include personal information swept up from public and partner sources3. That is incidental material from the open internet, not your account. And the recency edge is not one clean date, which the knowledge cutoff page takes up.

References

Quizzes
  1. Why does a chatbot not know what is in your own documents unless something is added?

    • They are outside the text it was trained on
    • It reads them only once you pay for a higher tier
    • It deletes them at the end of each conversation

    Vendors document that reaching your files takes an added file-retrieval tool; your account's documents are not part of what the trained model already has access to.

  2. One vendor's model pages give ____, rather than a single sharp line for what the model knows.

    • a reliable date and a later, broader one
    • one date shared by every model
    • a date each user picks at sign-up

    Anthropic publishes a reliable knowledge cutoff and a broader training data cutoff per model, and for some models the two differ by months.

  3. No major AI vendor publishes the exact dataset used to train its current flagship model.

    • True
    • False

    Both OpenAI and Anthropic mark their dataset lists, sources and composition as proprietary and undisclosed for their current models, citing competitive and safety reasons.

Comments

No comments yet. Start the conversation.