OpenAI’s models ‘memorized’ copyrighted content, new study suggests
image via TechCrunch
April 4, 2025, 6:42 PM
- •A new study suggests that OpenAI’s AI models may have memorized copyrighted content.
- •The study was co-authored by researchers at the University of Washington, the University of Copenhagen, and Stanford.
- •The study’s method relies on words that the co-authors call “high-surprisal” — that is, words that stand out as uncommon in the context of a larger body of work.
A new study suggests that OpenAI's AI models may have been trained on copyrighted content. The study's authors propose a new method for identifying training data that has been memorized by models, and they found that GPT-4 showed signs of having memorized portions of popular fiction books and New York Times articles.
Entities Mentioned
OpenAIUniversity of WashingtonUniversity of CopenhagenStanford
Topics Covered
AIcopyrightOpenAIstudy
Comments (0)
No comments yet.