AI Training on Public Data Is Mostly Legal in 2026, But Courts Are Still Deciding
AI models are largely built by scraping publicly available internet content — including blog posts, code, artwork, and social media — without explicit consent from creators. The central legal question is whether this constitutes copyright infringement or qualifies as fair use under existing law. In the United States, a court ruling in a case against Anthropic found AI training on books to be 'transformative' and thus leaning toward fair use, a significant win for the industry. However, legal experts note that 'legal' and 'something you agreed to' have quietly diverged, as most users never consented to their work being used this way. The law continues to evolve through ongoing litigation, and the final boundaries of what AI companies can and cannot do with public data remain unsettled.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in