Generative AI (genAI) training requires massive amounts of text. The better-written and more information-dense that text is, the more it helps. AI gets a lot smarter a lot faster when it’s trained on well-written books and magazine and newspaper articles than when it’s trained on social media banter (or most everything else you find on the internet).
Since the dawn of AI, companies have been hoovering up copyrighted material wherever they find it — on the open web, behind paywalls, even in manually scanned books — and then used the scanned text. And they do it all without asking authors’ or publishers’ permissions, and without paying them.
It’s the greatest intellectual property theft in history by a long shot — billions and billions of dollars worth. Books, articles, music, photographs, artwork, you name it. If a human being has created it, Big AI has likely grabbed it and ingested it without asking, then made big profits off it.
Read the full article here

