The world’s biggest AI companies are buying antiquarian books en masse so they can be scanned to train their large language models (LLMs) before being destroyed.
Aren’t some of the books considered rare ?
I mean that’s why people are mad.
And yeah I agree they already use stolen works.
I’m just saying I heard someone say that.
I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.
But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception).
And I was really surprised they had a photocopy of it.
But maybe Google wanted a hefty price for access. Maybe a subscription.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.
Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
And yeah I agree they already use stolen works.
I’m just saying I heard someone say that.
I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.
But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception). And I was really surprised they had a photocopy of it.
But maybe Google wanted a hefty price for access. Maybe a subscription.
I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.
Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.