Somewhere in Tokyo right now, there’s a warehouse worker weighing books like they’re scrap metal. Not counting them. Weighing them. Fifty tons, shipped off to the US, destined to get scanned page by page and then shredded into confetti. This is real. This is happening. And used bookstores across Japan are seeing sales jump 5x because of it.
Wait, People Are Buying Books By The Ton?
Yeah. By the ton. Not by the box, not by the crate – by the ton, like you’re buying gravel for a driveway. According to reports circulating out of Japan (this started making rounds on Reddit and honestly I had to read it twice to believe it), used bookstores that normally sold, I don’t know, a modest steady stream of secondhand paperbacks are suddenly getting cleaned out by bulk buyers. Whole shelves. Whole storerooms. Gone.

And here’s the part that makes this whole thing feel like something out of a weird near-future novel: a lot of these books aren’t going to readers. They’re going to AI companies. Or at least, that’s the strong suspicion – one 50-ton shipment reportedly got sent straight to the US where the working theory is it’ll be scanned for AI training data and then destroyed. Shredded. Turned into pulp or landfill or whatever happens to a book once a machine has already sucked the words out of it.
Let that sit for a second. The physical book becomes disposable the moment its content gets digitized. That’s… a lot to process, honestly.
The Suspicious Bulk Buy Pattern
What’s got people talking isn’t just the one shipment – it’s the pattern. Multiple large, suspicious purchases, all pointing toward the same destination: foreign facilities set up specifically to scan and shred. Not resell. Not archive. Scan and shred. That phrase alone tells you everything about what these books are worth to the buyers – which is to say, not the books themselves. The data inside them.
But Why Japan Specifically?
This is the question I keep coming back to. Why Japan? And honestly, the answer probably comes down to a mix of things that make total sense once you think about it for more than five seconds.
Japan has an enormous, incredibly well-preserved secondhand book market. Bookstores like Book Off have turned used book selling into something close to a science – organized, clean, insanely well-stocked. Decades of print culture, decades of people actually buying and keeping physical books, means there’s just… a lot of supply sitting around. High quality supply too. Not moldy basement boxes. Actual clean, readable, well-organized text.

And AI companies need text. Mountains of it. Clean, structured, ideally non-web-scraped text that isn’t already contaminated with SEO garbage or bot-generated nonsense (which, let’s be honest, is a growing percentage of the internet at this point). Books – real, professionally edited, human-written books – are basically gold for training data. High signal, low noise. You can’t say that about most of Reddit, no offense to Reddit.
“It’s not entirely clear who’s behind every purchase, but the pattern is too consistent to ignore – bulk orders, cash deals, and destinations that lead straight to data processing facilities.”
The Bookstores Are Thrilled. Should They Be?
Here’s the thing that’s kind of funny in a dark way – the used bookstores are having a moment. A 5x sales surge is massive by any measure. If you’re a shop owner who’s watched foot traffic slowly die for years while people migrated to e-readers and Amazon, this probably feels like a miracle. Free money falling from the sky. Who’s going to say no to that?
I get it. I really do. If someone showed up at my door offering to buy my entire garage sale inventory by weight, I wouldn’t ask a ton of questions either (pun very much intended). But I can’t shake the feeling that this is a short-term win wrapped around a long-term loss nobody’s really talking about yet.
Because once those books get shredded, they’re gone. Not “out of print” gone – actually, physically, permanently gone. And if enough of them disappear this way, at scale, across enough regions, you start losing something that’s genuinely hard to get back: physical cultural record. Old editions. Marginalia. Print runs that never got digitized properly the first time around. Sure, the AI companies are theoretically preserving the text – but text isn’t the same as the object. A scanned PDF of a 1987 paperback is not the same experience, culturally or historically, as the actual paperback sitting on a shelf in Shimokitazawa.
Nobody Asked The Books’ Original Owners
And there’s a copyright question buried in here that I don’t think enough people are asking loudly enough. If these books are getting scanned for AI training, are the authors getting paid? Are publishers being asked? From what I can tell – and I want to be careful here because the reporting on this is still pretty thin – a lot of this seems to be happening in a legal gray zone. Buy the physical book, you own the physical book. Does that give you the right to digitize its entire contents for commercial AI training? That’s… genuinely murky, and I don’t think courts have fully sorted it out yet, in Japan or anywhere else.
This Isn’t Really New, It’s Just Louder Now
Look, AI companies scraping content for training data isn’t some shocking revelation at this point. We’ve all watched this story play out with news articles, artwork, code repositories, you name it. The pattern is basically the same every time: find a huge pool of high-quality human-created content, ingest it, argue about whether that was legal after the fact. Books being physically bought and shredded is just… a more dramatic, more visceral version of the same move. It’s the meatspace equivalent of a web scraper running in the background you never see.
What makes this one hit different, though, is the physicality of it. Deleting a scraped webpage doesn’t feel like anything. Shredding an actual book – something a person once held, dog-eared, maybe wrote their name inside the front cover – that feels like something. Even if you’re someone who thinks AI training on books should absolutely be allowed (and there are reasonable arguments for that), there’s something a little unsettling about the literal grinding up of physical media once its “useful” data has been extracted. It’s efficient. It’s also kind of grim.
What This Actually Means
I think we’re watching the early, weird phase of something that’s going to look completely normal in five years. Bulk book buying for AI training data is probably going to become its own quiet little industry – brokers, scanning facilities, maybe even standardized pricing per ton depending on genre or condition. It sounds absurd typing it out, but so did a lot of things about the internet in 2005.
My honest take? The used bookstores cashing in right now should enjoy the windfall, but somebody – governments, publishers, authors’ guilds, somebody – needs to figure out the legal framework around this before it becomes an even bigger habit. Because right now it feels like nobody’s really in charge of deciding whether this is fine. It’s just happening. Fifty tons at a time.
And the next time you donate a box of old paperbacks to a thrift store thinking they’ll end up in someone’s hands, on someone’s shelf, being read again… maybe they will. Or maybe they’re getting weighed on a scale somewhere, on their way to becoming training data before getting shredded into nothing. We honestly don’t know anymore. That’s the part that stays with me.