Every conversation we've had about AI and music over the past two years has had the same shape. An artist suspects their catalog got scraped. A label's lawyers allege "mass infringement" in a court filing. An AI company says, more or less, don't worry about it, the data was public, it's fair use, move along. And the rest of us are stuck in the middle, never quite able to confirm who's right because nobody outside the AI labs could actually see what was inside these training sets.
That changed a couple of weeks ago, and it's worth slowing down on.
The Atlantic rolled out an expanded version of its AI Watchdog project, built by staff writer Alex Reisner, that does something almost embarrassingly simple: it lets you type in a song or an artist's name and find out whether that track shows up in any of four datasets currently being passed around the AI development world. No subpoena required. No NDA. Just a search bar.
And what it turned up is a lot. Across those four datasets, more than 21 million recordings have been catalogued, linked, and packaged up in ways that make them easy for developers to pull into a training pipeline. That's not a typo, and it's not a worst-case estimate from an advocacy group with a reason to round up. It's just what's sitting in the open, on places like Hugging Face, where this stuff gets shared.
You can search AI Watchdog yourself over at The Atlantic. And if you want to see how Sleeping-DISCO-9M was actually built, Sleeping AI has their research notes up at sleeping-ai.com, at least for now.
Okay, so what's actually in these datasets?
Two of the four are relatively modest, just over 100,000 tracks each, including one built from the Free Music Archive, a collection that predates generative AI entirely. It was assembled back in 2017 for academic research into how software analyzes and sorts music, using tracks artists had released under Creative Commons licenses. Nobody involved in building it had AI music generators in mind, which tells you something about how repurposed a lot of this "training data" really is.
The other two are where it gets serious.
LAION-DISCO-12M is the big one, over 12 million tracks, put together by LAION, the German nonprofit that's also behind the dataset that trained Stability AI's image generator. It went up in November 2024, explicitly marked for research and academic use only, with LAION warning against using it commercially. Whether that warning has been respected is, charitably, an open question.
Then there's the one we're actually here to talk about: Sleeping-DISCO-9M. This is a collection of roughly 9.7 million tracks pulled from YouTube, each one paired with lyrics scraped off Genius.com. It was put together by Sleeping AI, a decentralized, mostly volunteer research collective that builds datasets and publishes how they did it. According to their own write-up, the whole point of Sleeping-DISCO was that earlier open datasets like this were basically just YouTube links with barely any metadata attached, not useful for training anything sophisticated. So they built something more complete: full lyric data, artist details, songs spanning 169 languages. It's a genuinely impressive piece of engineering, which is sort of the uncomfortable part of this story. The people building these things are often not cartoon villains. They're researchers solving an interesting technical problem, and the ethical weight of where the underlying material came from seems to land somewhere further down their priority list.
One thing worth flagging if you go looking for it yourself: Sleeping AI's website currently just says "Sleeping DISCO was taken down." No explanation given as of this writing. But taking a dataset off your own site doesn't claw back the copies already downloaded elsewhere, and there's no shortage of evidence that this one had already been pulled into commercial use well before anyone hit delete.
Combined, these four datasets sweep up everyone from Taylor Swift, Bad Bunny, and the Beatles to jazz catalogs, classical recordings, and a long tail of independent artists who never agreed to any of this. An audit by the Australasian rights group APRA AMCOS found work by Kylie Minogue, AC/DC, Flume, Tame Impala, and Lorde buried in there too. If you make music and you've ever uploaded it anywhere with a "public" setting, there's a real chance you're in one of these.
Why this lands differently than the last two years of headlines
We've covered the Suno and Udio lawsuits in this space before, and if you've been following along, you know the broad strokes: RIAA sued both companies in June 2024 on behalf of the majors, alleging the kind of wholesale copying that doesn't really have a polite name. Since then the story's split. UMG and Warner walked away from the courtroom and toward the negotiating table, settling and licensing their way into "walled garden" AI platforms with filtering built in. Sony hasn't budged and is still suing both companies. Smaller players, like the instrumental duo The American Dollar, have filed their own suits claiming Suno gutted their licensing income by close to 80%.
What all of that litigation has had in common, up until now, is that it's mostly been an argument about things nobody outside the courtroom could see. Sealed records. Disputed disclosures. A lot of "trust us" from companies whose entire business model depends on you not asking too many questions about where the training data came from.
AI Watchdog doesn't end that argument, but it does change its shape. A song turning up in one of these datasets isn't airtight proof that it trained a specific commercial model, Reisner is upfront about that. Companies filter, mix, and exclude material in ways nobody can fully audit from outside. And a song's absence from the tool doesn't mean it's clean either, since there are almost certainly private datasets we'll never see. But for the first time, there's something concrete to point at instead of a hunch.
And the financial stakes underneath all this are not small. A CISAC-commissioned study put a number on what generative AI could cost music creators by 2028: 24% of their revenue, somewhere around $10.5 billion cumulatively. Meanwhile Deezer reported in April that it's now receiving close to 75,000 fully AI-generated tracks a day, which works out to something like 44% of everything new hitting the platform. For context, that number was 10,000 a day back in January 2025, when Deezer first turned on its detection tool.
What you can actually do with this if you're an artist
Here's the practical part. The tool is free, it's public, and it lives directly on The Atlantic's site under the AI Watchdog banner. You search your name or a song title, and you either find yourself in there or you don't.
If you do, that's not nothing. It's documentation you didn't have a month ago, and documentation matters whether you're talking to an entertainment lawyer, negotiating deals with AI clauses baked in, or just trying to understand your own exposure without paying someone to do forensic research on your behalf. Independent artists in particular have spent the last two years stuck arguing from a position of "I assume this happened to me," which is a weak place to negotiate from. This moves a lot of people from assuming to knowing.
It's not a legal strategy on its own, and we'd never tell you to treat a search result as a substitute for actual counsel. But it's a starting point that didn't exist before, and starting points are usually the hardest part.
Where this leaves us
What sticks with me about this story isn't really the 21 million number, even though it's a genuinely wild figure to sit with. It's that the secrecy itself just took a hit. For two years, this entire debate has run on competing claims nobody could verify, companies insisting their data was clean, artists insisting it wasn't, and the rest of us picking a side based on vibes more than evidence.
That's harder to do now. Not impossible, the fight over what these datasets actually prove is just getting started, but harder. And in an industry that's spent a long time telling artists to simply trust that their work wasn't used without permission, having something you can check yourself is a different kind of reassurance than being told to take someone's word for it.