throbber

`
`
`UNITED STATES DISTRICT COURT
`NORTHERN DISTRICT OF CALIFORNIA
`
`RICHARD KADREY, et al.,
`Plaintiffs,
`v.
`
`META PLATFORMS, INC.,
`Defendant.
`
`Case No. 23-cv-03417-VC
`
`
`ORDER DENYING THE PLAINTIFFS’
`MOTION FOR PARTIAL SUMMARY
`JUDGMENT AND GRANTING
`META’S CROSS-MOTION FOR
`PARTIAL SUMMARY JUDGMENT
`Re: Dkt. Nos. 482, 501
`
`
`Companies are presently racing to develop generative artificial intelligence models—
`software products that are capable of generating text, images, videos, or sound based on
`materials they’ve previously been “trained” on. Because the performance of a generative AI
`model depends on the amount and quality of data it absorbs as part of its training, companies
`have been unable to resist the temptation to feed copyright-protected materials into their
`models—without getting permission from the copyright holders or paying them for the right to
`use their works for this purpose. This case presents the question whether such conduct is illegal.
`Although the devil is in the details, in most cases the answer will likely be yes. What
`copyright law cares about, above all else, is preserving the incentive for human beings to create
`artistic and scientific works. Therefore, it is generally illegal to copy protected works without
`permission. And the doctrine of “fair use,” which provides a defense to certain claims of
`copyright infringement, typically doesn’t apply to copying that will significantly diminish the
`ability of copyright holders to make money from their works (thus significantly diminishing the
`incentive to create in the future). Generative AI has the potential to flood the market with endless
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 1 of 40
`
`
`
`
`
`
`
`
`2
`amounts of images, songs, articles, books, and more. People can prompt generative AI models to
`produce these outputs using a tiny fraction of the time and creativity that would otherwise be
`required. So by training generative AI models with copyrighted works, companies are creating
`something that often will dramatically undermine the market for those works, and thus
`dramatically undermine the incentive for human beings to create things the old-fashioned way.
`Take, for example, biographies. If a company uses copyrighted biographies to train a
`model, and if the model is thus capable of generating endless amounts of biographies, the market
`for many of the copied biographies could be severely harmed. Perhaps not the market for Robert
`Caro’s Master of the Senate, because that book is at the top of so many people’s lists of
`biographies to read. But you can bet that the market for lesser-known biographies of Lyndon B.
`Johnson will be affected. And this, in turn, will diminish the incentive to write biographies in the
`future.
`Or take magazine articles. If a company uses copyrighted magazine articles to train a
`model capable of generating similar articles, it’s easy to imagine the market for the copied
`articles diminishing substantially. Especially if the AI-generated articles are made available for
`free. And again, how will this affect the incentive for human beings to put in the effort necessary
`to produce high-quality magazine articles?
`With some types of works, the picture is a bit murkier. For example, it’s not clear how
`generative AI would affect the market for memoirs or autobiographies, since by definition people
`read those works because of who wrote them. With fiction, it might depend on the type of book.
`Perhaps classic works of literature like The Catcher in the Rye would not see their markets
`diminished. But the market for the typical human-created romance or spy novel could be
`diminished substantially by the proliferation of similar AI-created works. And again, the
`proliferation of such works would presumably diminish the incentive for human beings to write
`romance or spy novels in the first place.
`Some students of copyright law respond that none of this matters because when
`companies use copyrighted works to train generative AI models, they are using the works in a
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 2 of 40
`
`
`
`
`
`
`
`
`3
`way that’s highly creative in its own right. In the language of copyright law, the companies’ use
`of the works is “transformative.” As a factual matter, there’s no disputing that. And as a legal
`matter, it’s true that you’re less likely to be liable for copyright infringement if you’re copying
`the work for a transformative purpose. In that situation, you’re more likely to be protected by the
`fair use doctrine. But as the Supreme Court has emphasized, the fair use inquiry is highly fact
`dependent, and there are few bright-line rules. There is certainly no rule that when your use of a
`protected work is “transformative,” this automatically inoculates you from a claim of copyright
`infringement. And here, copying the protected works, however transformative, involves the
`creation of a product with the ability to severely harm the market for the works being copied, and
`thus severely undermine the incentive for human beings to create. Under the fair use doctrine,
`harm to the market for the copyrighted work is more important than the purpose for which the
`copies are made.
`Speaking of which, in a recent ruling on this topic, Judge Alsup focused heavily on the
`transformative nature of generative AI while brushing aside concerns about the harm it can
`inflict on the market for the works it gets trained on. Such harm would be no different, he
`reasoned, than the harm caused by using the works for “training schoolchildren to write well,”
`which could “result in an explosion of competing works.” Order on Fair Use at 28, Bartz v.
`Anthropic PBC, No. 24-cv-5417 (N.D. Cal. June 23, 2025), Dkt. No. 231. According to Judge
`Alsup, this “is not the kind of competitive or creative displacement that concerns the Copyright
`Act.” Id. But when it comes to market effects, using books to teach children to write is not
`remotely like using books to create a product that a single individual could employ to generate
`countless competing works with a miniscule fraction of the time and creativity it would
`otherwise take. This inapt analogy is not a basis for blowing off the most important factor in the
`fair use analysis.
`Another argument offered in support of the companies is more rhetorical than legal:
`Don’t rule against them, or you’ll stop the development of this groundbreaking technology. The
`technology is certainly groundbreaking. But the suggestion that adverse copyright rulings would
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 3 of 40
`
`
`
`
`
`
`
`
`4
`stop this technology in its tracks is ridiculous. These products are expected to generate billions,
`even trillions, of dollars for the companies that are developing them. If using copyrighted works
`to train the models is as necessary as the companies say, they will figure out a way to
`compensate copyright holders for it.
`The upshot is that in many circumstances it will be illegal to copy copyright-protected
`works to train generative AI models without permission. Which means that the companies, to
`avoid liability for copyright infringement, will generally need to pay copyright holders for the
`right to use their materials.
`But that brings us to this particular case. The above discussion is based in significant part
`on this Court’s general understanding of generative AI models and their capabilities. Courts can’t
`decide cases based on general understandings. They must decide cases based on the evidence
`presented by the parties.
`In this case, thirteen authors—mostly famous fiction writers—have sued Meta for
`downloading their books from online “shadow libraries” and using the books to train Meta’s
`generative AI models (specifically, its large language models, called Llama). The parties have
`filed cross-motions for partial summary judgment, with the plaintiffs arguing that Meta’s
`conduct cannot possibly be fair use, and with Meta responding that its conduct must be
`considered fair use as a matter of law. In connection with these fair use arguments, the plaintiffs
`offer two primary theories for how the markets for their works are affected by Meta’s copying.
`They contend that Llama is capable of reproducing small snippets of text from their books. And
`they contend that Meta, by using their works for training without permission, has diminished the
`authors’ ability to license their works for the purpose of training large language models. As
`explained below, both of these arguments are clear losers. Llama is not capable of generating
`enough text from the plaintiffs’ books to matter, and the plaintiffs are not entitled to the market
`for licensing their works as AI training data. As for the potentially winning argument—that Meta
`has copied their works to create a product that will likely flood the market with similar works,
`causing market dilution—the plaintiffs barely give this issue lip service, and they present no
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 4 of 40
`
`
`
`
`
`
`
`
`5
`evidence about how the current or expected outputs from Meta’s models would dilute the market
`for their own works.
`Given the state of the record, the Court has no choice but to grant summary judgment to
`Meta on the plaintiffs’ claim that the company violated copyright law by training its models with
`their books. But in the grand scheme of things, the consequences of this ruling are limited. This
`is not a class action, so the ruling only affects the rights of these thirteen authors—not the
`countless others whose works Meta used to train its models. And, as should now be clear, this
`ruling does not stand for the proposition that Meta’s use of copyrighted materials to train its
`language models is lawful. It stands only for the proposition that these plaintiffs made the wrong
`arguments and failed to develop a record in support of the right one.
`I. COPYRIGHT LAW AND FAIR USE
`The goal of copyright law is to promote “broad public availability of literature, music,
`and the other arts.” Twentieth Century Music Corp. v. Aiken, 422 U.S. 151, 156 (1975). To this
`end, copyright law incentivizes creativity by giving authors of original works a bundle of
`exclusive rights—for instance, the rights to prevent others from reproducing or distributing the
`works. 17 U.S.C. § 106. At the same time, however, copyright law “trades off the benefits of
`incentives to create against the costs of restrictions on copying.” Andy Warhol Foundation for
`the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508, 526 (2023). For example, copyright only
`protects expression, not underlying ideas, and the duration of copyright protection is limited. See
`id. (citing 17 U.S.C. §§ 102, 302–305).
`One major way the Copyright Act strikes a balance between protecting ownership and
`leaving room for innovation is through the affirmative defense of fair use. Under this doctrine,
`“the fair use of a copyrighted work . . . for purposes such as criticism, comment, news reporting,
`teaching . . . , scholarship, or research, is not an infringement of copyright.” 17 U.S.C. § 107.
`Fair use “permits courts to avoid rigid application of the copyright statute when, on occasion, it
`would stifle the very creativity which that law is designed to foster.” Google LLC v. Oracle
`America, Inc., 593 U.S. 1, 18 (2021) (quoting Stewart v. Abend, 495 U.S. 207, 236 (1990)).
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 5 of 40
`
`
`
`
`
`
`
`
`6
`The Copyright Act lists four factors to be considered in determining whether a given use
`is fair:
`1. the purpose and character of the use, including whether such use
`is of a commercial nature or is for nonprofit educational purposes;
`
`2. the nature of the copyrighted work;
`
`3. the amount and substantiality of the portion used in relation to the
`copyrighted work as a whole; and
`
`4. the effect of the use upon the potential market for or value of the
`copyrighted work.
`17 U.S.C. § 107.
` While the statute lists these four factors, fair use is a “flexible concept.” Warhol, 598 U.S.
`at 527 (quotation marks omitted) (quoting Oracle, 593 U.S. at 20). The list is not exhaustive. A
`particular factor “may prove more important in some contexts than in others.” Oracle, 593 U.S.
`at 19. And application of the factors “requires judicial balancing, depending upon relevant
`circumstances, including ‘significant changes in technology.’” Id. (quoting Sony Corp. of
`America v. Universal City Studios, Inc., 464 U.S. 417, 430 (1984)). The factors may also overlap
`such that facts relevant to one factor are also relevant to others. See A.V. ex rel. Vanderhye v.
`iParadigms, LLC, 562 F.3d 630, 642 (4th Cir. 2009). Overall, the factors are not meant to be
`applied mechanically, but to contribute “to a holistic inquiry”: whether the secondary work is
`likely to substitute for the original work in the marketplace and therefore undermine the
`incentive to create. See Romanova v. Amilus Inc., 138 F.4th 104, 117 n.9 (2d Cir. 2025) (Leval,
`J.); see also Warhol, 598 U.S. at 528 (referring to substitution as “copyright’s bête noire”).
`Because it “focuses on actual or potential market substitution,” Warhol, 598 U.S. at 536
`n.12, the fourth factor is “undoubtedly the single most important element of fair use,” Harper &
`Row Publishers, Inc. v. Nation Enterprises, 471 U.S. 539, 566 (1985). If the law allowed people
`to copy your creations in a way that would diminish the market for your works, this would
`diminish your incentive to create more in the future. Thus, the key question in virtually any case
`where a defendant has copied someone’s original work without permission is whether allowing
`people to engage in that sort of conduct would substantially diminish the market for the original
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 6 of 40
`
`
`
`
`
`
`
`
`7
`work. See Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569, 590 (1994).
` Because fair use is an affirmative defense, the burden of proof is on the party invoking it.
`Dr. Seuss Enterprises, L.P. v. ComicMix LLC, 983 F.3d 443, 459 (9th Cir. 2020), abrogated on
`other grounds by Jack Daniel’s Properties, Inc. v. VIP Products LLC, 599 U.S. 140 (2023). In
`particular, because the fourth factor is the most important, the secondary user (generally, the
`defendant) will “have difficulty carrying the burden of demonstrating fair use without favorable
`evidence about relevant markets.” Campbell, 510 U.S. at 590. But while the rightsholder need
`not prove or present evidence of market harm, they “may bear some initial burden of identifying
`relevant markets.” Hachette Book Group, Inc. v. Internet Archive, 115 F.4th 163, 194 (2d Cir.
`2024); see also Newegg Inc. v. Ezra Sutton, P.A., No. CV 15-01395, 2016 WL 6747629, at *2
`(C.D. Cal. Sep. 13, 2016). Moreover, because fair use is a holistic inquiry, the party invoking it
`“bears the burden on the defense as a whole,” not as to each individual factor. William F. Patry,
`Patry on Fair Use § 2:5 (May 2025 ed.).
` Fair use is a mixed question of law and fact, but the “question primarily involves legal
`work.” Oracle, 593 U.S. at 24. Therefore, fair use can be addressed at summary judgment where
`there are no genuine issues of material fact relevant to fair use. Leadsinger, Inc. v. BMG Music
`Publishing, 512 F.3d 522, 530 (9th Cir. 2008). By contrast, where there are genuine factual
`disputes that might affect whether the defendant’s use was fair, those disputes must be resolved
`by a jury. See Oracle, 593 U.S. at 23–25. Once a jury finds the facts, whether those facts support
`fair use “is a legal question for judges to decide.” Id. at 23–24.
` It bears emphasis that where a fair use defense fails, the consequence isn’t necessarily
`that the defendant must stop whatever they were doing. The consequence will often be that the
`defendant needs to pay the copyright owner for a license that grants them permission to do
`whatever they were doing. This way, the defendant compensates the copyright owner for the fact
`that the defendant’s conduct will otherwise harm the market for the original work. The defendant
`will only be forced to stop what they’re doing if they’re unwilling or unable to pay for the right
`to do it.
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 7 of 40
`
`
`
`
`
`
`
`
`8
`II. FACTS AND PROCEDURAL HISTORY
`A
`“Generative AI” is a type of artificial intelligence that creates new content, such as text,
`images, videos, or sound.1 Generative AI models do this, as Meta describes it, by extracting
`“increasingly complex mathematical patterns from training data, enabling the network to output
`a prediction or decision based on the patterns derived.” To put it more simply, generative AI
`models are “trained” to identify common patterns across large training datasets. They can then
`create, in response to user prompts, new content based on the patterns they have recognized in
`that training data. By the same token, a model’s outputs are limited based on the patterns that
`existed in its training data. For instance, if the only bridge in an image-generating model’s
`training data was the Golden Gate Bridge, and a user told that model to generate an image of a
`bridge, it would likely generate an orange-red suspension bridge, because that is the pattern of a
`bridge that would emerge from its training data.
`A large language model, or LLM, is a particular type of generative AI model designed to
`understand and generate text. Users can prompt LLMs to do a wide range of things, such as draft
`emails, summarize documents, or write computer code. Well-known LLMs include OpenAI’s
`ChatGPT models and Google’s Gemini models.
`LLMs learn to understand language by analyzing relationships among words and
`punctuation marks in their training data. The units of text—words and punctuation marks—on
`which LLMs are trained are often referred to as “tokens.” LLMs are trained on an immense
`amount of text and thereby learn an immense amount about the statistical relationships among
`words. Based on what they learned from their training data, LLMs can create new text by
`predicting what words are most likely to come next in sequences. This allows them to generate
`text responses to basically any user prompt. Model developers may also “post-train” or
`“finetune” their models to improve their performance at specific tasks or otherwise adjust their
`outputs, such as to prevent generation of offensive statements. Therefore, as with other
`
`1 Except as noted, the parties do not dispute the facts described in this section.
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 8 of 40
`
`
`
`
`
`
`
`
`9
`generative AI models, LLMs’ outputs are limited by their training data. To be able to generate a
`wide range of text—in different languages or styles, or regarding different subject matter—an
`LLM’s training dataset must be large and diverse. As one Meta witness put it, “If a model only
`saw social media posts, for example, it would not do well in generating source computer code.”
`But while a variety of text is necessary for training, books make for especially valuable
`training data. This is because they provide very high-quality data for training an LLM’s
`“memory” and allowing it to work with larger amounts of text at once. (The technical term for
`how many tokens an LLM can hold in its memory at once is its “context window.”) For instance,
`an LLM with a better memory will be able to process and respond to longer prompts, incorporate
`more information into outputs, and remember things from earlier in an exchange, resulting in
`smoother “conversations.” Books are good data for training LLMs’ memories because, in the
`words of one of Meta’s expert witnesses, they are “long but consistent,” maintaining a particular
`style and coherent structure. They are also high quality in the sense that they generally are well
`written and use proper grammar (especially compared to text from the internet, which varies
`widely on these metrics).
`B
`Meta Platforms owns and operates social media services including Facebook, Instagram,
`and WhatsApp. It is also the developer of a series of LLMs named “Llama.” Meta released
`Llama 1 in February 2023 and Llama 2 that July. Llama 3—along with Meta AI, an easily
`accessible AI chatbot (analogous to ChatGPT) that incorporates Llama 3—was released in April
`2024. Llama 4 is planned for release later in 2025. As Meta explains, each new Llama edition
`improved in certain ways over its predecessor: Llama 2 was finetuned “to improve the safety,
`quality, and consistency” of its outputs; Llama 3 made “significant improvements in
`performance and efficiency”; and Llama 4 is generally “larger” and “more advanced.” Subject to
`certain restrictions, members of the public can download all of the Llama models for free for
`noncommercial use; Llama 2 and 3 are also free to download for commercial use. While the
`Llama models are free to download, Meta estimates that its total revenue from generative AI will
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 9 of 40
`
`
`
`
`
`
`
`
`10
`range from $2 to $3 billion in 2025, and from $460 billion to $1.4 trillion over the next ten years.
`See Pls. MSJ Ex. 8 at 12.
`To get the varied and extensive text necessary to train its models, Meta cast a wide net.
`Approximately two-thirds of the data used to train Llama 1 and 2 came from Common Crawl, a
`nonprofit organization that collects and provides free access to website data, metadata, and text.
`The remainder came from websites and databases including Wikipedia, GitHub, ArXiv, Stack
`Exchange, and a combination of Project Gutenberg and Books3 (two book databases).2 With the
`exception of Books3, none of the sources contained any copyrighted material at issue in this
`case.
`Although Meta needed (and acquired and used) a wide range of training data, it
`especially needed books because, as discussed above, books make for high-quality data. Meta AI
`researchers and engineers repeatedly discussed the benefits of using books as training data, as
`well as the need to acquire more books for this use. One Meta employee said that the “best
`resources we can think of are definitely books.” Pls. MSJ Ex. 18 at 2. Another said it was “really
`important for us to get books data ASAP.” Id. Ex. 40 at 2. So as Meta expanded its datasets
`generally, it also continued to look for more books in particular.
`At first, Meta wanted to license books and so tried to negotiate licensing deals with
`several major publishers. Meta’s head of generative AI discussed spending up to $100 million on
`licensing. But as negotiations proceeded, Meta realized that licensing would be more difficult
`than anticipated. For one thing, publishers generally do not hold the subsidiary rights to license
`books for AI training. These rights are instead held by individual authors, and there is no
`organization for collective licensing of such rights. Sinkinson Decl. ISO Meta MSJ ¶¶ 58–59, 62.
`Even where publishers do hold AI training licensing rights, they do so regionally rather than
`globally. Meta MSJ Ex. 34 at 22:22–25:15. For another thing, some publishers apparently
`
`2 According to Meta, GitHub is “a leading cloud-based platform where coders store and share
`code, frequently on an open-source basis.” ArXiv is “a free online archive of math, science, and
`economics papers.” Stack Exchange is “a network of question-and-answer websites for sharing
`technical knowledge, geared toward the programming community.”
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 10 of 40
`
`
`
`
`
`
`
`
`11
`ignored Meta’s outreach, and only one gave Meta a pricing proposal. Id. at 23:11–14, 24:2–10.
`Eventually, Meta began investigating the possibility of procuring the books (and other
`text) needed for training by downloading them from “shadow libraries.” A shadow library is an
`online repository that provides things like books, academic journal articles, music, or films for
`free download, regardless of whether that media is copyrighted. Meta first used a shadow library
`in October 2022, when it downloaded the Library Genesis (“LibGen”) database to investigate
`whether there was value in training Llama on the works it contained. Pls. MSJ Ex. 32 at 3. If the
`answer was yes, the plan was to then set up licensing agreements for those or similar works. Id.
`But in spring 2023, after failing to acquire licenses and following escalation to CEO Mark
`Zuckerberg, Meta decided to just use the works acquired from LibGen as training data. Id. Ex. 61
`at 5. And after confirming that LibGen contained most of the works available for license from
`certain publishers with which it had been negotiating, Meta abandoned its licensing efforts. Id.
`Ex. 50 at 131:1–132:10, 383:5–384:12; id. Ex. 57 at 2; id. Ex. 58 at 3; see also id. Ex. 92 at 12.
`In early 2024, Meta also downloaded Anna’s Archive, a compilation of shadow libraries
`including LibGen, Z-Library, and others. See id. Ex. 66 at 2–3.
`To download these large datasets more quickly and without unnecessarily slowing down
`its networks, Meta torrented them. “Torrenting” is a filesharing technique that entails the
`simultaneous distribution of small portions of a larger file from many different sources. To be
`more precise, those sources are many other computer systems that also contain that file. So, for
`instance, one who torrented LibGen would download small pieces of each book LibGen contains
`from other users who had copies of LibGen on their computer and who were participating in the
`torrenting network. The torrenting software would then take those pieces and reassemble them
`into the original files on the downloader’s computer.3
`Certain torrenting protocols—including the one used by Meta, called BitTorrent—are, by
`
`3 See generally David Gerwitz, What Is Torrenting?, ZDNET (Aug. 6, 2024),
`https://www.zdnet.com/article/what-is-torrenting-and-how-does-it-work [https://perma.cc/8PG5-
`H7UW].
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 11 of 40
`
`
`
`
`
`
`
`
`12
`default, configured so that files downloaded via torrenting may also be reuploaded to other
`computer systems. This reuploading can occur both while files are still being downloaded (which
`the parties refer to as “leeching”) and after those files have been fully downloaded (which the
`parties refer to as “seeding”). Some torrenting protocols—including BitTorrent—are designed to
`prioritize downloads to users who are also uploading.
`There is no dispute that Meta torrented LibGen and Anna’s Archive, but the parties
`dispute whether and to what extent Meta uploaded (via leeching or seeding) the data it torrented.
`A Meta engineer involved in the torrenting wrote a script to prevent seeding, but apparently not
`leeching. See Pls. MSJ at 13; id. Ex. 71 ¶¶ 16–17, 19; id. Ex. 67 at 3, 6–7, 13–16, 24–26; see also
`Meta MSJ Ex. 38 at 4–5. Therefore, say the plaintiffs, because BitTorrent’s default settings allow
`for leeching, and because Meta did nothing to change those default settings, Meta must have
`reuploaded “at least some” of the data Meta downloaded via torrent. The plaintiffs assert further
`that Meta chose not to take any steps to prevent leeching because that would have slowed its
`download speeds. Meta responds that, even if it reuploaded some of what it downloaded, that
`doesn’t mean it reuploaded any of the plaintiffs’ books. It also notes that leeching was not clearly
`an issue in the case until recently, and so it has not yet had a chance to fully develop evidence to
`address the plaintiffs’ assertions.
`Either way, Meta added the books it downloaded to the datasets it used to train the Llama
`models. It also post-trained its models to prevent them from “memorizing” and outputting certain
`text from their training data, including copyrighted material. These training efforts, which Meta
`calls “mitigations,” appear to have been successful. Meta’s expert witness tested them using a
`method designed to get LLMs to regurgitate material from its training data (which Meta calls
`“adversarial prompting”). Even using that method, the expert could get no model to generate
`more than 50 words and punctuation marks (that is, “tokens”) from the plaintiffs’ books. And the
`plaintiffs’ expert could only get the Llama model best at regurgitation to generate 50 words and
`punctuation marks from the plaintiffs’ books in 60% of tests. She also testified that Llama was
`not able to reproduce “any significant percentage” of them. Meta MSJ Ex. 24 at 237:16–19; see
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 12 of 40
`
`
`
`
`
`
`
`
`13
`also Pls. Ex. 79 ¶¶ 70–72, 79, 82–83, 92; Meta MSJ Ex. 23 at 179:22–25, 180:17–181:16. In
`short, Llama cannot currently be used to read or otherwise meaningfully access the plaintiffs’
`books.
`C
`The plaintiffs are thirteen published authors who have written, and who hold copyright
`in, various works. Those works are mostly novels, but also include plays, short stories, memoirs,
`essays, and nonfiction books. Examples include Sarah Silverman’s The Bedwetter, a comic
`memoir; Rachel Louise Snyder’s No Visible Bruises: What We Don’t Know About Domestic
`Violence Can Kill Us, a nonfiction book about domestic violence and how to combat it; Junot
`Díaz’s Pulitzer Prize–winning novel, The Brief Wondrous Life of Oscar Wao; and Andrew Sean
`Greer’s Less, also a Pulitzer Prize–winning novel. All of the books in which the plaintiffs hold
`copyright can be found in the datasets Meta downloaded, including both Books3 and the Anna’s
`Archive databases. In total, Meta downloaded at least 666 copies of books whose copyrights the
`plaintiffs hold.
`Each plaintiff says that they would be open to licensing their books for use as generative
`AI training data, but that Meta did not approach them about this licensing. No plaintiff has
`licensed a book to any company for use as LLM training data or been asked by any company to
`license a book for that purpose.
`The plaintiffs filed this lawsuit seeking to represent a class of all owners of copyrighted
`works used as training data for Llama. They brought claims for direct copyright infringement
`(based on Meta’s reproduction of their books), vicarious copyright infringement, removal of
`copyright management information in violation of the Digital Millennium Copyright Act
`(DMCA), unfair competition, unjust enrichment, and negligence. The plaintiffs seek damages,
`restitution, and injunctive and declaratory relief, although it is not entirely clear what exactly
`they seek to enjoin. They did not, for instance, seek a preliminary injunction preventing Meta
`from using their works as training data or requiring Meta to retrain the existing Llama models on
`data excluding their books. Cf. Concord Music Group, Inc. v. Anthropic PBC, No. 24-cv-3811,
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 13 of 40
`
`
`
`
`
`
`
`
`14
`2025 WL 904333, at *3–4 (N.D. Cal. Mar. 25, 2025) (discussing music publishers’ motion for a
`preliminary injunction seeking relief based on future training of AI models).
`All of the claims except the direct copyright infringement claim were dismissed early on.
`The plaintiffs were later granted leave to amend to expand their copyright claim to encompass a
`theory of infringement by distribution (based on the allegation that Meta was reuploading the
`data it torrented), and to add both a different DMCA claim and a claim under the California
`Comprehensive Computer Data Access and Fraud Act (CDAFA). Meta moved to dismiss the
`new claims, and its motion was granted as to the CDAFA claim but denied as to the DMCA
`claim.
`Often, the next step in a case like this is a motion for class certification by the named
`plaintiffs. But sometimes the parties will first move for summary judgment regarding the
`individual claims of the named plaintiffs. For defendants in such cases, there is a trade-off. On
`the one hand, a defendant could benefit by getting a favorable ruling and disposing of the case
`before being subjected to expensive and burdensome class-related discovery and motion practice.
`On the other hand, this favorable ruling for the defendant binds only the individual named
`plaintiffs, leaving all other members of the proposed class free to sue on the same claims. In this
`case, Meta proposed doing summary judgment regarding the individual claims of the named
`plaintiffs first, and the Court accepted this approach.
`Thus, after the close of discovery

This document is available on Docket Alarm but you must sign up to view it.


Or .

Accessing this document will incur an additional charge of $.

After purchase, you can access this document again without charge.

Accept $ Charge
throbber

Still Working On It

This document is taking longer than usual to download. This can happen if we need to contact the court directly to obtain the document and their servers are running slowly.

Give it another minute or two to complete, and then try the refresh button.

throbber

A few More Minutes ... Still Working

It can take up to 5 minutes for us to download a document if the court servers are running slowly.

Thank you for your continued patience.

This document could not be displayed.

We could not find this document within its docket. Please go back to the docket page and check the link. If that does not work, go back to the docket and refresh it to pull the newest information.

Your account does not support viewing this document.

You need a Paid Account to view this document. Click here to change your account type.

Your account does not support viewing this document.

Set your membership status to view this document.

With a Docket Alarm membership, you'll get a whole lot more, including:

  • Up-to-date information for this case.
  • Email alerts whenever there is an update.
  • Full text search for other cases.
  • Get email alerts whenever a new case matches your search.

Become a Member

One Moment Please

The filing “” is large (MB) and is being downloaded.

Please refresh this page in a few minutes to see if the filing has been downloaded. The filing will also be emailed to you when the download completes.

Your document is on its way!

If you do not receive the document in five minutes, contact support at support@docketalarm.com.

Sealed Document

We are unable to display this document, it may be under a court ordered seal.

If you have proper credentials to access the file, you may proceed directly to the court's system using your government issued username and password.


Access Government Site

We are redirecting you
to a mobile optimized page.





Document Unreadable or Corrupt

Refresh this Document
Go to the Docket

We are unable to display this document.

Refresh this Document
Go to the Docket