`
`
`UNITED STATES DISTRICT COURT
`NORTHERN DISTRICT OF CALIFORNIA
`
`RICHARD KADREY, et al.,
`Plaintiffs,
`v.
`
`META PLATFORMS, INC.,
`Defendant.
`
`Case No. 23-cv-03417-VC
`
`
`ORDER DENYING THE PLAINTIFFS’
`MOTION FOR PARTIAL SUMMARY
`JUDGMENT AND GRANTING
`META’S CROSS-MOTION FOR
`PARTIAL SUMMARY JUDGMENT
`Re: Dkt. Nos. 482, 501
`
`
`Companies are presently racing to develop generative artificial intelligence models—
`software products that are capable of generating text, images, videos, or sound based on
`materials they’ve previously been “trained” on. Because the performance of a generative AI
`model depends on the amount and quality of data it absorbs as part of its training, companies
`have been unable to resist the temptation to feed copyright-protected materials into their
`models—without getting permission from the copyright holders or paying them for the right to
`use their works for this purpose. This case presents the question whether such conduct is illegal.
`Although the devil is in the details, in most cases the answer will likely be yes. What
`copyright law cares about, above all else, is preserving the incentive for human beings to create
`artistic and scientific works. Therefore, it is generally illegal to copy protected works without
`permission. And the doctrine of “fair use,” which provides a defense to certain claims of
`copyright infringement, typically doesn’t apply to copying that will significantly diminish the
`ability of copyright holders to make money from their works (thus significantly diminishing the
`incentive to create in the future). Generative AI has the potential to flood the market with endless
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 1 of 40
`
`
`
`
`
`
`
`
`2
`amounts of images, songs, articles, books, and more. People can prompt generative AI models to
`produce these outputs using a tiny fraction of the time and creativity that would otherwise be
`required. So by training generative AI models with copyrighted works, companies are creating
`something that often will dramatically undermine the market for those works, and thus
`dramatically undermine the incentive for human beings to create things the old-fashioned way.
`Take, for example, biographies. If a company uses copyrighted biographies to train a
`model, and if the model is thus capable of generating endless amounts of biographies, the market
`for many of the copied biographies could be severely harmed. Perhaps not the market for Robert
`Caro’s Master of the Senate, because that book is at the top of so many people’s lists of
`biographies to read. But you can bet that the market for lesser-known biographies of Lyndon B.
`Johnson will be affected. And this, in turn, will diminish the incentive to write biographies in the
`future.
`Or take magazine articles. If a company uses copyrighted magazine articles to train a
`model capable of generating similar articles, it’s easy to imagine the market for the copied
`articles diminishing substantially. Especially if the AI-generated articles are made available for
`free. And again, how will this affect the incentive for human beings to put in the effort necessary
`to produce high-quality magazine articles?
`With some types of works, the picture is a bit murkier. For example, it’s not clear how
`generative AI would affect the market for memoirs or autobiographies, since by definition people
`read those works because of who wrote them. With fiction, it might depend on the type of book.
`Perhaps classic works of literature like The Catcher in the Rye would not see their markets
`diminished. But the market for the typical human-created romance or spy novel could be
`diminished substantially by the proliferation of similar AI-created works. And again, the
`proliferation of such works would presumably diminish the incentive for human beings to write
`romance or spy novels in the first place.
`Some students of copyright law respond that none of this matters because when
`companies use copyrighted works to train generative AI models, they are using the works in a
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 2 of 40
`
`
`
`
`
`
`
`
`3
`way that’s highly creative in its own right. In the language of copyright law, the companies’ use
`of the works is “transformative.” As a factual matter, there’s no disputing that. And as a legal
`matter, it’s true that you’re less likely to be liable for copyright infringement if you’re copying
`the work for a transformative purpose. In that situation, you’re more likely to be protected by the
`fair use doctrine. But as the Supreme Court has emphasized, the fair use inquiry is highly fact
`dependent, and there are few bright-line rules. There is certainly no rule that when your use of a
`protected work is “transformative,” this automatically inoculates you from a claim of copyright
`infringement. And here, copying the protected works, however transformative, involves the
`creation of a product with the ability to severely harm the market for the works being copied, and
`thus severely undermine the incentive for human beings to create. Under the fair use doctrine,
`harm to the market for the copyrighted work is more important than the purpose for which the
`copies are made.
`Speaking of which, in a recent ruling on this topic, Judge Alsup focused heavily on the
`transformative nature of generative AI while brushing aside concerns about the harm it can
`inflict on the market for the works it gets trained on. Such harm would be no different, he
`reasoned, than the harm caused by using the works for “training schoolchildren to write well,”
`which could “result in an explosion of competing works.” Order on Fair Use at 28, Bartz v.
`Anthropic PBC, No. 24-cv-5417 (N.D. Cal. June 23, 2025), Dkt. No. 231. According to Judge
`Alsup, this “is not the kind of competitive or creative displacement that concerns the Copyright
`Act.” Id. But when it comes to market effects, using books to teach children to write is not
`remotely like using books to create a product that a single individual could employ to generate
`countless competing works with a miniscule fraction of the time and creativity it would
`otherwise take. This inapt analogy is not a basis for blowing off the most important factor in the
`fair use analysis.
`Another argument offered in support of the companies is more rhetorical than legal:
`Don’t rule against them, or you’ll stop the development of this groundbreaking technology. The
`technology is certainly groundbreaking. But the suggestion that adverse copyright rulings would
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 3 of 40
`
`
`
`
`
`
`
`
`4
`stop this technology in its tracks is ridiculous. These products are expected to generate billions,
`even trillions, of dollars for the companies that are developing them. If using copyrighted works
`to train the models is as necessary as the companies say, they will figure out a way to
`compensate copyright holders for it.
`The upshot is that in many circumstances it will be illegal to copy copyright-protected
`works to train generative AI models without permission. Which means that the companies, to
`avoid liability for copyright infringement, will generally need to pay copyright holders for the
`right to use their materials.
`But that brings us to this particular case. The above discussion is based in significant part
`on this Court’s general understanding of generative AI models and their capabilities. Courts can’t
`decide cases based on general understandings. They must decide cases based on the evidence
`presented by the parties.
`In this case, thirteen authors—mostly famous fiction writers—have sued Meta for
`downloading their books from online “shadow libraries” and using the books to train Meta’s
`generative AI models (specifically, its large language models, called Llama). The parties have
`filed cross-motions for partial summary judgment, with the plaintiffs arguing that Meta’s
`conduct cannot possibly be fair use, and with Meta responding that its conduct must be
`considered fair use as a matter of law. In connection with these fair use arguments, the plaintiffs
`offer two primary theories for how the markets for their works are affected by Meta’s copying.
`They contend that Llama is capable of reproducing small snippets of text from their books. And
`they contend that Meta, by using their works for training without permission, has diminished the
`authors’ ability to license their works for the purpose of training large language models. As
`explained below, both of these arguments are clear losers. Llama is not capable of generating
`enough text from the plaintiffs’ books to matter, and the plaintiffs are not entitled to the market
`for licensing their works as AI training data. As for the potentially winning argument—that Meta
`has copied their works to create a product that will likely flood the market with similar works,
`causing market dilution—the plaintiffs barely give this issue lip service, and they present no
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 4 of 40
`
`
`
`
`
`
`
`
`5
`evidence about how the current or expected outputs from Meta’s models would dilute the market
`for their own works.
`Given the state of the record, the Court has no choice but to grant summary judgment to
`Meta on the plaintiffs’ claim that the company violated copyright law by training its models with
`their books. But in the grand scheme of things, the consequences of this ruling are limited. This
`is not a class action, so the ruling only affects the rights of these thirteen authors—not the
`countless others whose works Meta used to train its models. And, as should now be clear, this
`ruling does not stand for the proposition that Meta’s use of copyrighted materials to train its
`language models is lawful. It stands only for the proposition that these plaintiffs made the wrong
`arguments and failed to develop a record in support of the right one.
`I. COPYRIGHT LAW AND FAIR USE
`The goal of copyright law is to promote “broad public availability of literature, music,
`and the other arts.” Twentieth Century Music Corp. v. Aiken, 422 U.S. 151, 156 (1975). To this
`end, copyright law incentivizes creativity by giving authors of original works a bundle of
`exclusive rights—for instance, the rights to prevent others from reproducing or distributing the
`works. 17 U.S.C. § 106. At the same time, however, copyright law “trades off the benefits of
`incentives to create against the costs of restrictions on copying.” Andy Warhol Foundation for
`the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508, 526 (2023). For example, copyright only
`protects expression, not underlying ideas, and the duration of copyright protection is limited. See
`id. (citing 17 U.S.C. §§ 102, 302–305).
`One major way the Copyright Act strikes a balance between protecting ownership and
`leaving room for innovation is through the affirmative defense of fair use. Under this doctrine,
`“the fair use of a copyrighted work . . . for purposes such as criticism, comment, news reporting,
`teaching . . . , scholarship, or research, is not an infringement of copyright.” 17 U.S.C. § 107.
`Fair use “permits courts to avoid rigid application of the copyright statute when, on occasion, it
`would stifle the very creativity which that law is designed to foster.” Google LLC v. Oracle
`America, Inc., 593 U.S. 1, 18 (2021) (quoting Stewart v. Abend, 495 U.S. 207, 236 (1990)).
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 5 of 40
`
`
`
`
`
`
`
`
`6
`The Copyright Act lists four factors to be considered in determining whether a given use
`is fair:
`1. the purpose and character of the use, including whether such use
`is of a commercial nature or is for nonprofit educational purposes;
`
`2. the nature of the copyrighted work;
`
`3. the amount and substantiality of the portion used in relation to the
`copyrighted work as a whole; and
`
`4. the effect of the use upon the potential market for or value of the
`copyrighted work.
`17 U.S.C. § 107.
` While the statute lists these four factors, fair use is a “flexible concept.” Warhol, 598 U.S.
`at 527 (quotation marks omitted) (quoting Oracle, 593 U.S. at 20). The list is not exhaustive. A
`particular factor “may prove more important in some contexts than in others.” Oracle, 593 U.S.
`at 19. And application of the factors “requires judicial balancing, depending upon relevant
`circumstances, including ‘significant changes in technology.’” Id. (quoting Sony Corp. of
`America v. Universal City Studios, Inc., 464 U.S. 417, 430 (1984)). The factors may also overlap
`such that facts relevant to one factor are also relevant to others. See A.V. ex rel. Vanderhye v.
`iParadigms, LLC, 562 F.3d 630, 642 (4th Cir. 2009). Overall, the factors are not meant to be
`applied mechanically, but to contribute “to a holistic inquiry”: whether the secondary work is
`likely to substitute for the original work in the marketplace and therefore undermine the
`incentive to create. See Romanova v. Amilus Inc., 138 F.4th 104, 117 n.9 (2d Cir. 2025) (Leval,
`J.); see also Warhol, 598 U.S. at 528 (referring to substitution as “copyright’s bête noire”).
`Because it “focuses on actual or potential market substitution,” Warhol, 598 U.S. at 536
`n.12, the fourth factor is “undoubtedly the single most important element of fair use,” Harper &
`Row Publishers, Inc. v. Nation Enterprises, 471 U.S. 539, 566 (1985). If the law allowed people
`to copy your creations in a way that would diminish the market for your works, this would
`diminish your incentive to create more in the future. Thus, the key question in virtually any case
`where a defendant has copied someone’s original work without permission is whether allowing
`people to engage in that sort of conduct would substantially diminish the market for the original
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 6 of 40
`
`
`
`
`
`
`
`
`7
`work. See Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569, 590 (1994).
` Because fair use is an affirmative defense, the burden of proof is on the party invoking it.
`Dr. Seuss Enterprises, L.P. v. ComicMix LLC, 983 F.3d 443, 459 (9th Cir. 2020), abrogated on
`other grounds by Jack Daniel’s Properties, Inc. v. VIP Products LLC, 599 U.S. 140 (2023). In
`particular, because the fourth factor is the most important, the secondary user (generally, the
`defendant) will “have difficulty carrying the burden of demonstrating fair use without favorable
`evidence about relevant markets.” Campbell, 510 U.S. at 590. But while the rightsholder need
`not prove or present evidence of market harm, they “may bear some initial burden of identifying
`relevant markets.” Hachette Book Group, Inc. v. Internet Archive, 115 F.4th 163, 194 (2d Cir.
`2024); see also Newegg Inc. v. Ezra Sutton, P.A., No. CV 15-01395, 2016 WL 6747629, at *2
`(C.D. Cal. Sep. 13, 2016). Moreover, because fair use is a holistic inquiry, the party invoking it
`“bears the burden on the defense as a whole,” not as to each individual factor. William F. Patry,
`Patry on Fair Use § 2:5 (May 2025 ed.).
` Fair use is a mixed question of law and fact, but the “question primarily involves legal
`work.” Oracle, 593 U.S. at 24. Therefore, fair use can be addressed at summary judgment where
`there are no genuine issues of material fact relevant to fair use. Leadsinger, Inc. v. BMG Music
`Publishing, 512 F.3d 522, 530 (9th Cir. 2008). By contrast, where there are genuine factual
`disputes that might affect whether the defendant’s use was fair, those disputes must be resolved
`by a jury. See Oracle, 593 U.S. at 23–25. Once a jury finds the facts, whether those facts support
`fair use “is a legal question for judges to decide.” Id. at 23–24.
` It bears emphasis that where a fair use defense fails, the consequence isn’t necessarily
`that the defendant must stop whatever they were doing. The consequence will often be that the
`defendant needs to pay the copyright owner for a license that grants them permission to do
`whatever they were doing. This way, the defendant compensates the copyright owner for the fact
`that the defendant’s conduct will otherwise harm the market for the original work. The defendant
`will only be forced to stop what they’re doing if they’re unwilling or unable to pay for the right
`to do it.
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 7 of 40
`
`
`
`
`
`
`
`
`8
`II. FACTS AND PROCEDURAL HISTORY
`A
`“Generative AI” is a type of artificial intelligence that creates new content, such as text,
`images, videos, or sound.1 Generative AI models do this, as Meta describes it, by extracting
`“increasingly complex mathematical patterns from training data, enabling the network to output
`a prediction or decision based on the patterns derived.” To put it more simply, generative AI
`models are “trained” to identify common patterns across large training datasets. They can then
`create, in response to user prompts, new content based on the patterns they have recognized in
`that training data. By the same token, a model’s outputs are limited based on the patterns that
`existed in its training data. For instance, if the only bridge in an image-generating model’s
`training data was the Golden Gate Bridge, and a user told that model to generate an image of a
`bridge, it would likely generate an orange-red suspension bridge, because that is the pattern of a
`bridge that would emerge from its training data.
`A large language model, or LLM, is a particular type of generative AI model designed to
`understand and generate text. Users can prompt LLMs to do a wide range of things, such as draft
`emails, summarize documents, or write computer code. Well-known LLMs include OpenAI’s
`ChatGPT models and Google’s Gemini models.
`LLMs learn to understand language by analyzing relationships among words and
`punctuation marks in their training data. The units of text—words and punctuation marks—on
`which LLMs are trained are often referred to as “tokens.” LLMs are trained on an immense
`amount of text and thereby learn an immense amount about the statistical relationships among
`words. Based on what they learned from their training data, LLMs can create new text by
`predicting what words are most likely to come next in sequences. This allows them to generate
`text responses to basically any user prompt. Model developers may also “post-train” or
`“finetune” their models to improve their performance at specific tasks or otherwise adjust their
`outputs, such as to prevent generation of offensive statements. Therefore, as with other
`
`1 Except as noted, the parties do not dispute the facts described in this section.
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 8 of 40
`
`
`
`
`
`
`
`
`9
`generative AI models, LLMs’ outputs are limited by their training data. To be able to generate a
`wide range of text—in different languages or styles, or regarding different subject matter—an
`LLM’s training dataset must be large and diverse. As one Meta witness put it, “If a model only
`saw social media posts, for example, it would not do well in generating source computer code.”
`But while a variety of text is necessary for training, books make for especially valuable
`training data. This is because they provide very high-quality data for training an LLM’s
`“memory” and allowing it to work with larger amounts of text at once. (The technical term for
`how many tokens an LLM can hold in its memory at once is its “context window.”) For instance,
`an LLM with a better memory will be able to process and respond to longer prompts, incorporate
`more information into outputs, and remember things from earlier in an exchange, resulting in
`smoother “conversations.” Books are good data for training LLMs’ memories because, in the
`words of one of Meta’s expert witnesses, they are “long but consistent,” maintaining a particular
`style and coherent structure. They are also high quality in the sense that they generally are well
`written and use proper grammar (especially compared to text from the internet, which varies
`widely on these metrics).
`B
`Meta Platforms owns and operates social media services including Facebook, Instagram,
`and WhatsApp. It is also the developer of a series of LLMs named “Llama.” Meta released
`Llama 1 in February 2023 and Llama 2 that July. Llama 3—along with Meta AI, an easily
`accessible AI chatbot (analogous to ChatGPT) that incorporates Llama 3—was released in April
`2024. Llama 4 is planned for release later in 2025. As Meta explains, each new Llama edition
`improved in certain ways over its predecessor: Llama 2 was finetuned “to improve the safety,
`quality, and consistency” of its outputs; Llama 3 made “significant improvements in
`performance and efficiency”; and Llama 4 is generally “larger” and “more advanced.” Subject to
`certain restrictions, members of the public can download all of the Llama models for free for
`noncommercial use; Llama 2 and 3 are also free to download for commercial use. While the
`Llama models are free to download, Meta estimates that its total revenue from generative AI will
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 9 of 40
`
`
`
`
`
`
`
`
`10
`range from $2 to $3 billion in 2025, and from $460 billion to $1.4 trillion over the next ten years.
`See Pls. MSJ Ex. 8 at 12.
`To get the varied and extensive text necessary to train its models, Meta cast a wide net.
`Approximately two-thirds of the data used to train Llama 1 and 2 came from Common Crawl, a
`nonprofit organization that collects and provides free access to website data, metadata, and text.
`The remainder came from websites and databases including Wikipedia, GitHub, ArXiv, Stack
`Exchange, and a combination of Project Gutenberg and Books3 (two book databases).2 With the
`exception of Books3, none of the sources contained any copyrighted material at issue in this
`case.
`Although Meta needed (and acquired and used) a wide range of training data, it
`especially needed books because, as discussed above, books make for high-quality data. Meta AI
`researchers and engineers repeatedly discussed the benefits of using books as training data, as
`well as the need to acquire more books for this use. One Meta employee said that the “best
`resources we can think of are definitely books.” Pls. MSJ Ex. 18 at 2. Another said it was “really
`important for us to get books data ASAP.” Id. Ex. 40 at 2. So as Meta expanded its datasets
`generally, it also continued to look for more books in particular.
`At first, Meta wanted to license books and so tried to negotiate licensing deals with
`several major publishers. Meta’s head of generative AI discussed spending up to $100 million on
`licensing. But as negotiations proceeded, Meta realized that licensing would be more difficult
`than anticipated. For one thing, publishers generally do not hold the subsidiary rights to license
`books for AI training. These rights are instead held by individual authors, and there is no
`organization for collective licensing of such rights. Sinkinson Decl. ISO Meta MSJ ¶¶ 58–59, 62.
`Even where publishers do hold AI training licensing rights, they do so regionally rather than
`globally. Meta MSJ Ex. 34 at 22:22–25:15. For another thing, some publishers apparently
`
`2 According to Meta, GitHub is “a leading cloud-based platform where coders store and share
`code, frequently on an open-source basis.” ArXiv is “a free online archive of math, science, and
`economics papers.” Stack Exchange is “a network of question-and-answer websites for sharing
`technical knowledge, geared toward the programming community.”
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 10 of 40
`
`
`
`
`
`
`
`
`11
`ignored Meta’s outreach, and only one gave Meta a pricing proposal. Id. at 23:11–14, 24:2–10.
`Eventually, Meta began investigating the possibility of procuring the books (and other
`text) needed for training by downloading them from “shadow libraries.” A shadow library is an
`online repository that provides things like books, academic journal articles, music, or films for
`free download, regardless of whether that media is copyrighted. Meta first used a shadow library
`in October 2022, when it downloaded the Library Genesis (“LibGen”) database to investigate
`whether there was value in training Llama on the works it contained. Pls. MSJ Ex. 32 at 3. If the
`answer was yes, the plan was to then set up licensing agreements for those or similar works. Id.
`But in spring 2023, after failing to acquire licenses and following escalation to CEO Mark
`Zuckerberg, Meta decided to just use the works acquired from LibGen as training data. Id. Ex. 61
`at 5. And after confirming that LibGen contained most of the works available for license from
`certain publishers with which it had been negotiating, Meta abandoned its licensing efforts. Id.
`Ex. 50 at 131:1–132:10, 383:5–384:12; id. Ex. 57 at 2; id. Ex. 58 at 3; see also id. Ex. 92 at 12.
`In early 2024, Meta also downloaded Anna’s Archive, a compilation of shadow libraries
`including LibGen, Z-Library, and others. See id. Ex. 66 at 2–3.
`To download these large datasets more quickly and without unnecessarily slowing down
`its networks, Meta torrented them. “Torrenting” is a filesharing technique that entails the
`simultaneous distribution of small portions of a larger file from many different sources. To be
`more precise, those sources are many other computer systems that also contain that file. So, for
`instance, one who torrented LibGen would download small pieces of each book LibGen contains
`from other users who had copies of LibGen on their computer and who were participating in the
`torrenting network. The torrenting software would then take those pieces and reassemble them
`into the original files on the downloader’s computer.3
`Certain torrenting protocols—including the one used by Meta, called BitTorrent—are, by
`
`3 See generally David Gerwitz, What Is Torrenting?, ZDNET (Aug. 6, 2024),
`https://www.zdnet.com/article/what-is-torrenting-and-how-does-it-work [https://perma.cc/8PG5-
`H7UW].
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 11 of 40
`
`
`
`
`
`
`
`
`12
`default, configured so that files downloaded via torrenting may also be reuploaded to other
`computer systems. This reuploading can occur both while files are still being downloaded (which
`the parties refer to as “leeching”) and after those files have been fully downloaded (which the
`parties refer to as “seeding”). Some torrenting protocols—including BitTorrent—are designed to
`prioritize downloads to users who are also uploading.
`There is no dispute that Meta torrented LibGen and Anna’s Archive, but the parties
`dispute whether and to what extent Meta uploaded (via leeching or seeding) the data it torrented.
`A Meta engineer involved in the torrenting wrote a script to prevent seeding, but apparently not
`leeching. See Pls. MSJ at 13; id. Ex. 71 ¶¶ 16–17, 19; id. Ex. 67 at 3, 6–7, 13–16, 24–26; see also
`Meta MSJ Ex. 38 at 4–5. Therefore, say the plaintiffs, because BitTorrent’s default settings allow
`for leeching, and because Meta did nothing to change those default settings, Meta must have
`reuploaded “at least some” of the data Meta downloaded via torrent. The plaintiffs assert further
`that Meta chose not to take any steps to prevent leeching because that would have slowed its
`download speeds. Meta responds that, even if it reuploaded some of what it downloaded, that
`doesn’t mean it reuploaded any of the plaintiffs’ books. It also notes that leeching was not clearly
`an issue in the case until recently, and so it has not yet had a chance to fully develop evidence to
`address the plaintiffs’ assertions.
`Either way, Meta added the books it downloaded to the datasets it used to train the Llama
`models. It also post-trained its models to prevent them from “memorizing” and outputting certain
`text from their training data, including copyrighted material. These training efforts, which Meta
`calls “mitigations,” appear to have been successful. Meta’s expert witness tested them using a
`method designed to get LLMs to regurgitate material from its training data (which Meta calls
`“adversarial prompting”). Even using that method, the expert could get no model to generate
`more than 50 words and punctuation marks (that is, “tokens”) from the plaintiffs’ books. And the
`plaintiffs’ expert could only get the Llama model best at regurgitation to generate 50 words and
`punctuation marks from the plaintiffs’ books in 60% of tests. She also testified that Llama was
`not able to reproduce “any significant percentage” of them. Meta MSJ Ex. 24 at 237:16–19; see
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 12 of 40
`
`
`
`
`
`
`
`
`13
`also Pls. Ex. 79 ¶¶ 70–72, 79, 82–83, 92; Meta MSJ Ex. 23 at 179:22–25, 180:17–181:16. In
`short, Llama cannot currently be used to read or otherwise meaningfully access the plaintiffs’
`books.
`C
`The plaintiffs are thirteen published authors who have written, and who hold copyright
`in, various works. Those works are mostly novels, but also include plays, short stories, memoirs,
`essays, and nonfiction books. Examples include Sarah Silverman’s The Bedwetter, a comic
`memoir; Rachel Louise Snyder’s No Visible Bruises: What We Don’t Know About Domestic
`Violence Can Kill Us, a nonfiction book about domestic violence and how to combat it; Junot
`Díaz’s Pulitzer Prize–winning novel, The Brief Wondrous Life of Oscar Wao; and Andrew Sean
`Greer’s Less, also a Pulitzer Prize–winning novel. All of the books in which the plaintiffs hold
`copyright can be found in the datasets Meta downloaded, including both Books3 and the Anna’s
`Archive databases. In total, Meta downloaded at least 666 copies of books whose copyrights the
`plaintiffs hold.
`Each plaintiff says that they would be open to licensing their books for use as generative
`AI training data, but that Meta did not approach them about this licensing. No plaintiff has
`licensed a book to any company for use as LLM training data or been asked by any company to
`license a book for that purpose.
`The plaintiffs filed this lawsuit seeking to represent a class of all owners of copyrighted
`works used as training data for Llama. They brought claims for direct copyright infringement
`(based on Meta’s reproduction of their books), vicarious copyright infringement, removal of
`copyright management information in violation of the Digital Millennium Copyright Act
`(DMCA), unfair competition, unjust enrichment, and negligence. The plaintiffs seek damages,
`restitution, and injunctive and declaratory relief, although it is not entirely clear what exactly
`they seek to enjoin. They did not, for instance, seek a preliminary injunction preventing Meta
`from using their works as training data or requiring Meta to retrain the existing Llama models on
`data excluding their books. Cf. Concord Music Group, Inc. v. Anthropic PBC, No. 24-cv-3811,
`Case 3:23-cv-03417-VC Document 598 Filed 06/25/25 Page 13 of 40
`
`
`
`
`
`
`
`
`14
`2025 WL 904333, at *3–4 (N.D. Cal. Mar. 25, 2025) (discussing music publishers’ motion for a
`preliminary injunction seeking relief based on future training of AI models).
`All of the claims except the direct copyright infringement claim were dismissed early on.
`The plaintiffs were later granted leave to amend to expand their copyright claim to encompass a
`theory of infringement by distribution (based on the allegation that Meta was reuploading the
`data it torrented), and to add both a different DMCA claim and a claim under the California
`Comprehensive Computer Data Access and Fraud Act (CDAFA). Meta moved to dismiss the
`new claims, and its motion was granted as to the CDAFA claim but denied as to the DMCA
`claim.
`Often, the next step in a case like this is a motion for class certification by the named
`plaintiffs. But sometimes the parties will first move for summary judgment regarding the
`individual claims of the named plaintiffs. For defendants in such cases, there is a trade-off. On
`the one hand, a defendant could benefit by getting a favorable ruling and disposing of the case
`before being subjected to expensive and burdensome class-related discovery and motion practice.
`On the other hand, this favorable ruling for the defendant binds only the individual named
`plaintiffs, leaving all other members of the proposed class free to sue on the same claims. In this
`case, Meta proposed doing summary judgment regarding the individual claims of the named
`plaintiffs first, and the Court accepted this approach.
`Thus, after the close of discovery



