Showing posts with label Internet Archive. Show all posts
Showing posts with label Internet Archive. Show all posts

Thursday, April 13, 2023

What is Controlled Digital Lending? The Origin Story

 The bulk of the reporting on the lawsuit between publishers (four of them, led by Hachette) and the Internet Archive's version of Controlled Digital Lending hits one of these points of view:

  •  The publishers are evil, money-grubbing idiots going after the generous, saintly Internet Archive
  •  The Internet Archive is evil, stealing from the poor publishers and even poorer writers

As is so often the case, it really is more compex than that. I will try to throw a bit of clarity into mix here, mainly by talking about some of the realities of library service in the 21st century, and the origins of controlled digital lending.

The Origins of CDL

Michelle M. Wu, a law librarian and law professor, wrote a piece for the Law Library Journal in 2011 explaining the dilemmas faced by law libraries and proposing a modest solution.

"Building a Collaborative Digital Collection: A Necessary Evolution in Libraries" LAW LIBRARY JOURNAL Vol. 103:4 [2011-34]) (online)

The solution is what became Controlled Digital Lending. The reasons she lays out are the key.

The main argument that Wu puts forth (and that I find convincing) is this: library users either want or actually need to be able to access materials remotely, which means in electronic formats over a network. Increasingly, materials that libraries wish to provide are available from publishers in those electronic formats. The catch, however, is that libraries are not able to own materials in electronic formats, but instead can only subscribe to access services. It is this lack of ownership that is the rub. If a library loses its digital subscription for some reason, such as no longer being able to afford it, it not only loses access to future materials, it loses access to all of the past materials that were included in that subscription. This puts libraries in the terrible position of having to decide between fulfilling their role as the reliable repository and archive of material in their subject area, or of serving the needs of library users. As Wu points out, libraries are already struggling to afford the materials that they feel they should be collecting, so purchasing these materials both in hard copy for archival purposes and also in digital form for user service is entirely beyond the pale.

What Wu suggests in her article is a variation on Inter-Library Loan, combined with a library collective purchasing plan. A cooperative group of law libraries would combine purchasing physical resources for those items that are rarely used but that should be available to the researchers who need them. This is not a revolutionary idea - library consortia have been making use of this kind of approach for a significant amount of time. The difference in Wu's plan is that as items are requested from the consortial holdings, they will be digitized and the digital format will be the one loaned. To stay within the intention of copyright law, in particular First Sale, Wu offers that the digital file will be loaned as a surrogate for the physical copy:

"Materials acquired would be digitized, and only the number of copies acquired in print for each subsequently digitized document would circulate at any given time. The print copy would be stored for archival purposes; only the digital copy would “circulate.”" Wu p.535
The physical resource would fulfill the need for an archival copy, and the digital resource would allow lending to any networked member of the cooperative group. Her solution assumes an effective digital rights management system that would make the loan a loan and not a pirate-able copy.

Wu carefully covers all of the potential legal objections, and points out the various areas of US copyright law that might be touched on with her proposal. Specific areas are First Sale, Fair Use, and the various exceptions to the copyright law that are applied to libraries. She defends the digitization as format-shifting, not unlike the format-shifting that is done for sound recordings as the technology for that medium has changed.
"It is the work itself that is copyrighted, not the form." Wu p. 541
She also addresses what would be an obvious objection of rights holders, that the digital copy is substituting for a purchase. The hard copy would be purchased by the consortium, and given her statement that these would primarily be low-use materials that many libraries would not themselves purchase, no harm would be done to the market which would be limited.

The argument that I find strongest is that of preservation: the US copyright law does allow libraries to make copies of works for the purposes of preservation if no equivalent copy is available for purchase. (Section 108 subsection c) Using the argument that the purpose of the library is to preserve as well as to make works available, Wu says:
"In cases where a digital version is available only for license, a library could argue that such a license is not equivalent to either the print copy or a digital copy they would make, because both of these items would be owned by the library and the licensed digital version would not." Wu p. 539

Context Counts

Controlled Digital Lending is the technology: the digitization of paper works and the lending of the digital copy using management software that prevents piracy of the digital file.  It is the context that makes Wu's proposal different to the implementation of controlled digital lending at the Internet Archive.
  • Wu's proposal was for a consortium of law libraries serving their own users; the IA's implementation was open to anyone on the web
  • Wu's proposal was for academic materials of low use; the IA's included popular works
  • The works in Wu's proposal would have been selected with specific research purposes in mind; the IA's collection was an opportunistic group of books that they had often obtained as second-hand - therefore no research purpose could be argued. (Purpose is one of the Fair Use factors.)
  • Wu's proposal argued for the need of libraries to preserve materials that otherwise would not be preserved; the Archive is indeed an archive, known for its preservation of web sites that otherwise would be lost. However, the popular books named in the lawsuit against the Archive are already "preserved" in thousands of libraries who have those physical books on their shelves - the preservation argument is not easily supplied. 

HathiTrust, which is a consortium of libraries that originally contributed to the Google Books project, is an example that follows Wu's approach. HathiTrust  stores digitized copies of books and follows the ruling related to Google Books that searching of in-copyright works is permitted, but not reading. HathiTrust developed its own controlled digital lending service as an emergency service for when a member library is temporarily closed due to a disaster. In that case, users from the member library can borrow digitized books held by the member library in hard copy.

Wu even suggests that libraries might share the burden of digitization by providing digitized copies to libraries that own the books in hard copy. This latter, though, was one of the things that got the Internet Archive in trouble because it became a "digital lending broker" for other libraries, adding their hard copy count to the Archive's lending "units" including some of books owned by the publishers in the lawsuit.

The Upshot

The argument presented by Wu is quite strong and is justified through her careful reading of copyright law, in particular as that law applies to libraries. The extension of her proposal to popular reading materials and to an unlimited user base changes everything. Libraries do have specific collections, identified users, and stated purposes that guide their acquisitions. Something I feel strongly about - that is absent from so many modern information activities - is that effective information use requires purposeful resource selection and organization. Any mass of stored resources is only as valuable as its organization and coherence. In some cases, a "less" that is well organized can be more informative than a "more" that may lack the key works in a subject area. It may be old-fashioned on my part, but I adhere to the concept of defined user goals and the deliberate collection of specific works in support of user learning. This is what I read in Wu's work but which I do not see in the Archive's activities.

I think Wu's context could be understood as falling within the confines of copyright law. I'm not sure that the Archive's case does. I do hope that this current lawsuit does not result in a rejection of digitization for lending for all libraries.

Friday, April 07, 2023

Libraries, the law, and equality

 


In the spirit of "everyone is equal under the law", it is equally illegal for both a starving man and a billionaire to steal a loaf of bread. Or to copy a book.

 Libraries for the People

It was not all that long ago when "library" often referred to the room in a rich man's home where he stored books that were only available to him, and perhaps members of his family (especially if they were not female). Other libraries, usually larger ones, were attached to prestigious educational institutions and accessible by people worthy of that prestige (which would not include non-white nor female people). We are fortunate  today that we have these things called "public libraries," libraries that serve everyone regardless of their wealth, their race, or their gender.

Here's the catch: public libraries are generally small and modestly funded by the local community. A moderately sized public library has 50,000 - 100,000 volumes. A large public library may have up to 500,00 volumes. A large university library has many more. Harvard University library claims to have 20 million book volumes, 400 million manuscripts, and 10 million photographs. Stanford University library may have at least 12 million book volumes. Michigan State University libraries have about 7 million book volumes. The British Museum Library lays claim to 170 and 200 million items of which 13.5 million are printed books and e-books. There is no question that the members of our community who are served solely by public libraries, while they have unprecedented access to books, are not able to study the full range of printed knowledge of our world. To whit, the university libraries are often referred to as "research libraries" while the local public libraries are called "reading libraries." This separates us into "readers" and "researchers," and while you might conclude that any literate person can read, only those associated with large libraries will be able to avail themselves of the tools to do research.

Digital Access

Much of the research done in academe consumes and creates journal articles. Originally issued only in paper, and mailed to libraries and departments, journal articles have been available in digital form from the mid-1990's and today it would be unusual for an article-based publication to be issued only in paper. Journal owners have digitized the full run of publications, as have cooperative projects based in academia. A researcher or student at a Western university is likely to have more than a century of scientific, technical or social science academic article output available through the Internet, any day, any time, and perhaps from any place. Anyone from less wealthy nations will have less access, although perhaps just a tad more than they had when the articles were issued only in paper.

The story is different with books. While most academic articles have been converted to digital form, the same cannot be said of books. It is only recently that publishers have issued their books in electronic form using the electronic files that are now part and parcel of the publishing process. That only takes care of current publications, however. Sitting in libraries are centuries of one-off publications in book form. Books from this vast backlog must be digitized from the existing physical copy.  Projects by libraries and educational institutions to digitize the monographic backlog, similar to those that succeeded in digitizing the journal output of the ages, have not been accomplished. There are various reasons why that is the case: the sheer number of book pages that would need to be digitized is huge; non-destructive digitization of bound volumes is difficult and often does not yield good results; partnering with publishers for this task is hampered by the fact that numerous books from the 20th century and older are "orphaned," meaning that although they may be under copyright their copyright holder cannot be found; and compared to modern ebooks, digitized books have little to recommend them for reading, although with their searchable text they may be useful for research.

The only efforts to digitize the backlog of books, Google Book Search and the digitizing by the Internet Archive, have resulted in lawsuits against those organizations. The suit against Google concluded that digitization is allowed as long as the digitized books are provided for purposes of searching but not reading. The Internet Archive took the view that books are for reading, an approach that I find hard to oppose.

Reading vs Research

Reading and research are related but different activities. Reading is often associated with books, and includes books on scientific and academic topics as well as fiction, from great literature to beach reads. While few non-researchers read academic articles, some members of academe do read books as part of their research. Of course, many people also read for pleasure; reading is a key means of acquiring culture, along side other activities like taking in performances of various arts.

If you are not at one of those institutions with a large research library, the only way you may have to see the content of many books is by accessing a digitized book. A digitized book is not the same reading experience as the ebook produced by publishers. A digitized book has not been produced from an electronic file of its contents as an ebook has been. Instead, each page of the physical book has been photographed, and those images have been analyzed using optical character recognition (OCR) software. The result of the OCR is a text file, and that file will be more or less "lossy" depending on things like the condition of the original book pages, the clarity of the font, the language of the text. 

Unlike an ebook, reading the digitized book usually means viewing pictures of the books' pages.

 


It's not a great reading experience, but imagine that the book is important for your studies or your work; it would be worth the effort.

On the other hand, if you are wanting some modern leisure reading and you are in North America, you will be much better served by checking out the book and ebook offerings of your public library. If you are not in North America, and if your locality has a limited public library or no public library at all, then the extra effort that you may need to make to read a digitized book may be worth it to you. If, however, you had the funds to purchase the materials you needed or were associated with an institution that made those materials available to you, it is unlikely that you would choose the less sophisticated and less available copy provided at the Archive.

Hachette, et al., v Internet Archive

The above sets out some of the social parameters that we should consider when thinking about the recent lawsuit relating to Controlled Digital Lending. (See previous post.) In brief, the Internet Archive has digitized many books and makes them available globally, lending one "copy" at a time. A group of publishers has sued the Archive based on a set of books for which the publishers hold the copyright. The issue is often presented as a test of the concept of Controlled Digital Lending, although only some books are in question in the lawsuit. Those books represent only a portion of the books available at the Archive or in libraries in general. Although one may think of a binary division of books into "still in copyright" and "no longer in copyright" the actual situation is more complex.

  1. There are the books that are out of copyright, which generally means books from 1924 and earlier in the US. These are not under discussion. However, there is no way to separate the basic copyrighted content of a book, like Mark Twain's Huck Finn from later reprintings that often add some bit of a preface so that the publisher can put a copyright notice on it and pretend to have the rights. Such "books" may be considered in copyright even though the primary content of those books is not. There is unfortunately no penalty for a publisher in slapping a copyright statement onto a book that is not under copyright, as can be seen in my favorite example of a blank journal sold with a copyright notice.
  2. There are the orphaned works, for which there is no one to assert rights. Either the rights holder (the publisher) no longer exists, or the documentation that would make it possible to assert rights does not exist. Because this is a category of unknowns, it is quite difficult to determine which books are in this category.
  3. There are works that are not orphaned but the publisher is not asserting rights in relation to Controlled Digital Lending. This may be the majority of the books being loaned by the Archive because there are only four publishers in the lawsuit. We don't know what the other publishers think about the lending.
  4. There are the books by the four publishers that are included in the lawsuit. These four publishers  are asserting that the Archive violated their rights and potentially deprived them of income.

It would be great to know the figures that would allow us to compare 1-3 with 4. It would also be great to know how many loans were actually made by the Archive of those books in the 4th category. Presumably that figure will inform the penalty that is imposed on the Archive.

The Archive's defense seems to be solid as it shows that in both the presence and the absence of its contested service no change was noted for publisher sales. It is chilling that the judge so readily dismissed the Archive's arguments, and especially chilling if you consider, as a hypothetical, applying this same argument to libraries in general.

"IA’s experts observed that print sales of the Works in Suit and general demand for library ebooks did not decrease while the Works in Suit were available on IA’s Website; that Amazon rankings for the Works in Suit improved when IA’s digital lending skyrocketed (and government lockdowns were in full effect) at the beginning of the Covid-19 pandemic; and that, despite the removal of the Works in Suit from IA’s library in June 2020, OverDrive checkouts of the Works in Suit did not increase." (Case 1:20-cv-04160-JGK-OTW Document 188 Filed 03/24/23 Page 42)

That sounds like a good defense, yet the judge dismisses it.

"But these metrics do not begin to meet IA’s burden to show a lack of market harm. Taking them at face value, they show at best that the presence of the Works in Suit in IA’s online library correlated, however weakly, with positive financial indicators for the Publishers in other areas. They do not show that IA’s conduct caused these benefits to the Publishers. In any event, IA cannot offset the harm it inflicts on the Publishers’ library ebook revenues, see, e.g., Andy Warhol Found., 11 F.4th at 48; TVEyes, 883 F.3d at 180, by pointing to other asserted benefits to the Publishers in other markets. Nor could those asserted benefits tip the scales in favor of fair use when the other factors point so strongly against fair use." (Case 1:20-cv-04160-JGK-OTW Document 188 Filed 03/24/23 Page 43)
Given this kind of reasoning, there is no "proof" that any library could provide that would clearly absolve the library of harm to publishers. That should be okay because "not harming publishers" is not how we should see the role of libraries in our world. Libraries exist for the same reasons that educational institutions exist: to further the abilities of citizens to participate in "science and the useful arts", as it is called in the constitution. Yet as Dan Cohen says in his article in the Atlantic:
On Friday, the judge sided almost entirely with the publishers. The Internet Archive “argues that its digital lending makes it easier for patrons who live far from physical libraries to access books and that it supports research, scholarship, and cultural participation by making books widely accessible on the Internet,” Judge John G. Koeltl wrote in his pointed ruling. “But these alleged benefits cannot outweigh the market harm to the Publishers.”
Thus, societal benefits, such as those of libraries and schools, take a back seat to profit. Or should I say "alleged benefits." Today, copyright law creates a basis for the legality of library lending through the first sale doctrine. Some library privileges relating to making copies are included in the US copyright law. But these do not add up to actual support for the work of libraries, only a limitation on culpability as they perform key functions such as preserving cultural materials that have been abandoned by their creators and providing access to recorded culture to all who request it. In the legal regime, libraries are allowed, but not encouraged, to provide a valuable service for society. Judge John G. Koeltl has little regard for that service.

Saturday, November 05, 2022

Cautions on ISBN and a bit on DOI

I have been reading through the documents relating to the court case that Hachette has brought against the Internet Archives "controlled digital lending" program. I wrote briefly about this before. In this recent reading I am once again struck by the use and over-use of ISBNs as identifiers. Most of my library peeps know this, but for others, I will lay out the issues and problems with these so-called "identifiers".

"BOOK"

The "BN" of the ISBN stands for "BOOK NUMBER." The "IS" is for "INTERNATIONAL STANDARD" which was issued by the International Standards Organization, whose documents are unfortunately paywalled. But the un-paywalled page defines the target of an ISBN as:

[A] unique international identification system for each product form or edition of a separately available monographic publication published or produced by a specific publisher that is available to the public.

What isn't said here in so many words is that the ISBN does not define a specific content; it defines a  salable product instance in the same way that a UPS code is applied to different sizes and "flavors" of Dawn dish soap. What many people either do not know or may have forgotten is that every book product is given a different ISBN. This means that the hardback book, the trade paperback, the mass-market paperback, the MOBI ebook, the EPUB ebook, even if all brought to market by a single publisher, all have different ISBNs.  

The word "book" is far from precise and it is a shame that the ISBN uses that term. Yes, it is applied to the book trade, but it is not applied to a "book" except in a common sense of that word. When you say "I read a book" you do not often mean the same thing as the B in ISBN. Your listener has no idea if you are referring to a hard back or a paperback copy of the text. It would be useful to think of the ISBN as the ISBpN - the International Standard Book product Number.

Emphasizing the ISBN's use as a product code, bookstores at one point were assigning ISBNs to non-book products like stuffed animals and other gift items. This was because the retail system that the stores used required ISBNs. I believe that this practice has been quashed, but it does illustrate that the ISBN is merely a bar code at a sales point.

1970

The ISBN became a standard product number in the book trade in 1970, in the era when the Universal Product Code (UPC) concept was being developed in a variety of sales environments. This means that every book product that appeared on the market before that date does not have an ISBN. This doesn't mean that a text from before that date cannot have an ISBN - as older works are re-issued for the current market, they, too, are given ISBNs as they are prepared for the retail environment. Even some works that are out of copyright (pre-1925) may be found to have ISBNs when they have been reissued. 

The existence of an ISBN on the physical or electronic version of a book tells you nothing about its copyright status and does not mean that the book content is currently in print. It has the same meaning as the bar code on your box of cereal - it is a product identifier that can be used in automated systems to ring up a purchase. 

The Controlled Digital Lending Lawsuit

The lawsuit between a group of publishers led by Hachette and the Internet Archive is an example of two different views: that of selling and that of reading.

 

In the lawsuit the publishers quantify the damage done to them by expressing the damage to them in terms of numbers of ISBNs. This Implies that the lawsuit is not including back titles that are pre-ISBN. Because the concern is economic, items that are long out of print don't seem to be included in the lawsuit.

The difference between the book as product and the book as content shows up in how ISBNs are used. The publisher’s expert notes that many metadata records at the archive have multiple ISBNs and surmises that the archive is adding these to the records. What this person doesn’t know is that library records, which the archive is using, often contain ISBNs for multiple book products which the libraries consider interchangeable. The library user is seeking specific content and is not concerned with whether the book is a hard back, has library binding, or is one of the possible soft covers. The “book “ that the library user is seeking is an information vessel.

It is the practice in libraries, where there is more than one physical book type available, to show the user a single metadata record that doesn’t distinguish between them. The record may describe a hard bound copy even though the library has only the trade paperback. This may not be ideal but the cost-benefit seems defensible. Users probably pay little attention to the publication details that would distinguish between these products. 

 

From a single library metadata record

 

Where libraries do differentiate is between forms that require special hardware or software. Even here however the ISBN cannot be used for the library’s purpose because services that manage these materials can provide the books in the primary digital reading formats based off a single metadata record, even though each ebook format is assigned its own ISBN for the purpose of sales.

The product view is what you see on Amazon. The different products have different prices which is one way they are distinguished. A buyer can see the different prices for hard copy, paperback, or kindle book, and often a range of prices for used copies. Unlike the library user, the Amazon customer has to make a choice, even if all of the options have the same content. For sales to be possible, each of the products has its own ISBN. 

Different products have different prices


Counting ISBNs may be the correct quantifier for the publishers, but they feature only minimally in the library environment. Multiple ISBNs on a single library metadata record is not an attempt to hide publisher products by putting them together; it's good library practice for serving its readers. Users coming to the library with an ISBN will be directed to the content they seek regardless of the particular binding the library owns. Counting the ISBNs in the Internet Archive's metadata will not be a good measure of the number of "books" there using the publisher's definition of "book."


Digital Object Identifier (DOI)


I haven't done a deep study of the use of DOIs, but again there seems to be a great enthusiasm for the DOI as an identifier yet I see little discussion of the limitations of its reach. DOI began in 2000 so it has a serious time limit. Although it has caught on big with academic and scientific publications, it has less reach with social sciences, political writing, and other journalism. Periodicals that do not use DOIs may well be covering topics that can also be found in the DOI-verse. Basing an article research system on the presence of DOIs is an arbitrary truncation of the knowledge universe.

 

The End

 

Identifiers are useful. Created works are messy. Metadata is often inadequate. As anyone who has tried to match up metadata from multiple sources knows, working without identifiers makes that task much more  difficult. However, we must be very clear, when using identifiers, to recognize what they identify.


Monday, March 01, 2021

Digitization Wars, Redux

 (NB: IANAL) 

 Because this is long, you can download it as a PDF here.

From 2004 to 2016 the book world (authors, publishers, libraries, and booksellers) was involved in the complex and legally fraught activities around Google’s book digitization project. Once known as “Google Book Search,” the company claimed that it was digitizing books to be able to provide search services across the print corpus, much as it provides search capabilities over texts and other media that are hosted throughout the Internet. 

Both the US Authors Guild and the Association of American Publishers sued Google (both separately and together) for violation of copyright. These suits took a number of turns including proposals for settlements that were arcane in their complexity and that ultimately failed. Finally, in 2016 the legal question was decided: digitizing to create an index is fair use as long as only minor portions of the original text are shown to users in the form of context-specific snippets. 

We now have another question about book digitization: can books be digitized for the purpose of substituting remote lending in the place of the lending of a physical copy? This has been referred to as “Controlled Digital Lending (CDL),” a term developed by the Internet Archive for its online book lending services. The Archive has considerable experience with both digitization and providing online access to materials in various formats, and its Open Library site has been providing digital downloads of out of copyright books for more than a decade. Controlled digital lending applies solely to works that are presumed to be in copyright. 

Controlled digital lending works like this: the Archive obtains and retains a physical copy of a book. The book is digitized and added to the Open Library catalog of works. Users can borrow the book for a limited time (2 weeks) after which the book “returns” to the Open Library. While the book is checked out to a user no other user can borrow that “copy.” The digital copy is linked one-to-one with a physical copy, so if more than one copy of the physical book is owned then there is one digital loan available for each physical copy. 

The Archive is not alone in experimenting with lending of digitized copies: some libraries have partnered with the Archive’s digitization and lending service to provide digital lending for library-owned materials. In the case of the Archive the physical books are not available for lending. Physical libraries that are experimenting with CDL face the added step of making sure that the physical book is removed from circulation while the digitized book is on loan, and reversing that on return of the digital book. 

Although CDL has an air of legality due to limiting lending to one user at a time, authors and publishers associations had raised objections to the practice. [nwu] However, in March of 2020 the Archive took a daring step that pushed their version of the CDL into litigation: using the closing of many physical libraries due to the COVID pandemic as its rationale, the Archive renamed its lending service the National Emergency Library [nel] and eliminated the one-to-one link between physical and digital copies. Ironically this meant that the Archive was then actually doing what the book industry had accused it of (either out of misunderstanding or as an exaggeration of the threat posed): it was making and lending digital copies beyond its physical holdings. The Archive stated that the National Emergency Library would last only until June of 2020, presumably because by then the COVID danger would have passed and libraries would have re-opened. In June the Archive’s book lending service returned to the one-to-one model. Also in June a suit was filed by four publishers (Hachette, HarperCollins, Penguin Random House, and Wiley) in the US District Court of the Southern District of New York. [suit] 

The Controlled Digital Lending, like the Google Books project, holds many interesting questions about the nature of “digital vs physical,” not only in a legal sense but in a sense of what it means to read and to be a reader today. The lawsuit not only does not further our understanding of this fascinating question; it sinks immediately into hyperbole, fear-mongering, and either mis-information or mis-direction. That is, admittedly, the nature of a lawsuit. What follows here is not that analysis but gives a few of the questions that are foremost in my mind.

 Apples and Oranges 

 Each of the players in this drama has admirable reasons for their actions. The publishers explain in their suit that they are acting in support of authors, in particular to protect the income of authors so that they may continue to write. The Authors’ Guild provides some data on author income, and by their estimate the average full-time author earns less than $20,000 per year, putting them at poverty level.[aghard] (If that average includes the earnings of highly paid best selling authors, then the actual earnings of many authors is quite a bit less than that.) 

The Internet Archive is motivated to provide democratic access to the content of books to anyone who needs or wants it. Even before the pandemic caused many libraries to close the collection housed at the Archive contained some works that are available only in a few research libraries. This is because many of the books were digitized during the Google Books project which digitized books from a small number of very large research libraries whose collections differ significantly from those of the public libraries available to most citizens. 

Where the pronouncements of both parties fail is in making a false equivalence between some authors and all authors, and between some books and all books, and the result is that this is a lawsuit pitting apples against oranges. We saw in the lawsuits against Google that some academic authors, who may gain status based on their publications but very little if any income, did not see themselves as among those harmed by the book digitization project. Notably the authors in this current suit, as listed in the bibliography of pirated books in the appendix to the lawsuit, are ones whose works would be characterized best as “popular” and “commercial,” not academic: James Patterson, J. D. Salinger, Malcolm Gladwell, Toni Morrison, Laura Ingalls Wilder, and others. Not only do the living authors here earn above the poverty level, all of them provide significant revenue for the publishers themselves. And all of the books listed are in print and available in the marketplace. No mention is made of out-of-print books, no academic publishers seem to be involved. 

On the part of the Archive, they state that their digitized books fill an educational purpose, and that their collection includes books that are not available in digital format from publishers:

“ While Overdrive, Hoopla, and other streaming services provide patrons access to latest best sellers and popular titles,  the long tail of reading and research materials available deep within a library’s print collection are often not available through these large commercial services.  What this means is that when libraries face closures in times of crisis, patrons are left with access to only a fraction of the materials that the library holds in its collection.”[cdl-blog]

This is undoubtedly true for some of the digitized books, but the main thesis of the lawsuit points out that the Archive has digitized and is also lending current popular titles. The list of books included in the appendix of the lawsuit shows that there are in-copyright and most likely in-print books of a popular reading nature that have been part of the CDL. These titles are available in print and may also be available as ebooks from the publishers. Thus while the publishers are arguing that current, popular books should not be digitized and loaned (apples), the Archive is arguing that they are providing access to items not available elsewhere, and for educational purposes (oranges). 

The Law 

The suit states that publishers are not questioning copyright law, only violations of the law.

“For the avoidance of doubt, this lawsuit is not about the occasional transmission of a title under appropriately limited circumstances, nor about anything permissioned or in the public domain. On the contrary, it is about IA’s purposeful collection of truckloads of in-copyright books to scan, reproduce, and then distribute digital bootleg versions online.” ([Suit] Page 3).

This brings up a whole range of legal issues in regard to distributing digital copies of copyrighted works. There have been lengthy arguments about whether copyright law could permit first sale rights for digital items, and the answer has generally been no; some copyright holders have made the argument that since transfer of a digital file is necessarily the making of a copy there can be no first sale rights for those files. [1stSale] [ag1] Some ebook systems, such as the Kindle, have allowed time-limited person-to-person lending for some ebooks. This is governed by license terms between Amazon and the publishers, not by the first sale rights of the analog world. 

Section 108 of the copyright law does allow libraries and archives to make a limited number of copies The first point of section 108 states that libraries can make a single copy of a work as long as 1) it is not for commercial advantage, 2) the collection is open to the public and 3) the reproduction includes the copyright notice from the original. This sounds to be what the Archive is doing. However, the next two sections (b and c) provide limitations on that first section that appear to put the Archive in legal jeopardy: section “b” clarifies that copies may be made for preservation or security; section “c” states that the copies can be made if the original item is deteriorating and a replacement can no longer be purchased. Neither of these applies to the Archive’s lending. 

 In addition to its lending program, the Archive provides downloads of scanned books in DAISY format for those who are certified as visually impaired by the National Library Service for the Blind and Physically Handicapped in the US. This is covered in 121A of the copyright law, Title17, which allows the distribution of copyrighted works in accessible formats. This service could possibly be cited as a justification of the scanning of in-copyright works at the Archive, although without mitigating the complaints about lending those copies to others. This is a laudable service of the Archive if scans are usable by the visually impaired, but the DAISY-compatible files are based on the OCR’d text, which can be quite dirty. Without data on downloads under this program it is hard to know the extent to which this program benefits visually impaired readers. 

 Lending 

Most likely as part of the strategy of the lawsuit, very little mention is made of “lending.” Instead the suit uses terms like “download” and “distribution” which imply that the user of the Archive’s service is given a permanent copy of the book

“With just a few clicks, any Internet-connected user can download complete digital copies of in-copyright books from Defendant.” ([suit] Page 2). “... distributing the resulting illegal bootleg copies for free over the Internet to individuals worldwide.” ([suit] Page 14).
Publishers were reluctant to allow the creation of ebooks for many years until they saw that DRM would protect the digital copies. It then was another couple of years before they could feel confident about lending - and by lending I mean lending by libraries. It appears that Overdrive, the main library lending platform for ebooks, worked closely with publishers to gain their trust. The lawsuit questions whether the lending technology created by the Archive can be trusted.
“...Plaintiffs have legitimate fears regarding the security of their works both as stored by IA on its servers” ([suit] Page 47).

In essence, the suit accuses IA of a lack of transparency about its lending operation. Of course, any collaboration between IA and publishers around the technology is not possible because the two are entirely at odds and the publishers would reasonably not cooperate with folks they see as engaged in piracy of their property. 

Even if the Archive’s lending technology were proven to be secure, lending alone is not the issue: the Archive copied the publishers’ books without permission prior to lending. In other words, they were lending content that they neither owned (in digital form) nor had licensed for digital distribution. Libraries pay, and pay dearly, for the ebook lending service that they provide to their users. The restrictions on ebooks may seem to be a money-grab on the part of publishers, but from their point of view it is a revenue stream that CDL threatens. 

Is it About the Money?

“... IA rakes in money from its infringing services…” ([suit] Page 40). (Note: publishers earn, IA “rakes in”)
“Moreover, while Defendant promotes its non-profit status, it is in fact a highly commercial enterprise with millions of dollars of annual revenues, including financial schemes that provide funding for IA’s infringing activities. ([suit] Page 4).

These arguments directly address section (a)(1) of Title 17, section 108: “(1) the reproduction or distribution is made without any purpose of direct or indirect commercial advantage”. 

At various points in the suit there are references to the Archive’s income, both for its scanning services and donations, as well as an unveiled show of envy at the over $100 million that Brewster Kahle and his wife have in their jointly owned foundation. This is an attempt to show that the Archive derives “direct or indirect commercial advantage” from CDL. Non-profit organizations do indeed have income, otherwise they could not function, and “non-profit” does not mean a lack of a revenue stream, it means returning revenue to the organization instead of taking it as profit. The argument relating to income is weakened by the fact that the Archive is not charging for the books it lends. However, much depends on how the courts will interpret “indirect commercial advantage.” The suit argues that the Archive benefits generally from the scanned books because this enhances the Archive’s reputation which possibly results in more donations. There is a section in the suit relating to the “sponsor a book” program where someone can donate a specific amount to the Archive to digitize a book. How many of us have not gotten a solicitation from a non-profit that makes statements like: “$10 will feed a child for a day; $100 will buy seed for a farmer, etc.”? The attempt to correlate free use of materials with income may be hard to prove. 

Reading 

Decades ago, when the service Questia was just being launched (Questia ceased operation December 21, 2020), Questia sales people assured a group of us that their books were for “research, not reading.” Google used a similar argument to support its scanning operation, something like “search, not reading.” The court decision in Google’s case decided that Google’s scanning was fair use (and transformative) because the books were not available for reading, as Google was not presenting the full text of the book to its audience.[suit-g] 

The Archive has taken the opposite approach, a “books are for reading” view. Beginning with public domain books, many from the Google books project, and then with in-copyright books, the Archive has promoted reading. It developed its own in-browser reading software to facilitate reading of the books online. [reader] (*See note below)

Although the publishers sued Google for its scanning, they lost due to the “search, not reading” aspect of that project. The Archive has been very clear about its support of reading, which takes the Google justification off the table. 

“Moreover, IA’s massive book digitization business has no new purpose that is fundamentally different than that of the Publishers: both distribute entire books for reading.” ([suit] Page 5). 

 However, the Archive's statistics on loaned books shows that a large proportion of the books are used for 30 minutes or less. 

“Patrons may be using the checked-out book for fact checking or research, but we suspect a large number of people are browsing the book in a way similar to browsing library shelves.” [ia1] 

 In its article on the CDL, the Center for Democracy and Technology notes that “the majority of books borrowed through NEL were used for less than 30 minutes, suggesting that CDL’s primary use is for fact-checking and research, a purpose that courts deem favorable in a finding of fair use.” [cdt] The complication is that the same service seems to be used both for reading of entire books and as a place to browse or to check individual facts (the facts themselves cannot be copyrighted). These may involve different sets of books, once again making it difficult to characterize the entire set of digitized books under a single legal claim. 

The publishers claim that the Archive is competing with them using pirated versions of their own products. That leads us to the question of whether the Archive’s books, presented for reading, are effectively substitutes for those of the publishers. Although the Archive offers actual copies, those copies that are significantly inferior to the original. However, the question of quality did not change the judgment in the lawsuit against copying of texts by Kinko’s [kinkos], which produced mediocre photocopies from printed and bound publications. It seems unlikely that the quality differential will serve to absolve the Archive from copyright infringement even though the poor quality of some of the books interferes with their readability. 

Digital is Different

Publishers have found a way to monetize digital versions, in spite of some risks, by taking advantage of the ability to control digital files with technology and by licensing, not selling, those files to individuals and to libraries. It’s a “new product” that gets around First Sale because, as it is argued, every transfer of a digital file makes a copy, and it is the making of copies that is covered by copyright law. [1stSale] 

The upshot of this is that because a digital resource is licensed, not sold, the right to pass along, lend, or re-sell a copy (as per Title 17 section 109) does not apply even though technology solutions that would delete the sender’s copy as the file safely reaches the recipient are not only plausible but have been developed. [resale] 

“Like other copyright sectors that license education technology or entertainment software, publishers either license ebooks to consumers or sell them pursuant to special agreements or terms.” ([suit] Page 15)

“When an ebook customer obtains access to the title in a digital format, there are set terms that determine what the user can or cannot do with the underlying file.”([suit] Page 16)

This control goes beyond the copyright holder’s rights in law: DRM can exercise controls over the actual use of a file, limiting it to specific formats or devices, allowing or not allowing text-to-speech capabilities, even limiting copying to the clipboard.

Publishers and Libraries 

The suit claims that publishers and libraries have reached an agreement, an equilibrium.

“To Plaintiffs, libraries are not just customers but allies in a shared mission to make books available to those who have a desire to read, including, especially, those who lack the financial means to purchase their own copies.” ([suit] Page 17).
In the suit, publishers contrast the Archive’s operation with the relationship that publishers have with libraries. In contrast with the Archive’s lending program, libraries are the “good guys.”
“... the Publishers have established independent and distinct distribution models for ebooks, including a market for lending ebooks through libraries, which are governed by different terms and expectations than print books.”([suit] Page 6).
These “different terms” include charging much higher prices to libraries for ebooks, limiting the number of times an ebook can be loaned. [pricing1] [pricing2]
“Legitimate libraries, on the other hand, license ebooks from publishers for limited periods of time or a limited number of loans; or at much higher prices than the ebooks available for individual purchase.” [agol]
The equilibrium of which publishers speak looks less equal from the library side of the equation: library literature is replete with stories about the avarice of publishers in relation to library lending of ebooks. Some authors/publishers even speak out against library lending of ebooks, claiming that this cuts into sales. (This same argument has been made for physical books.)
“If, as Macmillan has determined, 45% of ebook reads are occurring through libraries and that percentage is only growing, it means that we are training readers to read ebooks for free through libraries instead of buying them. With author earnings down to new lows, we cannot tolerate ever-decreasing book sales that result in even lower author earnings.” [agliblend][ag42]

The ease of access to digital books has become a boon for book sales, and ebook sales are now rising while hard copy sales fall. This economic factor is a motivator for any of those engaged with the book market. The Archive’s CDL is a direct affront to the revenue stream that publishers have carved out for specific digital products. There are indications that the ease of borrowing of ebooks - not even needing to go to the physical library to borrow a book - is seen as a threat by publishers. This has already played out in other media, from music to movies. 

It would be hard to argue that access to the Archive’s digitized books is merely a substitute for library access. Many people do not have actual physical library access to the books that the Archive lends, especially those digitized from the collections of academic libraries. This is particularly true when you consider that the Archive’s materials are available to anyone in the world with access to the Internet. If you don’t have an economic interest in book sales, and especially if you are an educator or researcher, this expanded access could feel long overdue. 

We need numbers 

We really do not know much about the uses of the Archive’s book collection. The lawsuit cites some statistics of “views” to show that the infringement has taken place, but the page in question does not explain what is meant by a “view”. Archive pages for downloadable files of metadata records also report “views” which most likely reflect views of that web page, since there is nothing viewable other than the page itself. Open Library book pages give “currently reading” and “have read” stats, but these are tags that users can manually add to the page for the work. To compound things, the 127 books cited in the suit have been removed from the lending service (and are identified in the Archive as being in the collection “litigation works

Although numbers may not affect the legality of the controlled digital lending, the social impact of the Archive’s contribution to reading and research would be clearer if we had this information. Although the Archive has provided a small number of testimonials, a proof of use in educational settings would bolster the claims of social benefit which in turn could strengthen a fair use defense. 

Notes

(*) The NWU has a slide show [nwu2] that explains what it calls Controlled Digital Lending at the Archive. Unfortunately this document conflates the Archive's book Reader with CDL and therefore muddies the water. It muddies it because it does not distinguish between sending files to dedicated devices (which is what Kindle is) or dedicated software like what libraries use via software like Libby, and the Archive's use of a web-based reader. It is not beyond reason to suppose that the Archive's Reader software does not fully secure loaned items. The NWU claims that files are left in the browser cache that represent all book pages viewed: "There’s no attempt whatsoever to restrict how long any user retains these images". (I cannot reproduce this. In my minor experiments those files disappear at the end of the lending period, but this requires more concerted study.) However, this is not a fault of CDL but a fault of the Reader software. The reader is software that works within a browser window. In general, electronic files that require secure and limited use are not used within browsers, which are general purpose programs.

Conflating the Archive's Reader software with Controlled Digital Lending will only hinder understanding. Already CDL has multiple components:

  1. Digitization of in-copyright materials
  2. Lending of digital copies of in-copyright materials that are owned by the library in a 1-to-1 relation to physical copies

We can add #3, the leakage of page copies via the browser cache, but I maintain that poorly functioning software does not automatically moot points 1 and 2. I would prefer that we take each point on its own in order to get a clear idea of the issues.

The NWU slides also refer to the Archive's API which allows linking to individual pages within books. This is an interesting legal area because it may be determined to be fair use regardless of the legality of the underlying copy. This becomes yet another issue to be discussed by the legal teams, but it is separate from the question of controlled digital lending. Let's stay focused.

The International Federation of Library Associations has issued its own statement on Controlled Digital Lending at https://www.ifla.org/publications/node/93954

Citations

[1stSale] https://abovethelaw.com/2017/11/a-digital-take-on-the-first-sale-doctrine/ 

[ag1]https://www.authorsguild.org/industry-advocacy/reselling-a-digital-file-infringes-copyright/ 

[ag42] https://www.authorsguild.org/industry-advocacy/authors-guild-survey-shows-drastic-42-percent-decline-in-authors-earnings-in-last-decade/ 

[aghard] https://www.authorsguild.org/the-writing-life/why-is-it-so-goddamned-hard-to-make-a-living-as-a-writer-today/

[aglibend] https://www.authorsguild.org/industry-advocacy/macmillan-announces-new-library-lending-terms-for-ebooks/

[agol] https://www.authorsguild.org/industry-advocacy/update-open-library/ 

[cdl-blog] https://blog.archive.org/2020/03/09/controlled-digital-lending-and-open-libraries-helping-libraries-and-readers-in-times-of-crisis

[cdt] https://cdt.org/insights/up-next-controlled-digital-lendings-first-legal-battle-as-publishers-take-on-the-internet-archive/ 

[kinkos] https://law.justia.com/cases/federal/district-courts/FSupp/758/1522/1809457

[nel] http://blog.archive.org/national-emergency-library/

[nwu] "Appeal from the victims of Controlled Digital Lending (CDL)". (Retrieved 2021-01-10) 

[nwu2] "What is the Internet Archive doing with our books?" https://nwu.org/wp-content/uploads/2020/04/NWU-Internet-Archive-webinar-27APR2020.pdf

[pricing1] https://www.authorsguild.org/industry-advocacy/e-book-library-pricing-the-game-changes-again/ 

[pricing2] https://americanlibrariesmagazine.org/blogs/e-content/ebook-pricing-wars-publishers-perspective/ 

[reader] Bookreader 

[resale] https://www.hollywoodreporter.com/thr-esq/appeals-court-weighs-resale-digital-files-1168577 

[suit] https://www.courtlistener.com/recap/gov.uscourts.nysd.537900/gov.uscourts.nysd.537900.1.0.pdf 

[suit-g] https://cases.justia.com/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.pdf?ts=1445005805