Showing posts with label copyright. Show all posts
Showing posts with label copyright. Show all posts

Monday, March 01, 2021

Digitization Wars, Redux

 (NB: IANAL) 

 Because this is long, you can download it as a PDF here.

From 2004 to 2016 the book world (authors, publishers, libraries, and booksellers) was involved in the complex and legally fraught activities around Google’s book digitization project. Once known as “Google Book Search,” the company claimed that it was digitizing books to be able to provide search services across the print corpus, much as it provides search capabilities over texts and other media that are hosted throughout the Internet. 

Both the US Authors Guild and the Association of American Publishers sued Google (both separately and together) for violation of copyright. These suits took a number of turns including proposals for settlements that were arcane in their complexity and that ultimately failed. Finally, in 2016 the legal question was decided: digitizing to create an index is fair use as long as only minor portions of the original text are shown to users in the form of context-specific snippets. 

We now have another question about book digitization: can books be digitized for the purpose of substituting remote lending in the place of the lending of a physical copy? This has been referred to as “Controlled Digital Lending (CDL),” a term developed by the Internet Archive for its online book lending services. The Archive has considerable experience with both digitization and providing online access to materials in various formats, and its Open Library site has been providing digital downloads of out of copyright books for more than a decade. Controlled digital lending applies solely to works that are presumed to be in copyright. 

Controlled digital lending works like this: the Archive obtains and retains a physical copy of a book. The book is digitized and added to the Open Library catalog of works. Users can borrow the book for a limited time (2 weeks) after which the book “returns” to the Open Library. While the book is checked out to a user no other user can borrow that “copy.” The digital copy is linked one-to-one with a physical copy, so if more than one copy of the physical book is owned then there is one digital loan available for each physical copy. 

The Archive is not alone in experimenting with lending of digitized copies: some libraries have partnered with the Archive’s digitization and lending service to provide digital lending for library-owned materials. In the case of the Archive the physical books are not available for lending. Physical libraries that are experimenting with CDL face the added step of making sure that the physical book is removed from circulation while the digitized book is on loan, and reversing that on return of the digital book. 

Although CDL has an air of legality due to limiting lending to one user at a time, authors and publishers associations had raised objections to the practice. [nwu] However, in March of 2020 the Archive took a daring step that pushed their version of the CDL into litigation: using the closing of many physical libraries due to the COVID pandemic as its rationale, the Archive renamed its lending service the National Emergency Library [nel] and eliminated the one-to-one link between physical and digital copies. Ironically this meant that the Archive was then actually doing what the book industry had accused it of (either out of misunderstanding or as an exaggeration of the threat posed): it was making and lending digital copies beyond its physical holdings. The Archive stated that the National Emergency Library would last only until June of 2020, presumably because by then the COVID danger would have passed and libraries would have re-opened. In June the Archive’s book lending service returned to the one-to-one model. Also in June a suit was filed by four publishers (Hachette, HarperCollins, Penguin Random House, and Wiley) in the US District Court of the Southern District of New York. [suit] 

The Controlled Digital Lending, like the Google Books project, holds many interesting questions about the nature of “digital vs physical,” not only in a legal sense but in a sense of what it means to read and to be a reader today. The lawsuit not only does not further our understanding of this fascinating question; it sinks immediately into hyperbole, fear-mongering, and either mis-information or mis-direction. That is, admittedly, the nature of a lawsuit. What follows here is not that analysis but gives a few of the questions that are foremost in my mind.

 Apples and Oranges 

 Each of the players in this drama has admirable reasons for their actions. The publishers explain in their suit that they are acting in support of authors, in particular to protect the income of authors so that they may continue to write. The Authors’ Guild provides some data on author income, and by their estimate the average full-time author earns less than $20,000 per year, putting them at poverty level.[aghard] (If that average includes the earnings of highly paid best selling authors, then the actual earnings of many authors is quite a bit less than that.) 

The Internet Archive is motivated to provide democratic access to the content of books to anyone who needs or wants it. Even before the pandemic caused many libraries to close the collection housed at the Archive contained some works that are available only in a few research libraries. This is because many of the books were digitized during the Google Books project which digitized books from a small number of very large research libraries whose collections differ significantly from those of the public libraries available to most citizens. 

Where the pronouncements of both parties fail is in making a false equivalence between some authors and all authors, and between some books and all books, and the result is that this is a lawsuit pitting apples against oranges. We saw in the lawsuits against Google that some academic authors, who may gain status based on their publications but very little if any income, did not see themselves as among those harmed by the book digitization project. Notably the authors in this current suit, as listed in the bibliography of pirated books in the appendix to the lawsuit, are ones whose works would be characterized best as “popular” and “commercial,” not academic: James Patterson, J. D. Salinger, Malcolm Gladwell, Toni Morrison, Laura Ingalls Wilder, and others. Not only do the living authors here earn above the poverty level, all of them provide significant revenue for the publishers themselves. And all of the books listed are in print and available in the marketplace. No mention is made of out-of-print books, no academic publishers seem to be involved. 

On the part of the Archive, they state that their digitized books fill an educational purpose, and that their collection includes books that are not available in digital format from publishers:

“ While Overdrive, Hoopla, and other streaming services provide patrons access to latest best sellers and popular titles,  the long tail of reading and research materials available deep within a library’s print collection are often not available through these large commercial services.  What this means is that when libraries face closures in times of crisis, patrons are left with access to only a fraction of the materials that the library holds in its collection.”[cdl-blog]

This is undoubtedly true for some of the digitized books, but the main thesis of the lawsuit points out that the Archive has digitized and is also lending current popular titles. The list of books included in the appendix of the lawsuit shows that there are in-copyright and most likely in-print books of a popular reading nature that have been part of the CDL. These titles are available in print and may also be available as ebooks from the publishers. Thus while the publishers are arguing that current, popular books should not be digitized and loaned (apples), the Archive is arguing that they are providing access to items not available elsewhere, and for educational purposes (oranges). 

The Law 

The suit states that publishers are not questioning copyright law, only violations of the law.

“For the avoidance of doubt, this lawsuit is not about the occasional transmission of a title under appropriately limited circumstances, nor about anything permissioned or in the public domain. On the contrary, it is about IA’s purposeful collection of truckloads of in-copyright books to scan, reproduce, and then distribute digital bootleg versions online.” ([Suit] Page 3).

This brings up a whole range of legal issues in regard to distributing digital copies of copyrighted works. There have been lengthy arguments about whether copyright law could permit first sale rights for digital items, and the answer has generally been no; some copyright holders have made the argument that since transfer of a digital file is necessarily the making of a copy there can be no first sale rights for those files. [1stSale] [ag1] Some ebook systems, such as the Kindle, have allowed time-limited person-to-person lending for some ebooks. This is governed by license terms between Amazon and the publishers, not by the first sale rights of the analog world. 

Section 108 of the copyright law does allow libraries and archives to make a limited number of copies The first point of section 108 states that libraries can make a single copy of a work as long as 1) it is not for commercial advantage, 2) the collection is open to the public and 3) the reproduction includes the copyright notice from the original. This sounds to be what the Archive is doing. However, the next two sections (b and c) provide limitations on that first section that appear to put the Archive in legal jeopardy: section “b” clarifies that copies may be made for preservation or security; section “c” states that the copies can be made if the original item is deteriorating and a replacement can no longer be purchased. Neither of these applies to the Archive’s lending. 

 In addition to its lending program, the Archive provides downloads of scanned books in DAISY format for those who are certified as visually impaired by the National Library Service for the Blind and Physically Handicapped in the US. This is covered in 121A of the copyright law, Title17, which allows the distribution of copyrighted works in accessible formats. This service could possibly be cited as a justification of the scanning of in-copyright works at the Archive, although without mitigating the complaints about lending those copies to others. This is a laudable service of the Archive if scans are usable by the visually impaired, but the DAISY-compatible files are based on the OCR’d text, which can be quite dirty. Without data on downloads under this program it is hard to know the extent to which this program benefits visually impaired readers. 

 Lending 

Most likely as part of the strategy of the lawsuit, very little mention is made of “lending.” Instead the suit uses terms like “download” and “distribution” which imply that the user of the Archive’s service is given a permanent copy of the book

“With just a few clicks, any Internet-connected user can download complete digital copies of in-copyright books from Defendant.” ([suit] Page 2). “... distributing the resulting illegal bootleg copies for free over the Internet to individuals worldwide.” ([suit] Page 14).
Publishers were reluctant to allow the creation of ebooks for many years until they saw that DRM would protect the digital copies. It then was another couple of years before they could feel confident about lending - and by lending I mean lending by libraries. It appears that Overdrive, the main library lending platform for ebooks, worked closely with publishers to gain their trust. The lawsuit questions whether the lending technology created by the Archive can be trusted.
“...Plaintiffs have legitimate fears regarding the security of their works both as stored by IA on its servers” ([suit] Page 47).

In essence, the suit accuses IA of a lack of transparency about its lending operation. Of course, any collaboration between IA and publishers around the technology is not possible because the two are entirely at odds and the publishers would reasonably not cooperate with folks they see as engaged in piracy of their property. 

Even if the Archive’s lending technology were proven to be secure, lending alone is not the issue: the Archive copied the publishers’ books without permission prior to lending. In other words, they were lending content that they neither owned (in digital form) nor had licensed for digital distribution. Libraries pay, and pay dearly, for the ebook lending service that they provide to their users. The restrictions on ebooks may seem to be a money-grab on the part of publishers, but from their point of view it is a revenue stream that CDL threatens. 

Is it About the Money?

“... IA rakes in money from its infringing services…” ([suit] Page 40). (Note: publishers earn, IA “rakes in”)
“Moreover, while Defendant promotes its non-profit status, it is in fact a highly commercial enterprise with millions of dollars of annual revenues, including financial schemes that provide funding for IA’s infringing activities. ([suit] Page 4).

These arguments directly address section (a)(1) of Title 17, section 108: “(1) the reproduction or distribution is made without any purpose of direct or indirect commercial advantage”. 

At various points in the suit there are references to the Archive’s income, both for its scanning services and donations, as well as an unveiled show of envy at the over $100 million that Brewster Kahle and his wife have in their jointly owned foundation. This is an attempt to show that the Archive derives “direct or indirect commercial advantage” from CDL. Non-profit organizations do indeed have income, otherwise they could not function, and “non-profit” does not mean a lack of a revenue stream, it means returning revenue to the organization instead of taking it as profit. The argument relating to income is weakened by the fact that the Archive is not charging for the books it lends. However, much depends on how the courts will interpret “indirect commercial advantage.” The suit argues that the Archive benefits generally from the scanned books because this enhances the Archive’s reputation which possibly results in more donations. There is a section in the suit relating to the “sponsor a book” program where someone can donate a specific amount to the Archive to digitize a book. How many of us have not gotten a solicitation from a non-profit that makes statements like: “$10 will feed a child for a day; $100 will buy seed for a farmer, etc.”? The attempt to correlate free use of materials with income may be hard to prove. 

Reading 

Decades ago, when the service Questia was just being launched (Questia ceased operation December 21, 2020), Questia sales people assured a group of us that their books were for “research, not reading.” Google used a similar argument to support its scanning operation, something like “search, not reading.” The court decision in Google’s case decided that Google’s scanning was fair use (and transformative) because the books were not available for reading, as Google was not presenting the full text of the book to its audience.[suit-g] 

The Archive has taken the opposite approach, a “books are for reading” view. Beginning with public domain books, many from the Google books project, and then with in-copyright books, the Archive has promoted reading. It developed its own in-browser reading software to facilitate reading of the books online. [reader] (*See note below)

Although the publishers sued Google for its scanning, they lost due to the “search, not reading” aspect of that project. The Archive has been very clear about its support of reading, which takes the Google justification off the table. 

“Moreover, IA’s massive book digitization business has no new purpose that is fundamentally different than that of the Publishers: both distribute entire books for reading.” ([suit] Page 5). 

 However, the Archive's statistics on loaned books shows that a large proportion of the books are used for 30 minutes or less. 

“Patrons may be using the checked-out book for fact checking or research, but we suspect a large number of people are browsing the book in a way similar to browsing library shelves.” [ia1] 

 In its article on the CDL, the Center for Democracy and Technology notes that “the majority of books borrowed through NEL were used for less than 30 minutes, suggesting that CDL’s primary use is for fact-checking and research, a purpose that courts deem favorable in a finding of fair use.” [cdt] The complication is that the same service seems to be used both for reading of entire books and as a place to browse or to check individual facts (the facts themselves cannot be copyrighted). These may involve different sets of books, once again making it difficult to characterize the entire set of digitized books under a single legal claim. 

The publishers claim that the Archive is competing with them using pirated versions of their own products. That leads us to the question of whether the Archive’s books, presented for reading, are effectively substitutes for those of the publishers. Although the Archive offers actual copies, those copies that are significantly inferior to the original. However, the question of quality did not change the judgment in the lawsuit against copying of texts by Kinko’s [kinkos], which produced mediocre photocopies from printed and bound publications. It seems unlikely that the quality differential will serve to absolve the Archive from copyright infringement even though the poor quality of some of the books interferes with their readability. 

Digital is Different

Publishers have found a way to monetize digital versions, in spite of some risks, by taking advantage of the ability to control digital files with technology and by licensing, not selling, those files to individuals and to libraries. It’s a “new product” that gets around First Sale because, as it is argued, every transfer of a digital file makes a copy, and it is the making of copies that is covered by copyright law. [1stSale] 

The upshot of this is that because a digital resource is licensed, not sold, the right to pass along, lend, or re-sell a copy (as per Title 17 section 109) does not apply even though technology solutions that would delete the sender’s copy as the file safely reaches the recipient are not only plausible but have been developed. [resale] 

“Like other copyright sectors that license education technology or entertainment software, publishers either license ebooks to consumers or sell them pursuant to special agreements or terms.” ([suit] Page 15)

“When an ebook customer obtains access to the title in a digital format, there are set terms that determine what the user can or cannot do with the underlying file.”([suit] Page 16)

This control goes beyond the copyright holder’s rights in law: DRM can exercise controls over the actual use of a file, limiting it to specific formats or devices, allowing or not allowing text-to-speech capabilities, even limiting copying to the clipboard.

Publishers and Libraries 

The suit claims that publishers and libraries have reached an agreement, an equilibrium.

“To Plaintiffs, libraries are not just customers but allies in a shared mission to make books available to those who have a desire to read, including, especially, those who lack the financial means to purchase their own copies.” ([suit] Page 17).
In the suit, publishers contrast the Archive’s operation with the relationship that publishers have with libraries. In contrast with the Archive’s lending program, libraries are the “good guys.”
“... the Publishers have established independent and distinct distribution models for ebooks, including a market for lending ebooks through libraries, which are governed by different terms and expectations than print books.”([suit] Page 6).
These “different terms” include charging much higher prices to libraries for ebooks, limiting the number of times an ebook can be loaned. [pricing1] [pricing2]
“Legitimate libraries, on the other hand, license ebooks from publishers for limited periods of time or a limited number of loans; or at much higher prices than the ebooks available for individual purchase.” [agol]
The equilibrium of which publishers speak looks less equal from the library side of the equation: library literature is replete with stories about the avarice of publishers in relation to library lending of ebooks. Some authors/publishers even speak out against library lending of ebooks, claiming that this cuts into sales. (This same argument has been made for physical books.)
“If, as Macmillan has determined, 45% of ebook reads are occurring through libraries and that percentage is only growing, it means that we are training readers to read ebooks for free through libraries instead of buying them. With author earnings down to new lows, we cannot tolerate ever-decreasing book sales that result in even lower author earnings.” [agliblend][ag42]

The ease of access to digital books has become a boon for book sales, and ebook sales are now rising while hard copy sales fall. This economic factor is a motivator for any of those engaged with the book market. The Archive’s CDL is a direct affront to the revenue stream that publishers have carved out for specific digital products. There are indications that the ease of borrowing of ebooks - not even needing to go to the physical library to borrow a book - is seen as a threat by publishers. This has already played out in other media, from music to movies. 

It would be hard to argue that access to the Archive’s digitized books is merely a substitute for library access. Many people do not have actual physical library access to the books that the Archive lends, especially those digitized from the collections of academic libraries. This is particularly true when you consider that the Archive’s materials are available to anyone in the world with access to the Internet. If you don’t have an economic interest in book sales, and especially if you are an educator or researcher, this expanded access could feel long overdue. 

We need numbers 

We really do not know much about the uses of the Archive’s book collection. The lawsuit cites some statistics of “views” to show that the infringement has taken place, but the page in question does not explain what is meant by a “view”. Archive pages for downloadable files of metadata records also report “views” which most likely reflect views of that web page, since there is nothing viewable other than the page itself. Open Library book pages give “currently reading” and “have read” stats, but these are tags that users can manually add to the page for the work. To compound things, the 127 books cited in the suit have been removed from the lending service (and are identified in the Archive as being in the collection “litigation works

Although numbers may not affect the legality of the controlled digital lending, the social impact of the Archive’s contribution to reading and research would be clearer if we had this information. Although the Archive has provided a small number of testimonials, a proof of use in educational settings would bolster the claims of social benefit which in turn could strengthen a fair use defense. 

Notes

(*) The NWU has a slide show [nwu2] that explains what it calls Controlled Digital Lending at the Archive. Unfortunately this document conflates the Archive's book Reader with CDL and therefore muddies the water. It muddies it because it does not distinguish between sending files to dedicated devices (which is what Kindle is) or dedicated software like what libraries use via software like Libby, and the Archive's use of a web-based reader. It is not beyond reason to suppose that the Archive's Reader software does not fully secure loaned items. The NWU claims that files are left in the browser cache that represent all book pages viewed: "There’s no attempt whatsoever to restrict how long any user retains these images". (I cannot reproduce this. In my minor experiments those files disappear at the end of the lending period, but this requires more concerted study.) However, this is not a fault of CDL but a fault of the Reader software. The reader is software that works within a browser window. In general, electronic files that require secure and limited use are not used within browsers, which are general purpose programs.

Conflating the Archive's Reader software with Controlled Digital Lending will only hinder understanding. Already CDL has multiple components:

  1. Digitization of in-copyright materials
  2. Lending of digital copies of in-copyright materials that are owned by the library in a 1-to-1 relation to physical copies

We can add #3, the leakage of page copies via the browser cache, but I maintain that poorly functioning software does not automatically moot points 1 and 2. I would prefer that we take each point on its own in order to get a clear idea of the issues.

The NWU slides also refer to the Archive's API which allows linking to individual pages within books. This is an interesting legal area because it may be determined to be fair use regardless of the legality of the underlying copy. This becomes yet another issue to be discussed by the legal teams, but it is separate from the question of controlled digital lending. Let's stay focused.

The International Federation of Library Associations has issued its own statement on Controlled Digital Lending at https://www.ifla.org/publications/node/93954

Citations

[1stSale] https://abovethelaw.com/2017/11/a-digital-take-on-the-first-sale-doctrine/ 

[ag1]https://www.authorsguild.org/industry-advocacy/reselling-a-digital-file-infringes-copyright/ 

[ag42] https://www.authorsguild.org/industry-advocacy/authors-guild-survey-shows-drastic-42-percent-decline-in-authors-earnings-in-last-decade/ 

[aghard] https://www.authorsguild.org/the-writing-life/why-is-it-so-goddamned-hard-to-make-a-living-as-a-writer-today/

[aglibend] https://www.authorsguild.org/industry-advocacy/macmillan-announces-new-library-lending-terms-for-ebooks/

[agol] https://www.authorsguild.org/industry-advocacy/update-open-library/ 

[cdl-blog] https://blog.archive.org/2020/03/09/controlled-digital-lending-and-open-libraries-helping-libraries-and-readers-in-times-of-crisis

[cdt] https://cdt.org/insights/up-next-controlled-digital-lendings-first-legal-battle-as-publishers-take-on-the-internet-archive/ 

[kinkos] https://law.justia.com/cases/federal/district-courts/FSupp/758/1522/1809457

[nel] http://blog.archive.org/national-emergency-library/

[nwu] "Appeal from the victims of Controlled Digital Lending (CDL)". (Retrieved 2021-01-10) 

[nwu2] "What is the Internet Archive doing with our books?" https://nwu.org/wp-content/uploads/2020/04/NWU-Internet-Archive-webinar-27APR2020.pdf

[pricing1] https://www.authorsguild.org/industry-advocacy/e-book-library-pricing-the-game-changes-again/ 

[pricing2] https://americanlibrariesmagazine.org/blogs/e-content/ebook-pricing-wars-publishers-perspective/ 

[reader] Bookreader 

[resale] https://www.hollywoodreporter.com/thr-esq/appeals-court-weighs-resale-digital-files-1168577 

[suit] https://www.courtlistener.com/recap/gov.uscourts.nysd.537900/gov.uscourts.nysd.537900.1.0.pdf 

[suit-g] https://cases.justia.com/federal/appellate-courts/ca2/13-4829/13-4829-2015-10-16.pdf?ts=1445005805

Thursday, November 14, 2013

It's FAIR!

"In my view, Google Books provides significant public benefits. It advances the progress of the arts and sciences, while maintaining respectful consideration for the rights of authors and other creative individuals, and without adversely impacting the rights of copyright holders. It has become an invaluable research tool that permits students, teachers, librarians, and others to more efficiently identify and locate books. It has given scholars the ability, for the first time, to conduct full-text searches of tens of millions of books. It preserves books, in particular out-of-print and old books that have been forgotten in the bowels of libraries, and it gives them new life. It facilitates access to books for print-disabled and remote or underserved populations. It generates new audiences and creates new sources of income for authors and publishers. Indeed, all society benefits." p. 26
With that statement, Judge Denny Chin has ruled (PDF) that Google's digitization of books from libraries is a fair use.  And a very long saga ends.

Google was first brought to court in 2005 by the Author's Guild in a copyright infringement suit for its mass digitization of library holdings. Since then the matter has gone back to the court a number of times. Most significantly, Google, authors, and publishers developed two complex proposed settlements that were, however, so fraught with problems that the Department of Justice weighed in. Finally, the publishers bowed out and the original Author's Guild suit was revived. At that point, the question became: Is Google's digitization of books for the purposes of indexing (and showing snippets as search results) fair use?

Of course, much happened between 2005 and 2013. One important thing that happened was the development of HathiTrust, the digital repository where libraries can store the digital copies that they received from Google of their own books. The same Authors Guild sued HathiTrust for copyright infringement, but Judge Baer in that case decided for fair use.

I cannot over-emphasize either the role of libraries in this case nor the support that both judges expressed for libraries and for their promotion of "progress and the useful arts." Chin refers frequently to the amicus brief (PDF) presented by the American Library Association, as well as the conclusions in the HathiTrust case. Both judges clearly admire the mission of libraries, and it seems clear to me that the educational use of the materials by libraries was seen to offset the for-profit use by Google. In fact, Judge Chin reverses the roles of Google and the libraries when he says:
"Google provides the libraries with the technological means to make digital copies of books that they already own. The purpose of the library copies is to advance the libraries' lawful uses of the digitized books consistent with the copyright law." p. 26
In those terms, Google has simply helped libraries do what they do, better. Google's digitization of the library books is thus a public service.
"Google Books helps to preserve books and give them new life. Older books, many of which are out-of-print books that are falling apart buried in library stacks, are being scanned and saved." p. 12
Note that Google and the libraries (in HathiTrust) are exceedingly careful to stay within the letter of the law. Google's snippet display algorithm is rococo in design, making it literally impossible to reconstruct a book from the snippets it displays. So much so that it would probably take less time to re-scan the book at home on your page-at-a-time desktop scanner.

The full impact of this ruling is impossible (for me) to predict, but there are many among us who are breathing a great sigh of relief today. This opens the door for us to rethink digital scholarship based on materials produced before information was in digital form. 

I do have a wishlist, however, and at the top of that is for us to turn our attention to making the digitized texts even more useful by turning that uncorrected OCR into a more faithful reproduction of the original book. While large-scale linguistic studies may be valid in spite of a small percentage of errors, the use of the digitized materials for reading, in the case of those works in the public domain, and for listening, in the case of works made available to VIPs (visually impaired persons), is greatly hampered by the number and kinds of errors that result. In a future post I will give the results of a short study that I have done in that area.

See all my posts on Google Books

Sunday, September 22, 2013

Copyright, Metadata, and Attribution

The Berkeley Center for Law and Technology (BCLT) has done some interesting research on copyright, including a white paper that details the issues of performing "due diligence" in a determination of orphan works.

Recently I attended a small meeting co-sponsored by BCLT and the DPLA to begin a discussion of the issues around copyright in metadata, with a particular emphasis on bibliographic metadata. Much of the motivation for this is the uncertainty in the library and archival community about whether they can freely share their metadata. As long as this question remains un-answered, there are barriers to the free flow of data from and between cultural heritage institutions.

At the conclusion of the meeting it was clear that it will take some research to fully define the problem space. Fortunately for all of us, BCLT may be able to devote resources to undertake such a study, similar to what they have done around orphan works.

One of the first questions to undertake is whether bibliographic metadata is copyrightable in the first place. If not, then no further steps need to be taken -- not even putting a CC0 license on the data. In fact, some knowledgeable folks worry that using CC0 implies that there do exist intellectual property rights that must be addressed.

However, before you can attempt to determine if bibliographic metadata can be argued to be a set of facts which, under US copyright law, do not enjoy protection, you must be able to define "bibliographic metadata." During the meeting we did not attempt to create such a definition, but discussion ranged from "anything about a resource" to a specific set of descriptive elements. As there were representatives of archives in the room, we also talked about some of the implications of describing unpublished materials, which have a different legal standing but also provide less self-identification than resources that have been published. Drawing the line between fact and embellishment in bibliographic metadata is not going to be easy. Nor will the determination of level of creativity of the data, a necessary part of the analysis for US law. Note that other types of metadata were also discussed, such as rights metadata and preservation metadata, as well as a recognition that the exchange of metadata will of course cross national boundaries. Any study will have to determine where it will draw the "metadata" line, and also whether one can address the the question with an international scope.

Another complexity is that bibliographic data is already "crowd-sourced" in a sense. For any given bibliographic record,  different contributions have been made by different librarians from different institutions and at different times. This recognition makes it hard to ascribe intellectual ownership to any one party. And while library catalog data may be considered to be factual, it is much more than a simple rendering of facts, as the complexity of the cataloging rules attests. I likened library cataloging to a medical diagnosis: the end result (some scribbles in a file and perhaps a prescription given to the patient) does not reveal all of the knowledge and judgment that went into the decision. Metadata is the tip of an iceberg. That may not change its legal status, but I think that unless you have delved into the intricacies of cataloging it is hard to appreciate all that goes into the fairly simple display that users see on the screen.

The legal question is difficult, and to me it isn't entirely clear that solving the question on the legality of bibliographic data exchange will be sufficient to break the logjam. In a sense, projects like DPLA and Europeana, both of which have declared their metadata to be available with a CC0 license, might have more real impact than a determination based in law. Significant discussion at the meeting was about the need for attribution on the part of cultural heritage institutions. Like academics, the reputation and standing of such institutions depends on their getting recognition for their work. Releasing metadata (including thumbnails in the case of visual materials) needs to increase the visibility of those institutions, and to raise public awareness of the value of their collections. It is possible that solving the attribution problem could essentially dissolve the barriers to metadata sharing, since the gain to the institutions would be obvious.

Perhaps my one unique contribution to the group discussion was this:

We all know the © symbol and what it means. What we need now is an equally concise and recognizable symbol for attribution. Something like "(@)The Bancroft Library" or "(@)Dr. Seuss Collection". This would shorten attribution statements but also make them quickly recognizable, and a statement could also be a link to the appropriate web page. Standardizing attribution in this way should make adding attributions easier, and would demonstrate a culture of "giving credit where credit is due." The symbol needs to be simple, and should be easy to understand. It's time to comb through the Unicode charts for just the right one. Any suggestions?

See Also:


Unicode 1F6A9 - Triangular flag meaning "location"

Friday, August 09, 2013

Green paper on copyright, II


"To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries" (US Constitution)
 In 1993 the US government issued a green paper on copyright and the Internet. At the time the latter was being referred to as the "National Information Infrastructure" or NII. That green paper led to a white paper called: Intellectual Property and the National Information Infrastructure. The conclusions in this paper led to government efforts like the DMCA, as well as the as-yet unresolved questions about the library exceptions in section 108 of the copyright law.

Twenty years later we have another green paper on copyright and the Internet. It is, as are nearly all government documents, long and complex, and I hope to find time to do a comparison of the two papers to see if we have progressed in this area. But from a first reading I believe I can say that there is something that the two papers have in common: who they consider to be a creator and a rights holder.

It won't surprise you to learn that the emphasis in papers I and II is on the commercial content production communities: books, movies, music, newspapers.   As quoted in the press release,
“We see a digital future in which the relationship among digital technology, the Internet, and creative industries becomes increasingly symbiotic,” said Assistant Secretary of Commerce for Communications and Information and NTIA Administrator Lawrence E. Strickling. “In this digital future, the rights of creators and copyright owners are appropriately protected; creative industries continue to make their substantial contributions to the nation’s economic competitiveness; online service providers continue to expand the variety and quality of their offerings; technological innovation continues to thrive; and consumers have access to the broadest possible range of creative content.”
Oddly, he seems to be describing the Internet that I interact with today, but calls it the digital future. Nowhere, however, is there a mention of the rights of creators of Facebook pages, Google+ entries, Youtube videos, or tweets. A distinction is often made between creators and consumers, yet through social media and Internet publication methods (like blogging), that division is less valid than it was in the past.

Nor does the report admit that there is a the vast amount of content that is user-supplied and highly consumed, like Youtube, and that this content exceeds, both in publication and in consumption, the commercial offerings that the authors of the report are so concerned about. Hearings held to discuss the first green paper had industry representatives saying that the Internet would not be a "success" if commercial and entertainment content was not available, as if the Internet at the time (1994) were just an empty shell waiting for Time-Warner and Disney and Thomson-Reuters to come along and fill it up. Bruce Lehman, chairman of the committee producing the green and white papers, said to Congress in 1995:
 Creators, publishers and distributors of works will be wary of the electronic marketplace unless the law provides them the tools to protect their property against unauthorized use. Thus, the full potential of the NII will not be realized if the education, information and entertainment products protected by intellectual property laws are not protected effectively when disseminated via the NII.
In case you weren't there at the time, the Internet in 1994 was a thriving community with huge amounts of content. It was less flashy than today's Internet, since the technology did not yet allow for the efficient streaming of video and sound, and downloading a photograph could take a while. But it was not languishing for lack of content. This goes many-fold for the Internet today, but the representatives of the commercial content industries are somehow oblivious to any content that isn't making money for them. And that includes every Youtube upload, every Facebook page, bazillions of Flickr photographs, countless tweets.

Here are some stats from Youtube and Hulu, taken from their sites. Admittedly, Youtube is not 100% non-commercial and Hulu is not 100% pay-for-view, but this still shows that Youtube and its crowd-sourced content should be counted as real content in the sense of the green paper on copyright. But it isn't.

Hulu
 Number of Hulu video views in the past year     457 million
Total number of people who watched Hulu at least once in the past year     38 million
Number of Hulu video views in the past year     457 million
Average number of videos a Hulu watcher views     12
Average length of time a person spends on Hulu     1hr 13min
Number of devices in use that Hulu is available on     120 million
Percentage of people who use Hulu and only watch television shows     73 %
Percent of people who use Hulu to watch movies     9 %
Percent of videos viewed online that Hulu makes up     4 %
Percent growth for Hulu from 2010 to 2011     60 %
Total revenue made in 2011 by Hulu     $420 million
Youtube
More than 1 billion unique users visit YouTube each month
Over 6 billion hours of video are watched each month on YouTube—that's almost an hour for every person on Earth, and 50% more than last year
100 hours of video are uploaded to YouTube every minute
70% of YouTube traffic comes from outside the US
YouTube is localized in 56 countries and across 61 languages
According to Nielsen, YouTube reaches more US adults ages 18-34 than any cable network
So Youtube got 12 billion unique visitors in a year, to Hulu's 457 million. You'd think data like that would have an impact on the thinking of these folks. And the fact that Youtube reaches more viewers in the key demographic than cable TV should also shake up some media executives. In addition, those who upload content to Youtube are rights holders, as the Youtube terms of use makes clear:
For clarity, you retain all of your ownership rights in your Content. However, by submitting Content to YouTube, you hereby grant YouTube a worldwide, non-exclusive, royalty-free, sublicenseable and transferable license to use, reproduce, distribute, prepare derivative works of, display, and perform the Content in connection with the Service
You would also think that the report would address these Youtube content providers as creators, just as those who add content to Facebook pages are creators. But these creators are somehow not real creators in the minds of the writers of the report, and their needs are not addressed. No where does the report decry the exploitation, without compensation, of the work of these creators by industry giants like Facebook and Google. If someone else is making money off of your content it's bad, unless you aren't one of us, and therefore it's ok. The 99.99% of us who do not own the means of production and distribution must cede our rights in order to participate in the creation of content.

The report urges Congress to criminalize the unauthorized streaming of copyrighted works. Copyright law violations are subject to civil penalties, so this greatly ups the threat level for violations. This also revives one of the disputed and hated aspects of the famed "Stop Online Piracy Act," (SOPA) which was defeated through massive activism.

The upshot is that copyright will benefit those who create industrially, but not the millions who create individually. We should just admit that copyright law no longer has anything to do with creativity or social impact, and instead rename it to the "copyright industries law." As the 2013 green paper explains:
The industries that rely on copyright law are today an integral part of our economy, accounting for 5.1 million U.S. jobs in 2010—a figure that has grown dramatically over the past two decades. In that same year, these industries contributed 4.4 percent of U.S. GDP, or approximately $641 billion.
That's what this green paper, and the previous green paper, and all of the changes to copyright in the centuries since the time of the US Constitution, are really saying. Very little of what is protected today, and very little of what contributes to that $641 billion, is either Science or a useful Art. The founding fathers were writing at a time when science and technology defined progress, and progress defined prosperity. They did not, and could not have, anticipated the rise of leisure time and the media that would allow the creation of a huge economic center based on entertainment (including infotainment, which covers much news reporting today). Ironically, real Science is being encouraged to provide its content as Open Access. It's time to openly state that copyright law today is not what the founding fathers had in mind.

Update: Excellent blog post on the Green Paper and copyright by Kevin Smith.


Addendum: I just re-discovered the CPSR statement on the NII from 1994, and it contains this paragraph which is so sadly insightful:
An imaginative view of the risks of an NII designed without sufficient attention to public-interest needs can be found in the modern genre of dystopian fiction known as "cyberpunk." Cyberpunk novelists depict a world in which a handful of multinational corporations have seized control, not only of the physical world, but of the virtual world of cyberspace. The middle-class in these stories is sedated by a constant stream of mass-market entertainment that distracts them from the drudgery and powerlessness of their lives. It doesn't take a novelist's imagination to recognize the rapid concentration of power and the potential danger in the merging of major corporations in the computer, cable, television, publishing, radio, consumer electronics, film, and other industries. We would be distressed to see an NII shaped solely by the commercial needs of the entertainment, finance, home shopping, and advertising industries.

Friday, November 23, 2012

Fair Use(-ful)

The beauty and the aggravation of Fair Use in US copyright law is that one cannot pre-define particular uses as "fair." The countries that have, instead, the legal concept of "Fair Dealing" have an enumerated set of uses that are considered fair, although there is obviously still some need for interpretation. The advantage to Fair Use is that it can be re-interpreted with the times without the need for modification of the law. As new technologies come along, such as digitization of previously analog works, courts can make a decision based on the same four factors that have been used for earlier technologies. However, until such a decision is made in a court of law, it isn't possible to be sure whether a use is fair or not.

We have recently seen a court case that decided that HathiTrust's use of digitized books to provide an index to those books is fair. There is another court case that will decide a similar question regarding Google's digitization of books for its Google Book Search. Note, however, that even if both of these are determined to be fair use, each is a particular situation in a particular context. Both organizations have developed their services in an attempt to meet what they judged to be the letter of the law, and yet there is a considerable difference in the services they provide.

HathiTrust stores copies of digitized books from the collections of member libraries. In this case, HT is not itself doing the digitization but is storing files for books mostly digitized by Google. A search in the full text database of OCR'd page images returns, for in-copyright items, the page numbers on which the terms were found, and the number of hits found on each page. There are no snippets and no view of the text unless the text itself is deemed to be out of copyright.

Google has a different approach. To begin with, Google has performed mass digitization of books (estimated at about 20 million) without first obtaining permission from rights holders. So the Google case includes the act of digitization, whereas the HathiTrust case begins with digital files obtained from Google. Therefore the act of digitizing was not a factor in that case. In terms of use of the digitized works, Google also provides keyword searching of the OCR'd digital images, but takes a different approach to the results viewable by the searchers. Google provides short (about 3-5 lines) snippets that show the search terms in context on a page.
Google, however, places specific restrictions to avoid letting users "game" the search to gain access to enough of the text to substitute for actually acquiring access to the book. Here is how Google describes this in its recent legal response:
"The information that appears in Google Books does not substitute for reading the book. Google displays no more than three snippets from a book in response to a search query, even if the search term appears many times in the book. ... Google also prevents users from view a full page, or even several contiguous snippets, by displaying only one snippet per page in response to a given search and by 'blacking' (i.e. making unable for snippet view in response to any search) at least one snippet per page and one out of ten pages in a book." p.8
Google also exempts some types of books, like reference works, cookbooks, and poetry, from snippet display entirely.

The differences in the results returned by these two services reflect the differences in their contexts and their goals. HathiTrust has member institutions and their authorized users. The collection within HathiTrust reflects the holdings of the member institutions' libraries which means that the authorized users should have access, either in their library or through inter-library loan, to the physical book that was scanned. The HathiTrust full text is a search on the members' "stuff." The decision to give only page numbers makes some sense in this context, although providing snippets to scholars might have been acceptable to the judge. The return of page numbers and full word counts within pages reflects, IMO, the interest in quantitative analysis of term use. It also gives scholars some idea of the weight the term has within the text.

Google's situation is different. Google has no institutions, no members, no libraries; it provides its service to the general public (at least to the US public). There is no reason to assume that all of the members of that public will have access to the hard copy of any particular digitized book. Google seems to have decided that promoting its service as having primarily a marketing function, with the snippets as "teasers," would mollify the various intellectual property owners. In its brief of November 9, Google reiterates that it does not put advertising on the Google Book Search results pages, nor does Google make any money off of its referrals to book purchasing sites.

So here are two organizations that have bent over backwards to stay within what they deemed to be the boundaries of fair use, and they have done so in significantly different ways. This means that the fair use determination of each of these could have different outcomes, and each will provide different clues as to how fair use is viewed for digitized works.

It of course bears mentioning that both of these solutions provide hurdles for users. The HathiTrust user who is searching on a term that could have more than one meaning ("iron" "dive" "foot") does not have any context to help her understand if the results are relevant. The Google user, on the other hand, gets some context but cannot see all of the results and therefore does not know if there are key retrievals among those that have been blocked algorithmically. A use that is "fair" within copyright law may not seem "fair" to the user who is doing research. It makes you wonder if our idea of "fair use" couldn't be extended to be fair but also "useful."

Related posts
http://kcoyle.blogspot.com/2012/10/copyright-victories-part-ii.html

Tuesday, October 16, 2012

Copyright Victories, Part II

I did a short factual piece for InfoToday on the Authors Guild v. HathiTrust decision that was issued last week. The Authors Guild brought the suit against HathiTrust because HathiTrust is storing copies of books, digitized by Google, that are still under copyright. Fortunately for HathiTrust, its partners, and all of us in libraries, the judge decided:
  • The digitization of books for the purposes of providing a searchable index is transformative, and therefore is a Fair Use under copyright law.
  • The provision of these search capabilities “promotes the Progress of Science and useful Arts” and thus supports the goals of U.S. copyright policy and law.
  • The provision of in-copyright texts for visually impaired students and researchers is in direct support of the Americans With Disabilities Act.
The decision in the case of the Authors Guild v. Hathitrust echoes some of the same thinking as the GSU case, in particular on the educational and research use of intellectual property. This case hinged on the use of the digitized texts for indexing rather than for reading. The judge determined that the books in HathiTrust were not substitutes for the books on the library shelves, since they are not presented to users as texts to be read. The "transformation" of the readable texts to a searchable index that returns only page numbers and the number of times a term appears on the page results in a new product, not an imitation of the hard copy.

The judge decided this for HathiTrust, but this is the same question that is being asked in the Authors Guild lawsuit against Google. There are some obvious differences between the two situations, however. First, unlike HathiTrust Google is a for-profit company, so it loses points on the first factor of the fair use test:

(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
Because Google is digitizing works primarily from university libraries, both HathiTrust and Google do well on the second factor:

(2) the nature of the copyrighted work;
 Works of a creative nature (defined as "prose fiction, poetry, and drama") are given greater protection than works of fact. HathiTrust reports that only 9% of its digital collection meets the "creative" definition.

The third factor:

(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole

would seem to go against HathiTrust (and Google), but the judge looked at the two primary uses for the digital texts, keyword indexing and providing digital copies to members of the community with sight disabilities, and determined that they could not be done with anything less than a complete copy. If the argument of transformation is made in the case against Google, this factor should be the same.

Factor four is about the effect on the market:

(4) the effect of the use upon the potential market for or value of the copyrighted work

This is a bit tricky because presumably HathiTrust will point its users to the library-owned hard copies of the book, especially since many of the digitized books will be out of print and not available from publishers. Therefore there isn't much interaction with the market at all. The judge added that, if anything, the greater amount of discovery might lead to sales, but I wouldn't hold out much hope for that. The other use is to provide access to the blind; this is a non-market for print materials if there ever was one. Google, on the other hand, has partnered with publishers to sell digitized books as ebooks, and therefore the positive market force should be stronger in that case if Google can show that previously out-of-print books can be sold through its service.

Not mentioned anywhere that I can find is the question of digital "photographs" of pages vs. OCR'd text. The suit and the decision blend these together as "a digital copy." Having seen some of the results of Google's digitization I can say that the text resulting from the OCR can be quite lossy depending on the page layout (tables of contents in particular come out quite badly) and the quality of the original book. It is also the "transformation" part of the copying, since the photographs of the pages are simply copies of the page and are by their nature human-readable substitutes for the page itself. The judge seems to consider these "transitory" but in fact they are quite solidly real, and are stored in the HathiTrust repository. I suspect it is these pictures of the pages that the Authors Guild fears will be pirated should HathiTrust be hacked, less so the OCR'd pages which are unattractive plain text. However, HathiTrust was able to show the judge that it takes security quite seriously, and the Authors Guild was unable to demonstrate any quantifiable risk. 

What is heartening in this decision is the judge's enthusiasm for the role of libraries in further science and knowledge, and his great admiration of HathiTrust's service to scholars and to the blind. His decision is both factual and moral: he refers to the "invaluable contribution to the progress of science and the cultivation of the arts that at the same time effectuates the ideas espoused by the ADA." We could not have hoped for a better advocate for digital libraries than Judge Harold Baer Jr.

Copyright Victories, Part I

While we are awaiting the results of the long-standing Google Books digitization copyright suit, there have been some important copyright battles libraries have won. The first was the Georgia State University digital reserves. The second is the recent decision regarding HathiTrust. I'll cover GSU in this post, and HathiTrust in a subsequent one, since my comments are long.

Earlier this year the case of publishers vs. Georgia State University regarding their e-reserves program resulted in a win for GSU as 69 out of 74 copyright infringement claims were denied by the judge. The case was about the provision of course readings in digital form. The suit was brought by three academic publishers (Oxford, Cambridge and Sage) but was bankrolled by the Association of American Publishers and the Copyright Clearance Center. Most readings were individual book chapters, and these were digitized by the library. Students in the class were able to access these through a password-protected site. The judge decided that for all but a few of the works the use was a fair use based on the nature of the use (educational) and the amount being used (one chapter). Unfortunately the judge also decided to enforce a "bright line" test of not more than 10% of the work (and the amount used of the works in question averaged 10.1%). As we know, bright lines are not in the spirit of fair use, yet the judge clearly needed some way to make her decisions.

One of the more striking things in the GSU case was that the publishers were unable to prove their ownership of the copyrights for over a third of the original items. They had published the books, but many of them were essentially anthologies and they could not find the appropriate paperwork for all of the individual pieces. This lack of proof in essence revealed these particular documents to be orphan works, even though the publisher of the book itself is known and some of the books may have been in print. I suspect that if you were to require actual proof of rights ownership for books or journal articles that the number of orphans would grow considerably. This would be especially true for articles, at least based on my experience: journal publishers are rather casual about getting signed agreements, and I have often modified agreements through strike-outs which were never contested.

This is just more evidence that our copyright system is a huge mess. Proving ownership requires expensive research (the copyright office charges $165 per hour) and often does not solidly determine who holds the rights at this moment in time. Most of our action around intellectual property rights is based on claims and suppositions, not facts, and we often act as if there were evidence of held rights even though we have no such proof. In contrast, the patent system is fully documented with descriptions and drawings and references to other patents, although by its own admission the patent office has about a 3-year backlog, and filing and researching patents is time-consuming and expensive.

Another interesting aspect of the GSU case is that some of the works being copied were not covered by any available licensing scheme. This is especially interesting since the publisher plaintiffs were backed by the Copyright Clearance Center, presumably the agency that one would turn to when desiring to license a work for use. Licensing is a relatively big business: CCC earns over $200 million per year. The publishers included in this case each earn something shy of $500K per year in fees from CCC licensing. Much of that, however, comes from the commercial printing of course packs, not from direct educational institution use. The judge determined that the percentage of publisher revenue from electronic course content would be .00046 (five one-hundredths of one percent) of the average net revenue for any one of the publishers.

Reading this it becomes rather obvious that the move from traditional course packs, which are produced by commercial copy shops, to digital course readings, which are produced by the library or the professor, would mean a loss of revenue for the academically-oriented publishers. Course packs got slammed by the copyright holders not because copies were being made but because copies were being sold and all profit was going to the copy shops, none to the rights holders. In fact, in this lawsuit there were files on reserve that were never downloaded by students in the class, and the judge removed these titles from the suit because they were not read. This is an interesting answer to the question: "What if you make a copy and no one sees it?" Another way of wording this is: "Is it a copy if it has only been viewed by a computer?" With course packs, every student purchases every item in the course pack and you have no idea if any of those are read. With digital copies, every download can be counted. Although a download does not guarantee that the item has been read by the downloadee, it is a quantifiable use in the same way that the number of course packs printed is quantifiable. A file online does not seem, in and of itself, to be the same as a physical copy. This could have implications for library digitization projects, and relates to the decision in the Authors Guild v. HathiTrust case.

There are some gotchas to the use of a copyright licensing agency because of the inherent nature of the US fair use law. When one approaches CCC to license a work there is no fair use determination that is made as part of that request. It is up to the requestor to decide whether a license is needed or not. There are annual licenses available for educational institutions that cover a set of materials. None of these licenses are needed if the use is a fair use, and for educational institutions many uses are indeed fair uses, in particular classroom use. Therefore the CCC annual educational license may be paying for uses that do not require payment under copyright law. In essence, the license may be seen as a kind of insurance policy against infringement claims, but it may not be money well-spent. As the judge in the GSU case states (p. 66) "In the absence of judicial precedent concerning the limits of fair use for nonprofit educational uses, colleges and universities have been guessing about the permissible extent of fair use."

The decision itself runs to 350 pages, much of which is taken up with the decisions about the 74 documents in question. The judge does a very nice job of defining the nature of a work, and why individual chapters are viable on their own as part of a course syllabus. The decision that 10% of a "work" is permissible makes me wonder, however, if publishers won't see the light and begin digital publishing of individual chapters rather than creating book-length anthologies. 

More legal analysis is available on James Grimmelman's blog.

Friday, September 16, 2011

Due diligence do-over

In what I see as both a brave and an appropriate move, the University of Michigan admitted publicly that the Authors Guild had found some serious flaws in its process for identifying orphan works. The statement reaffirms the need to identify orphan works, and promises to revise its procedures.
"Having learned from our mistakes—we are, after all, an educational institution—we have already begun an examination of our procedures to identify the gaps that allowed volumes that are evidently not orphan works to be added to the list. Once we create a more robust, transparent, and fully documented process, we will proceed with the work, because we remain as certain as ever that our proposed uses of orphan works are lawful and important to the future of scholarship and the libraries that support it."
Among other things, what I find interesting in all this is that no one seems to be wondering why our copyright registration process is so broken that sometimes even the rights holders themselves don't know that they are the rights holders. It really shouldn't be this hard to find out if a work is covered by copyright. Larry Lessig covered this in his book "Free Culture," which is available online. The basic process of identifying copyrights is broken, and the burden is being placed on those who wish to make use of works. This is a clear anti-progress, pro-market bias in our copyright system.

Thursday, September 15, 2011

Diligence due

Oooof! Talk about making a BIG, public mistake.

HathiTrust's new Orphan Works Project proposed to do due diligence on works, then post them on the HT site for 90 days, after which those works would be assumed to be orphans and would then be made available (in full text) to members of the HT cooperating institutions. Sounds good, right? (Well, maybe other than the fact of posting the works on a site that few people even know about...)

The Authors Guild blog posted yesterday that it had found the rights holder of one of the books on HTs orphan works list in a little over 2 minutes using Google. (It's hard to believe that they didn't know this when the suit was filed on September 13 -- this is brilliant PR, if I ever saw it.) They then reported finding two others.

James Grimmelman, Associate Professor at New York Law School and someone considered expert on the Google Books case, has titled his blog post on this: "HathiTrust Single-Handedly Sinks Orphan Works Reform," stating that this incident will be brought up whenever anyone claims to have done due diligence on orphan works. I'm not quite as pessimistic as James, but I do believe this will be brought up in court and will work against HT.

Wednesday, September 14, 2011

Authors Guild in Perspective

In its suit against HathiTrust the three authors guilds claim that there are digitized copies of millions of copyrighted books in HathiTrust, and that these should be removed from the database and stored in escrow off-line.

A relevant question is: who do the authors guilds represent, and how many of those books belong to the represented authors?

The combined members of the three authors guilds is about 13,000. That seems like a significant number, but the Library of Congress name authority file has about 8 million names. That file also contains name/title combinations, and I don't have any statistics that tell me how many of those there are. (If anyone out there has a copy of the file and can run some stats on it, I'd greatly appreciate it.) Some of the names are for writers whose works are all in the public domain. Yet no matter how we slice it, the authors guilds of the lawsuit represent a small percentage of authors whose in-copyright works are in the HathiTrust database.

The legal question then is: does this lawsuit pertain to all in-copyright works in HathiTrust, or only those by the represented authors? Could I, for example, sue HathiTrust for violating Fay Weldon's copyright?


Reply to this from James Grimmelman on his blog:
Good question, Karen, and one I plan to address in more detail in a civil procedure post in the next few days. In brief, you couldn’t sue to enforce Fay Weldon’s copyright, as you aren’t an “owner” of any of the rights in it. The Authors Guild and other organizations can sue on behalf of their members, but the details of associational standing are complicated. There is also the question of the scope of a possible injunction (e.g. could Fay Weldon win as to one of her works and obtain an injunction covering others, or works by others), where there are also significant limits on how far the court can go. Again, more soon.

As I suspected, the legal issues are complex. Keep an eye on James' blog for more on this.

Monday, September 12, 2011

Authors Guild Sues HathiTrust

There has been a period of limbo since Judge Chin rejected the proposed settlement between the Author's Guild/Association of American publishers and Google. In fact, a supposedly final meeting between the parties is scheduled for this Thursday, 9/15, in the judge's court.

Monday, 9/12, the Author's Guild (and partners) filed suit against HathiTrust (and partners) for some of the same "crimes" of which it had accused Google: essentially making unauthorized copies of in-copyright texts. In addition, the recent announcement that the libraries would allow their users to access items that had been deemed to be orphan works figures in the suit. That this suit has come over 6 years since the original suit against Google is in itself interesting. Nearly all of the actions of HathiTrust and its member libraries fall within what would have been allowed if the agreement that came out of that suit had been approved by the court. Although we do not know the final outcome of that suit (and anxiously await Thursdays meeting to see if it is revelatory), this suit against the libraries is surely a sign that AG/AAP and Google have not come to a reconciliation.

The Suit

First, the suit establishes that the libraries received copies of Google-digitized items from Google, and have sent copies of these items to HathiTrust, which in turn makes some number of copies as part of its archival function. This is followed by a somewhat short exposition of the areas of copyright law that are pertinent, with an emphasis on section 108, which allows libraries to make limited copies to replace deteriorating works. The suit states that the copying being done is not in accord with section 108. Then it refers to the Orphan Works Project that several libraries are partnering in, and the plan on the part of the libraries to make the full text of orphan works available to institutional users.

Since most of these institutions (if not all of them) are state institutions that have protections against paying out large sums in a lawsuit of this nature, the goal is to regain the control of the works by forcing HathiTrust (and the named libraries) to transfer their digital copies of in-copyright works to a "commercial grade" escrow agency with the files held off network "pending an appropriate act of Congress."

As James Grimmelman comments in his blog post on the suit, there's a lot of mixing up between the orphan works and owned works in the suit. He points out that a group of organizations representing authors could hardly make a case for orphan works since, by definition, the lack of ownership of the orphans means they can't be represented by a guild of people defending their own works.

The Problems

There are numerous problems that I see in the text of this suit. (IANAL, just a Librarian.)
  • The suit mentions large numbers of books that have been copied without permission, but makes no attempt to state how many of those books belong to the members of the plaintiff organizations.
  • The suit throws around large numbers without clearly stating that none of the statements include Public Domain works. It isn't clear, therefore, what the numbers represent: the entire holdings of HathiTrust, or just the in-copyright holdings. Also, in relation to the latter, unless one has done a considerable amount of work there are many works that are post-1923 that are also in the Public Domain. Cutting off at that year does not account for works that were not renewed, or were never copyrighted. I also doubt if anyone has a clear idea how many of the works in question are Public Domain because they are US Federal documents. This imprecision on the copyright status of works is very frustrating, but HathiTrust is not to blame for this state of affairs.
  • Some of their claims do not seem to me to be within legal bounds. For example, in one section they claim that although HathiTrust is not giving users access to in-copyright works, they potentially could. Where does that fit in?
  • They also claim that there is a risk of unauthorized access. However, the security at HathiTrust meets the security standards that the Author's Guild agreed to in the (unapproved) settlement with Google. If it was good enough then, why is it now too risky?
  • They claim that the libraries themselves have been digitizing in-copyright books. I wasn't aware of this, and would like to know if this is the case.
  • They state that the libraries said that before Google it was costing them $100 a book for digitization. Then the plaintiffs say that this means that the value of the digital files is in the hundreds of millions of dollars. First, I have heard figures that are more like $30 a book. Second, I don't see how the cost to digitize can translate into a value that is relevant to the complaint.
  • Although the legislature has failed to pass an orphan works law that would allow the use of these materials and still benefit owners if they do come forth, it seems like a poor strategy to complain about a well-designed program of due diligence and notification, which is what the libraries have designed. Orphan works are the least available works: if you have an owner you can ask permission; if there is no owner you cannot ask permission and therefore there is no way to use the work if your use falls outside of fair use. It's hard to argue for taking these works entirely out of the cultural realm simply because we have a poorly managed copyright ownership record.
  • There are a few odd sections where they make reference to bibliographic data as though that were part of the "unauthorized digitization" rather than data that was created by and belongs to the libraries. There's an odd attempt to make bibliographic data searching seem nefarious.
Parties
Plaintiffs: The Author's Guild, Inc.; The Australian Society of Authors limited ; Union des Erivaines Quebecois; Pat Cummings; Angelo Loukakis; Roxana Robinson; Andre Roy; James Shapiro; Daniele Simpson; T.J. Stiles; and Fay Weldon. (Links are to some sample HathiTrust records.)

Defendants: HathiTrust; The Regents of the University of California, The Board of Regents of the University of Wisconsin System; The Trustees of Indiana University; and Cornell University.

Links
Boing Boing: Authors Guild declares war on university effort to rescue orphaned books
Library Journal: Copyright Clash

Saturday, January 24, 2009

Obama administration embraces Creative Commons

The copyright notice on the Obama Whitehouse site states:

Pursuant to federal law, government-produced materials appearing on this site are not copyright protected. The United States Government may receive and hold copyrights transferred to it by assignment, bequest, or otherwise.

Except where otherwise noted, third-party content on this site is licensed under a Creative Commons Attribution 3.0 License. Visitors to this website agree to grant a non-exclusive, irrevocable, royalty-free license to the rest of the world for their submissions to Whitehouse.gov under the Creative Commons Attribution 3.0 License.
Wow, that feels good!

Thursday, December 04, 2008

Google and Fair Use

There's some background to the Google/AAP settlement that I believe is key to understanding the subtext around it. This won't be news to most folks, but I thought it would be good to re-articulate it in the context of the settlement, lest we forget.

Google's first business is that of indexing resources that are on the web. I'll talk about them as if they were all texts because it's easier, but the same thing could be said for images and other resources.

To do the indexing, Google must make a copy of the web page or document. Using this copy, it adds the page to its search engine. As a good citizen, Google pays attention to the robots.txt file, and does not index pages where the site owner has opted out of being included in search engines.

This is all fine and unremarkable until you look at it from the point of view of copyright law. Copyright is specifically about... making copies, and it gives the right to make copies, or to authorize the making of copies, to the copyright holder. That can be the author, or someone to whom the author has passed along the right. Copyright holders must opt in to the making copies: they have to give permission. The default in copyright law is that copies cannot be made unless the copyright holder gives approval.

So the big question is: Is Google violating copyright law by making copies of web pages without the permission of the copyright holders? There are two main ways of looking at this:
  1. The web is different from the print environment. Anyone who has put their works out on the web has agreed to copying because no one can even view the work without making a copy. If they don't want people copying, they need to hide their works behind a security screen. However, there is no such exception or wording in copyright law that would support this.
  2. The web is not different from the print environment. But Google is just producing an index and there is nothing in copyright law that would prevent someone from producing an index of words in texts. The incidental copies that Google makes in order to produce the index are allowed under the Fair Use aspects of the copyright law.
So then we move on to the Google Books project. Initially, Google claimed that it was doing the same thing with books as it does with the web: making incidental copies in order to create keyword indexes to the texts. In terms of copyright law, argument #1 is pretty much out because these works can be read without making a copy, so the copyright holders haven't agreed to let their works be copied. This leaves us with argument #2: it must be fair use.

In fact, Google did and does make the fair use argument. The libraries that partnered with Google also came to the fair use conclusion in at least some cases. The CIC project FAQ says:

University of Michigan said this in 2007:

Does this project comply with copyright law?

Yes. This project was undertaken with careful attention to the law and to the rights and responsibilities of the various parties involved. The purpose of copyright law is to promote progress in society. We are confident that the Books Library project is fully consistent with the fair use doctrine under U.S. copyright law and the principles underlying copyright law itself. Copyright law strikes a balance between rewarding creators of intellectual property for their creations and facilitating public access to these works in ways that do not create a business harm. For books, this means ensuring authors write books, publishers sell them and libraries lend them. By making books more discoverable, Google is enhancing the ability of authors and publishers to sell books to an audience beyond the traditional book market.

What was at stake with the AAP lawsuit was exactly this decision about Fair Use. If copying the books for the purpose of indexing were determined to not be fair use, then this decision could bleed over into the web. And of course it would mean the end of Google Book Search (which has now become Google Book Store). Although Google has always provided a confident posture to the public, declaring unwaveringly that what it does as a search engine is perfectly within copyright law, the idea of going to court over the issue would have put their entire operation at risk.

Now back to libraries. Fair use is not a list of things you can do but a judgment call relating to some complex factors. Some key factors have to do with whether your use is commercial in nature or could compete with the exploitation of works by the copyright holders. There are, in addition, exceptions in the copyright law relating to research and study, and special exceptions for libraries. In fact, in relation to copyright law, libraries and educational institutions get considerably more latitude in using works than do commercial enterprises. As an example, a teacher can make copies of an article for her students as part of a lesson, and that is generally considered fair use. A company manager who wants his staff to read an article cannot rely on fair use for copying, but must apply to the copyright holder (usually through an intermediary such as CCC) and pay a fee. (See the Texaco case.)

What happened with Google Book Search and the AAP is that the digitization of the libraries' books and subsequent use of those was judged not by the criteria that would be used normally for libraries, of course, but by the criteria that would be used for a commercial entity. That's totally logical, since although Google was partnered with the libraries, the primary use of the materials was to fuel Google Book Search, an obviously for-profit activity.

Libraries have gotten the short end of the stick because their use of their own materials became commercialized through their partnership with Google. If instead libraries had managed to digitize the books on their own, the outcome would have likely have been entirely different (if any lawsuit had been brought, which might not have happened). I believe that libraries could be found to have a fair use case for digitizing their works for the purposes of searching, and could be allowed to use those digitized copies for the exceptions spelled out in section 108 of the copyright law (such as providing access to the sight impaired, or for replacement of deteriorated originals). Unfortunately, the concept of digitization of the contents of libraries has now been tainted with the air of commercialization and has earned the wrath of the publishers and authors. The Google/AAP settlement has created a mechanism that ignores the inherent rights of the libraries, but also makes it more difficult for them to justify undertaking their own digitization project.

This is why I disagree heartily when I hear statements like:

We're delighted that this agreement creates new opportunities for libraries and universities to offer their patrons and students access to millions of books beyond their own collections. (from Google)
The settlement might look good from the point of view of a commercial entity facing copyright law, but it binds the non-profit educational and cultural heritage community to legal decisions designed for the for-profit sector. This is not only not a win for libraries, but it will hinder libraries in their efforts to make use of current technologies to further the arts and sciences.

Saturday, November 22, 2008

More on Google/AAP

Here are some more bits and thoughts on the agreement between Google and the AAP.

Library Involvement

Some librarians were involved in the settlement talks. The only one I have found so far who has come out about this is Georgia Harper. The librarians were working under a non-disclosure agreement (NDA), and therefore will not be able to reveal any details of the discussions. I have heard statements from others who I believe were privy to the negotiations, and they all seem to feel that the outcome was better for libraries due to the involvement of members of our "class." (Note that Google and AAP had high-end lawyers arguing their side, and we had hard-working librarians. I don't know how many of "our" representatives were also lawyers, but you can just imagine how greatly out-gunned they were.) Unfortunately that doesn't change my mind about the bait and switch move.

Google Books as Library

Some have begun to refer to Google Books as a library. We have to do some serious thinking about what the Google Book database really is. To begin with, it's not a research collection, at least not at this point. It's really a somewhat odd, almost random bunch of book "stuff." As you know, neither Google nor the libraries are selecting particular books for digitization. This is a "mass digitization" project that starts at one end of a library and plows through blindly to the other end. Some libraries have limited Google to public domain works, so in terms of any area of study there is an artificial cut-off of knowledge. Not to mention that some libraries, mainly the University of California, have been working with Google primarily to digitize books in their two storage facilities; that is, they have been digitizing the low use books that were stored remotely.

So the main reason why Google Books is not a library is that it isn't what we would call a "collection." The books have not been chosen to support a particular discipline or research area. Yet it will become a de facto collection because people will begin using it for research. Thus "all human knowledge" becomes something more like the elephant and the blind man: research in online resources and research that uses print materials will get very different views of human knowledge. (This is not a new phenomenon. I wrote about this in terms of some early digital projects I was involved in.) One of the big gaps in Google Books will be current materials, those that are still in print. Google will need to convince the publishers that it can increase their revenue stream for current books in order to get them to participate.

Subscribing to Google Books: Just Say No?


Beyond the (undoubtedly hard-won by library representatives) single terminal access in each public library in the US, libraries will be asked to subscribe to the Google Book service in order to give their users access to the text of the books (not just the search capability). This is one of the more painful aspects of the agreement because it seems to ignore the public costs that went in to the purchase, organization, and storage of those works by libraries. (I'm not includng privately funded libraries here, but many of the participants are publicly funded.) The parallels with the OCLC mess are ironic: libraries paying for access to their own materials. So, couldn't the libraries just refuse to subscribe? Not really. Publicly funded libraries have a mission to provide access to the world's intellectual output in a way that best serves their users. When something new comes along -- films on DVD, music on CD, the Internet -- libraries must do what they can to make sure that their users are not informationally underpriviledged. Google now has the largest body of digitized full text, and there will be a kind of "information arms race" as institutions work to make sure that their users can compete using these new resources.

The (Somewhat Hidden) Carrot

I can't imagine that anyone thought that libraries and Google were digitizing books primarily so that people could read what are essentially photographs of book pages on a computer screen. Google initially stated that they were only interested in searching the full text of books. While interesting in itself, keyword searching of rather poor OCR text is not a killer app. What we gain by having a large number of digitized books is a large corpus on which we can do computational research. We can experiment with ideas like: can we follow the flow of knowledge through these texts? Can we create topic maps of fields of study? Can we identify the seminal works in some area? The ability to do this research is included in the agreement (section 7.2(d), The Research Corpus). There will be two copies of this corpus allowed under the agreement, although I don't see any detail as to what the "corpus" will consist of. Will it just be a huge file of digitized books and OCR? Will it be a set of services?

I have suspected for a while that Google was already doing research on the digital files that it holds. It only makes sense. For academics in areas like statistics, computer science, and linguistics, this corpus opens up a whole range of possibilities for research; and research means grants, and grants mean jobs (or tenure, as the case may be). This will be a strong motivation for institutions to want to participate in the Google Book product. Research will NOT be limited to participants; others can request access. What I haven't yet found is anything relating to pricing for the use of the research collection, nor if being a participating library grants less expensive access for your institution. If the latter is the case, then one motivation for libraries to agree to allow Google to scan their books (at some continuing cost to the library) will be that it favors the institution's researchers in this new and exciting area. Full participant libraries (the ones that get to keep the digital copies of their works) can treat their own corpus as research fodder. The other costs of being a full participant are such that I'll still be surprised if any libraries go that route, but if they do I think that this "hidden carrot" will be a big part of it.

----

There's lots of good blogging going on out there on this topic. It needs a cumulative page to help people find the posts. Please tell me you have time to work on that, so I don't have to take it on! (Or that it exists already and I've missed it.) (The PureInformation Blog has a good list.)

Note: the Internet Archive/OCA may take this on. I'll post if/when they do.

Previous posts: