Showing posts with label googlebooks. Show all posts
Showing posts with label googlebooks. Show all posts

Tuesday, October 10, 2017

Google Books and Mein Kampf

I hadn't look at Google Books in a while, or at least not carefully, so I was surprised to find that Google had added blurbs to most of the books. Even more surprising (although perhaps I should say "troubling") is that no source is given for the book blurbs. Some at least come from publisher sites, which means that they are promotional in nature. For example, here's a mildly promotional text about a literary work, from a literary publisher:



This gives a synopsis of the book, starting with:

"Throughout a single day in 1892, John Shawnessy recalls the great moments of his life..." 

It ends by letting the reader know that this was a bestseller when published in 1948, and calls it a "powerful novel."

The blurb on a 1909 version of Darwin's The Origin of Species is mysterious because the book isn't a recent publication with an online site providing the text. I do not know where this description comes from, but because the  entire thrust of this blurb is about the controversy of evolution versus the Bible (even though Darwin did not press this point himself) I'm guessing that the blurb post-dates this particular publication.


"First published in 1859, this landmark book on evolutionary biology was not the first to deal with the subject, but it went on to become a sensation -- and a controversial one for many religious people who could not reconcile Darwin's science with their faith."
That's a reasonable view to take of Darwin's "landmark" book but it isn't what I would consider to be faithful to the full import of this tome.

The blurb on Hitler's Mein Kampf is particularly troubling. If you look at different versions of the book you get both pro- and anti- Nazi sentiments, neither of which really belong  on a site that claims to be a catalog of books. Also note that because each book entry has only one blurb, the tone changes considerably depending on which publication you happen to pick from the list.


First on the list:
"Settling Accounts became Mein Kampf, an unparalleled example of muddled economics and history, appalling bigotry, and an intense self-glorification of Adolf Hitler as the true founder and builder of the National Socialist movement. It was written in hate and it contained a blueprint for violent bloodshed."

Second on the list:
"This book has set a path toward a much higher understanding of the self and of our magnificent destiny as living beings part of this Race on our planet. It shows us that we must not look at nature in terms of good or bad, but in an unfiltered manner. It describes what we must do if we want to survive as a people and as a Race."
That's horrifying. Note that both books are self-published, and the blurbs are the ones that I find on those books in Amazon, perhaps indicating that Google is sucking up books from the Amazon site. There is, or at least at one point there once was, a difference between Amazon and Google Books. Google, after all, scanned books in libraries and presented itself as a search engine for published texts; Amazon will sell you Trump's tweets on toilet paper. The only text on the Google Books page still claims that Google Books is about  search: "Search the world's most comprehensive index of full-text books." Libraries partnered with Google with lofty promises of gains in scholarship:
"Our participation in the Google Books Library Project will add significantly to the extensive digital resources the Libraries already deliver. It will enable the Libraries to make available more significant portions of its extraordinary archival and special collections to scholars and researchers worldwide in ways that will ultimately change the nature of scholarship." Jim Neal, Columbia University
I don't know how these folks now feel about having their texts intermingled with publications they would never buy and described by texts that may come from shady and unreliable sources.

Even leaving aside the grossest aspects of the blurbs and Google's hypocrisy about its commercialization of its books project, adding blurbs to the book entries with no attribution and clearly not vetting the sources is extremely irresponsible. It's also very Google to create sloppy algorithms that illustrate their basic ignorance of the content their are working with -- in this case, the world's books.

Monday, November 28, 2016

All the Books

I just joined the Book of the Month Club. This is a throwback to my childhood, because my parents were members when I was young, and I still have some of the books they received through the club. I joined because my reading habits are narrowing, and I need someone to recommend books to me. And that brings me to "All the Books."

"All the Books" is a writing project I've had on my computer and in notes ever since Google announced that it was digitizing all the books in the world. (It did not do this.) The project was lauded in an article by Kevin Kelley in the New York Times Magazine of May 14, 2006, which he prefaced with:

"What will happen to books? Reader, take heart! Publisher, be very, very afraid. Internet search engines will set them free. A manifesto."

There are a number of things to say about All the Books. First, one would need to define "All" and "Books". (We can probably take "the" as it is.) The Google scanning projects defined this as "all the bound volumes on the shelves of certain libraries, unless they had physical problems that prevented scanning." This of course defines neither "All" nor "Books".

Next, one would need to gather the use cases for this digital corpus. Through the HathiTrust project we know that a small number of scholars are using the digital files for research into language usage over time. Others are using the the files to search for specific words or names, discovering new sources of information about possibly obscure topics. As far as I can tell, no one is using these files to read books. The Open Library, on the other hand, is lending digitized books as ebooks for reading. This brings us to the statement that was made by a Questia sales person many years ago, when there were no ebooks and screens were those flickery CRTs: "Our books are for research, not reading." Given that their audience was undergraduate students trying to finish a paper by 9:30 a.m. the next morning, this was an actual use case with actual users. But the fact that one does research in texts one does not read is, of course, not ideal from a knowledge acquisition point of view.

My biggest beef with "All the Books" is that it treats them as an undifferentiated mass, as if all the books are equal. I always come back to the fact that if you read one book every week for 60 years (which is a good pace) you will have read 3,120. Up that to two books a week and you've covered 6,240 of the estimated 200-300 million books represented in WorldCat. The problem isn't that we don't have enough books to read; the problem is finding the 3-6,000 books that will give us the knowledge we need to face life, and be a source of pleasure while we do so. "All the Books" ignores the heights of knowledge, of culture, and of art that can be found in some of the books. Like Sarah Palin's response to the question "Which newspapers form your world view?", "all of them" is inherently an anti-intellectual answer, either by someone who doesn't read any of them, or who isn't able to distinguish the differences.

"All the Books" is a complex concept. It includes religious identity; the effect of printing on book dissemination; the loss of Latin as a universal language for scholars; the rise of non-textual media. I hope to hunker down and write this piece, but meanwhile, this is a taste.

Thursday, November 14, 2013

It's FAIR!

"In my view, Google Books provides significant public benefits. It advances the progress of the arts and sciences, while maintaining respectful consideration for the rights of authors and other creative individuals, and without adversely impacting the rights of copyright holders. It has become an invaluable research tool that permits students, teachers, librarians, and others to more efficiently identify and locate books. It has given scholars the ability, for the first time, to conduct full-text searches of tens of millions of books. It preserves books, in particular out-of-print and old books that have been forgotten in the bowels of libraries, and it gives them new life. It facilitates access to books for print-disabled and remote or underserved populations. It generates new audiences and creates new sources of income for authors and publishers. Indeed, all society benefits." p. 26
With that statement, Judge Denny Chin has ruled (PDF) that Google's digitization of books from libraries is a fair use.  And a very long saga ends.

Google was first brought to court in 2005 by the Author's Guild in a copyright infringement suit for its mass digitization of library holdings. Since then the matter has gone back to the court a number of times. Most significantly, Google, authors, and publishers developed two complex proposed settlements that were, however, so fraught with problems that the Department of Justice weighed in. Finally, the publishers bowed out and the original Author's Guild suit was revived. At that point, the question became: Is Google's digitization of books for the purposes of indexing (and showing snippets as search results) fair use?

Of course, much happened between 2005 and 2013. One important thing that happened was the development of HathiTrust, the digital repository where libraries can store the digital copies that they received from Google of their own books. The same Authors Guild sued HathiTrust for copyright infringement, but Judge Baer in that case decided for fair use.

I cannot over-emphasize either the role of libraries in this case nor the support that both judges expressed for libraries and for their promotion of "progress and the useful arts." Chin refers frequently to the amicus brief (PDF) presented by the American Library Association, as well as the conclusions in the HathiTrust case. Both judges clearly admire the mission of libraries, and it seems clear to me that the educational use of the materials by libraries was seen to offset the for-profit use by Google. In fact, Judge Chin reverses the roles of Google and the libraries when he says:
"Google provides the libraries with the technological means to make digital copies of books that they already own. The purpose of the library copies is to advance the libraries' lawful uses of the digitized books consistent with the copyright law." p. 26
In those terms, Google has simply helped libraries do what they do, better. Google's digitization of the library books is thus a public service.
"Google Books helps to preserve books and give them new life. Older books, many of which are out-of-print books that are falling apart buried in library stacks, are being scanned and saved." p. 12
Note that Google and the libraries (in HathiTrust) are exceedingly careful to stay within the letter of the law. Google's snippet display algorithm is rococo in design, making it literally impossible to reconstruct a book from the snippets it displays. So much so that it would probably take less time to re-scan the book at home on your page-at-a-time desktop scanner.

The full impact of this ruling is impossible (for me) to predict, but there are many among us who are breathing a great sigh of relief today. This opens the door for us to rethink digital scholarship based on materials produced before information was in digital form. 

I do have a wishlist, however, and at the top of that is for us to turn our attention to making the digitized texts even more useful by turning that uncorrected OCR into a more faithful reproduction of the original book. While large-scale linguistic studies may be valid in spite of a small percentage of errors, the use of the digitized materials for reading, in the case of those works in the public domain, and for listening, in the case of works made available to VIPs (visually impaired persons), is greatly hampered by the number and kinds of errors that result. In a future post I will give the results of a short study that I have done in that area.

See all my posts on Google Books

Tuesday, September 24, 2013

Hopes and fears for Google Books case

We're back in the saddle of the now epic lawsuit against Google for its massive scanning of the books held by libraries. I have very mixed feelings about the case and its outcomes, and the news reports from yesterday's hearing (transcript) in Judge Denny Chin's court are not making me feel any better about it. In brief, the Author's Guild is claiming that Google violated fair use by scanning in-copyright books. Since that act alone is not sufficient to address a defense of fair use, they also state (correctly, in my view) that although Google is not providing advertising on the individual book pages, that it overall makes money off of the scanned books because that digital corpus enforces its position against other search engines. There are some things that the Authors Guild has right (such as, that Google makes money off of search results pages that can include links to Google Books), but they miss the mark in other arguments:
"For all intents and purposes, it paid libraries for the right to digitize and copy much of our nation’s literary heritage and then used the resulting digital library to gain a competitive advantage over search engine competitors that respected the rights of authors by limiting their digitization programs to books that were either licensed or were no longer protected by copyright. Aided by its infringing conduct, Google’s search engine has proven remarkably successful—to the point where “google” has become a widely used verb in the English language.
First, the addition of Google Books to the search took place long after we were all "googling." Google's main value still comes from providing access to open web resources that otherwise would just be a massive digital junk heap. I suspect that those who are interested in using Google to search within the text of "closed" books (ones that are not available as full text online) consciously go to the Google Book Search pages. I don't know this for a fact, but I'd be willing to bet that user intent behind most Google searches is to access the actual content of a web page or document, not to be given a reference to an off-line resource.

Next is the statement that Google "paid libraries for the right to digitize..." This makes it sound like Google gave the libraries money, and that there was no cost to the libraries. The agreement between Google and libraries was an exchange that had costs for both (less for the libraries, more for Google) and benefits for both (less for the libraries, more for Google). In the end, Google got the better part of the deal, but libraries got something, even though something they have not yet been able to greatly benefit from: libraries got copies of the scans at a lower price than had they done the digitization themselves. Unfortunately, due to both copyright issues and the nature of the agreement between Google and the libraries, there are significant barriers to making the kind of uses that would make this a truly transformative corpus for research.

All of the news reports emphasized some comments by Judge Chin to the effect that Google Books appears to be both transformative (in the copyright law sense) and a benefit to society. What worries me a bit is that Judge Chin is not looking beyond the use of the resulting digital texts for search. I consider search to be the tip of the iceberg, and the visible part of Google Books that Google would like everyone to focus on. My assumption is that Google has a research interest in having exclusive access to 20 million non-Web digital texts in a myriad of languages, and that this research is aimed not only at search but at Google's desire to be THE interface between man and machine, which means that machines have to get better at human languages.

If Judge Chin rules that Google's book digitization is fair use, it's a huge win, not only for Google but also for libraries. After all, if it is fair use for Google to digitize works for the purposes of searching, there is no question that it is also fair use for libraries to do the same. If Judge Chin rules that Google's book digitization is NOT fair use because of profit-making, then we still do not know for sure whether library digitization would be considered fair use (although much would depend on exactly how the decision is worded). This of course makes me want to cheer on Chin toward the "is fair use" decision, but at the same time I know that this means that any research that Google is doing on its private cache of digital texts will continue, giving them great advantages over competitors in the arms race of technology advancement.

Once again, I so wish that large-scale digitization for search and research had been undertaken by libraries, not Google. The questions of "not for profit" and social value would be a slam-dunk, and I'd not be harboring this fear that there is a hidden agenda behind the project. Maybe if libraries had done this we'd only have two or three million digitized books, not 20 million (as is claimed for Google), but they'd be untainted, in my mind, and I could still consider them a cultural heritage resource rather than a commercial product.

Friday, November 23, 2012

Fair Use(-ful)

The beauty and the aggravation of Fair Use in US copyright law is that one cannot pre-define particular uses as "fair." The countries that have, instead, the legal concept of "Fair Dealing" have an enumerated set of uses that are considered fair, although there is obviously still some need for interpretation. The advantage to Fair Use is that it can be re-interpreted with the times without the need for modification of the law. As new technologies come along, such as digitization of previously analog works, courts can make a decision based on the same four factors that have been used for earlier technologies. However, until such a decision is made in a court of law, it isn't possible to be sure whether a use is fair or not.

We have recently seen a court case that decided that HathiTrust's use of digitized books to provide an index to those books is fair. There is another court case that will decide a similar question regarding Google's digitization of books for its Google Book Search. Note, however, that even if both of these are determined to be fair use, each is a particular situation in a particular context. Both organizations have developed their services in an attempt to meet what they judged to be the letter of the law, and yet there is a considerable difference in the services they provide.

HathiTrust stores copies of digitized books from the collections of member libraries. In this case, HT is not itself doing the digitization but is storing files for books mostly digitized by Google. A search in the full text database of OCR'd page images returns, for in-copyright items, the page numbers on which the terms were found, and the number of hits found on each page. There are no snippets and no view of the text unless the text itself is deemed to be out of copyright.

Google has a different approach. To begin with, Google has performed mass digitization of books (estimated at about 20 million) without first obtaining permission from rights holders. So the Google case includes the act of digitization, whereas the HathiTrust case begins with digital files obtained from Google. Therefore the act of digitizing was not a factor in that case. In terms of use of the digitized works, Google also provides keyword searching of the OCR'd digital images, but takes a different approach to the results viewable by the searchers. Google provides short (about 3-5 lines) snippets that show the search terms in context on a page.
Google, however, places specific restrictions to avoid letting users "game" the search to gain access to enough of the text to substitute for actually acquiring access to the book. Here is how Google describes this in its recent legal response:
"The information that appears in Google Books does not substitute for reading the book. Google displays no more than three snippets from a book in response to a search query, even if the search term appears many times in the book. ... Google also prevents users from view a full page, or even several contiguous snippets, by displaying only one snippet per page in response to a given search and by 'blacking' (i.e. making unable for snippet view in response to any search) at least one snippet per page and one out of ten pages in a book." p.8
Google also exempts some types of books, like reference works, cookbooks, and poetry, from snippet display entirely.

The differences in the results returned by these two services reflect the differences in their contexts and their goals. HathiTrust has member institutions and their authorized users. The collection within HathiTrust reflects the holdings of the member institutions' libraries which means that the authorized users should have access, either in their library or through inter-library loan, to the physical book that was scanned. The HathiTrust full text is a search on the members' "stuff." The decision to give only page numbers makes some sense in this context, although providing snippets to scholars might have been acceptable to the judge. The return of page numbers and full word counts within pages reflects, IMO, the interest in quantitative analysis of term use. It also gives scholars some idea of the weight the term has within the text.

Google's situation is different. Google has no institutions, no members, no libraries; it provides its service to the general public (at least to the US public). There is no reason to assume that all of the members of that public will have access to the hard copy of any particular digitized book. Google seems to have decided that promoting its service as having primarily a marketing function, with the snippets as "teasers," would mollify the various intellectual property owners. In its brief of November 9, Google reiterates that it does not put advertising on the Google Book Search results pages, nor does Google make any money off of its referrals to book purchasing sites.

So here are two organizations that have bent over backwards to stay within what they deemed to be the boundaries of fair use, and they have done so in significantly different ways. This means that the fair use determination of each of these could have different outcomes, and each will provide different clues as to how fair use is viewed for digitized works.

It of course bears mentioning that both of these solutions provide hurdles for users. The HathiTrust user who is searching on a term that could have more than one meaning ("iron" "dive" "foot") does not have any context to help her understand if the results are relevant. The Google user, on the other hand, gets some context but cannot see all of the results and therefore does not know if there are key retrievals among those that have been blocked algorithmically. A use that is "fair" within copyright law may not seem "fair" to the user who is doing research. It makes you wonder if our idea of "fair use" couldn't be extended to be fair but also "useful."

Related posts
http://kcoyle.blogspot.com/2012/10/copyright-victories-part-ii.html

Tuesday, October 16, 2012

Copyright Victories, Part II

I did a short factual piece for InfoToday on the Authors Guild v. HathiTrust decision that was issued last week. The Authors Guild brought the suit against HathiTrust because HathiTrust is storing copies of books, digitized by Google, that are still under copyright. Fortunately for HathiTrust, its partners, and all of us in libraries, the judge decided:
  • The digitization of books for the purposes of providing a searchable index is transformative, and therefore is a Fair Use under copyright law.
  • The provision of these search capabilities “promotes the Progress of Science and useful Arts” and thus supports the goals of U.S. copyright policy and law.
  • The provision of in-copyright texts for visually impaired students and researchers is in direct support of the Americans With Disabilities Act.
The decision in the case of the Authors Guild v. Hathitrust echoes some of the same thinking as the GSU case, in particular on the educational and research use of intellectual property. This case hinged on the use of the digitized texts for indexing rather than for reading. The judge determined that the books in HathiTrust were not substitutes for the books on the library shelves, since they are not presented to users as texts to be read. The "transformation" of the readable texts to a searchable index that returns only page numbers and the number of times a term appears on the page results in a new product, not an imitation of the hard copy.

The judge decided this for HathiTrust, but this is the same question that is being asked in the Authors Guild lawsuit against Google. There are some obvious differences between the two situations, however. First, unlike HathiTrust Google is a for-profit company, so it loses points on the first factor of the fair use test:

(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
Because Google is digitizing works primarily from university libraries, both HathiTrust and Google do well on the second factor:

(2) the nature of the copyrighted work;
 Works of a creative nature (defined as "prose fiction, poetry, and drama") are given greater protection than works of fact. HathiTrust reports that only 9% of its digital collection meets the "creative" definition.

The third factor:

(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole

would seem to go against HathiTrust (and Google), but the judge looked at the two primary uses for the digital texts, keyword indexing and providing digital copies to members of the community with sight disabilities, and determined that they could not be done with anything less than a complete copy. If the argument of transformation is made in the case against Google, this factor should be the same.

Factor four is about the effect on the market:

(4) the effect of the use upon the potential market for or value of the copyrighted work

This is a bit tricky because presumably HathiTrust will point its users to the library-owned hard copies of the book, especially since many of the digitized books will be out of print and not available from publishers. Therefore there isn't much interaction with the market at all. The judge added that, if anything, the greater amount of discovery might lead to sales, but I wouldn't hold out much hope for that. The other use is to provide access to the blind; this is a non-market for print materials if there ever was one. Google, on the other hand, has partnered with publishers to sell digitized books as ebooks, and therefore the positive market force should be stronger in that case if Google can show that previously out-of-print books can be sold through its service.

Not mentioned anywhere that I can find is the question of digital "photographs" of pages vs. OCR'd text. The suit and the decision blend these together as "a digital copy." Having seen some of the results of Google's digitization I can say that the text resulting from the OCR can be quite lossy depending on the page layout (tables of contents in particular come out quite badly) and the quality of the original book. It is also the "transformation" part of the copying, since the photographs of the pages are simply copies of the page and are by their nature human-readable substitutes for the page itself. The judge seems to consider these "transitory" but in fact they are quite solidly real, and are stored in the HathiTrust repository. I suspect it is these pictures of the pages that the Authors Guild fears will be pirated should HathiTrust be hacked, less so the OCR'd pages which are unattractive plain text. However, HathiTrust was able to show the judge that it takes security quite seriously, and the Authors Guild was unable to demonstrate any quantifiable risk. 

What is heartening in this decision is the judge's enthusiasm for the role of libraries in further science and knowledge, and his great admiration of HathiTrust's service to scholars and to the blind. His decision is both factual and moral: he refers to the "invaluable contribution to the progress of science and the cultivation of the arts that at the same time effectuates the ideas espoused by the ADA." We could not have hoped for a better advocate for digital libraries than Judge Harold Baer Jr.

Saturday, October 06, 2012

Google and publishers settle

Recent news announcements state that the AAP and Google have settled their lawsuit. Essentially this is a formality rather than an actual change in status between Google and the publishers. As you probably recall, the AAP engaged in a lawsuit against Google, in partnership with the Author's Guild, in 2005, and five specific publishers were named in that suit. Since then the lawsuit has gone through two failed attempts to settle and a massive number of pages of legal parrying. Meanwhile, the publishers realized that Google provided them with a new sales opportunity and about 40,000 of them have entered into agreements with Google in which the publisher provides (or allows Google to provide) a digital copy of the book which then can be sold as a Google eBook. Each contract specifies the amount of the book that potential buyers can browse. This browsing is designed to prevent users from satisfying their needs without a purchase: although up to 20% of a digital book may be browsable, Google denies access to sequences of more than 9 pages in a row.

When the Author's Guild revived their suit against Google in late 2011, the AAP was notably absent from the list of plaintiffs. By then the publishers had made a satisfactory agreement with Google and had no interest in suing their business partner. This current settlement appears to be a pro forma legal action to terminate the previous lawsuit, and, as far as I can tell, makes no change to the current status of the business relationship between Google and publishers.

There is one question that remains which is what this means in terms of copyright, if anything. The contractual arrangements between Google and the publishers are standard business agreements which the publishers engage in as representatives of the rights holder through their contract with that rights holder. Sometimes the publisher also holds the copyright, but that isn't the salient point here. So unless I am missing something, this agreement has absolutely no effect on the questions of copyright and digitization that some of us have been so eager to hear about. It's just business as usual.

[After-note: Some authors and journalists question the settlement and want details made public.]

Tuesday, September 18, 2012

Seven years, and waiting

Just a quick note to say that the Google Book Search lawsuit has entered a new "pause" and will thus be delayed further. Among the salvos that Google and the Author's Guild have fired back and forth are those around the question of whether the Author's Guild can legally represent all authors against Google. As I recall, the Author's Guild has something on the order of 5,000 members, all alive (as far as I know), while Google's digitization now covers somewhere upwards to 10 million books and probably nearly as many authors, from the days of early printed works to the present.

Google challenged the Author's Guild as class representative of all authors, but Judge Denny Chin, the judge who has seen this case through most of its life, allowed the suit to go forward with the Guild as class representative. This has been reversed by an appeals court judge, which means that the question of class representation must be decided before the suit can continue. 

Sunday, July 29, 2012

Fair Use Dejà Vu

In its July 27 court filing,[1] Google has made the case for its fair use defense for the digitization of books in its Google Book Search (GBS) project. [2] As many of us have hoped, the case it makes appears strong. That it was necessary to throw libraries under the bus to achieve this is unfortunate, but I honestly do not see a an alternative that wouldn't weaken the case a bit.

Fair Use is Fair

The argument that Google has made from the beginning of its book scanning project is that copying for the purpose of providing keyword access to full texts is fair use. They are fortunately able to cite case law to defend this, including case law allowing the copying of entire images by image search engines.

Among the reasons that they give for their fair use defense are:

1. Keyword search is not a substitute for the text itself. In fact, the copy of the text is necessary to provide a means for users to discover the existence of books and therefore for the books to fulfill their purpose of being read.
"Books exist to be read. Google Books exists to help readers find those books. Like a paper index or a card catalogue, it does not substitute for reading the books themselves..." (p. 2)

2. Google has elaborate protections in place to prevent users from reconstructing the text from its products. They reveal some of these protections, such as disabling snippet display for one instance of the keyword on each page, and disabling display of one page out of ten.
"One of the snippets on each page is blacklisted (meaning that it will not be shown). In addition, at least one out of ten entire pages in each book is blacklisted." (p. 10)
3. No advertising appears on the GBS pages. This implies that Google is not making any money that could be claimed by authors as being theirs.

4. The Authors Guild has no proof of harm that has come from the digitization of the books. It is suggested that a thorough study might show that there have been gains rather than losses in terms of book sales. Even the Authors Guild (the Plaintiff in this case) advises authors to provide some of the text of their books (usually the first chapter) for browsing in online bookstores, and many rights holders participate voluntarily in Amazon's "Look inside" feature that shows considerably more than the disputed snippets that are displayed in GBS. And Google notes that 45,000 (!) publishers have signed up to have their in-print books searchable in GBS, with varying amounts of text available to the searcher prior to purchase. This makes the case that search and some text display is good for authors, not harmful.

5. Digital copies of books have never been "distributed to the public" (key wording in the copyright law). Only the libraries themselves that held the actual hard copies could receive a copy of the files resulting from the digitization.

Of course, all of this is done citing court cases in support of these arguments. The Authors Guild undoubtedly has counter-cases to present.

Libraries Under the Bus

One of the key copyright-related arguments that Google makes is that its full text search within books provides a public service and support of research that is unprecedented. In making these claims Google decided to particularly emphasize its superiority to library catalogs. (Google refers multiple times to "card catalogues" which seems oddly antiquated, but perhaps that was the intent.)
"The tool is not a substitute for the books themselves -- readers still must buy a book from a store or borrow it from a library to read it. Rather, Google Books is an important advance on the card-catalogue method of finding books. The advance is simply stated: unlike card catalogues, which are limited to a very small amount of bibliographic information, Google Books permits full-text search, identifying books that could never be found using even the most thorough card catalog." (p.1) [sic uses of "catalogue" and "catalog" in the same paragraph.]
"Google Books was born of the realization that much of the store of human knowledge lies in books on library shelves where it is very difficult to find....Despite the importance of this vast store of human knowledge, there exists no centralized way to search these texts to identify which might be germane to the interests of a particular reader." (p. 4)
As a librarian, I have to say that this dismissal of the library as inadequate really hurts. Yet I believe that Google is expressing an opinion that is probably quite common among information searchers today. One could counter with many examples where the library catalog entry succeeds and GBS fails, but of course that wouldn't bolster Google's arguments here. A reasonable analysis would put the two methods (full text and standards-based metadata) as complementary.

Google also argues that it did not give copies of the digital files resulting from its scanning to the libraries. How this plays out is not only clever, but it shows some real foresight on Google's part. They developed a portal where the libraries could request that a copy of the files be made "on demand" for the library, and using an encryption specific to that library. The transmission of the files from Google to the libraries was then an act of the libraries, not of Google.
"Moreover, the undisputed facts show that it is the libraries that make the library copies, not Google, and that Google provides only a technological system that enables libraries to create digital copies of books in their collections. Under established Second Circuit precedent, Google cannot be held directly liable for infringement because Google itself has not engaged in any volitional act constituting distribution." (p. 33)
Clearly, Google designed the system (with goes by the acronym "GRIN") with this in mind.

I don't mind this, but wish that Google hadn't included a dig at HathiTrust as part of this argument. The document would not have suffered, in my opinion, if Google had left the parenthetical phrase off of this sentence:
"No library may obtain a digital copy created from another library's book -- even if both libraries own identical copies of that book (although libraries may delegate that task to a technical service provider such as HathiTrust)." (p. 15)
It's one thing to claim innocence, but another to point the finger at others.

Omissions

There a few glaring omissions from the document, some of which would weaken Google's case.

There is no mention of the computational uses that can be made of the digital corpus, something that was a strong focus in the failed settlement between Google and the authors and publishers. I have no doubts that Google is currently engaged in research using this corpus -- I don't see how they could resist doing so. They do mention the "n-gram" feature briefly, but as this is based on what appears to be a simple use of term frequency, it may not attract the court's attention.

In another omission, Google states that:
"Informed by the results of a search of that index, users can click on links in Google Books to locate a library from which to borrow those books ... " (p. 4)
Google fails to state that this is not a service provided by Google but one provided by OCLC using exactly those card catalogues that Google finds so inadequate. Credit should be given where credit is due, but there is an important battle to be won.

Bottom Line

The ability to create full text searches of printed works (and other physical materials) is so important to research and learning -- and should be such an obvious modern approach to searching these materials -- that a win for Google is a win for us all. Although some aspects of this document shot arrows into my librarian-ly heart, I hope with all of that wounded heart that they prevail in this suit.


[1] This points to the ScribD site which unfortunately is now connected to Facebook and therefore is a huge privacy monster. The document should appear on the Public Index site shortly, with no login required.
[2] The term "product" could also be used to describe GBS.

Sunday, July 15, 2012

Friends of HathiTrust

I have written before about the lawsuit of the Author's Guild (AG) against HathiTrust (HT). The tweet-sized explanation is that the AG claims that the corpus of digitized books in the HathiTrust that are not in the public domain are infringements of copyright. HathiTrust claims that the digitized copies are justified under fair use. (It may be relevant that many of the digitized texts stored in HT are the result of the mass digitization done by Google.)

For analysis of the legal issues, please see James Grimmelman's blog, in particular his post summarizing how the various arguments fit into the copyright law's "four factors."

I want to focus on some issues that I think are of particular interest to librarians and scholars. In particular, I want to bring up some of the points from the amicus brief from the digital humanities and law scholars.

While scientists and others who work with quantifiable data (social scientists using census data, business researchers with huge amounts of data from stock markets, etc.), those working in the humanities whose raw material is in printed texts have not been able to make use of the massive data mining techniques that are moving other areas of research forward. If you want to study how language has changed over time, or when certain concepts entered the vocabulary of mass media, the physical storage of this information makes it impossible to run these as calculations, and the size of the corpus makes it very difficult, if not impossible, to do the research in "human time". Thus, the only way for the "Digital Humanities" to engage in modern research is after the digitization of their primary materials.

This presumably speaks to the first factor of fair use:

(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;

As Grimmelman says "The Authors Guild focuses on the corpus itself; HathiTrust focuses on its uses." It may make sense that scholars should be allowed to make copies of any material they need to use in their research, but I can imagine objections, some of which the AG has already made: 1) you don't need to systematically copy every book in every library to do your research and 2) that's fine, but can you guarantee that infringing copies will not be distributed?

It's a hard sell, yet it's also hard not to see the point of view of the humanities scholars who feel that they could make great progress (ok, and some good career moves) if they had access to this material.

The other argument that the digital humanities scholars make is that the data derived from the digitization process is not infringing because it is non-expressive metadata. Here it gets a bit confusing because although they refer to the data derived from digitization as "metadata," the examples that they give vary from the digitized copies themselves, to a database where all of this is stored, and to the output from Google n-grams. If the database consists of metadata, then the Google n-grams are an example of the use of that metadata, but are not an example of the metadata itself. In fact the "metadata" that is produced from digitization is a good graphic copy of each page of the book, plus a reproduction, word for word (with unfortunate but not deliberate imprecision) of the text itself. That this copy is essential for the research uses desired is undeniable, and the brief gives many good examples of quantitative research in the humanities. But I fear that their insistence that digitization produces mere "metadata" may not be convincing.

Here's a short version from the text:

"In ruling on the parties’ motions, the Court should recognize that text mining is a non-expressive use that presents no legally cognizable conflict with the statutory rights or interests of the copyright holders. Where, as here, the output of a database—i.e., the data it produces and displays—is noninfringing, this Court should find that the creation and operation of the database itself is likewise noninfringing. The copying required to convert paper library books into a searchable digital database is properly considered a “nonexpressive use” because the works are copied for reasons unrelated to their protectable expressive qualities; none of the works in question are being read by humans as they would be if sitting on the shelves of a library or bookstore." p. 2

They also talk about transformation of works, and the legal issues here are complex and my impression is that the various past legal decisions may not provide a clear path. They then end a section with this quote:

"By contrast, the many forms of metadata produced by the library digitization at the heart of this litigation do not merely recast copyrightable expression from underlying works; rather, the metadata encompasses numerous uncopyrightable facts about the works, such as author, title, frequency of particular words or phrases, and the like." (p.17)

This, to me, comes completely out of left field. Anyone who has done digitization projects is aware that most projects use human-produced library metadata for the authors and titles of the digitized works. In addition, the result of the OCR step of the digitization process is a large text file that is the text, from first word to last, in that order, and possibly a mapping file that gives the coordinates of the location of each word on each OCR'd page. Any term frequency data is a few steps away from the actual digitization process and its immediate output, and fits in perfectly with the earlier arguments around the use of datamining.

I do sincerely hope that digitization of texts will be permitted by the court for the purposes argued in this paper. An attempt at justification, after the fact, of Google's mass digitization project may, however, suffer weaknesses inherent in that project, in particular that no prior negotiation was attempted with either authors nor publishers, and once the amended settlement between Google and the suing parties was denied by court, there is no mutual agreement on uses, security, nor compensation.

In addition, the economic and emotional impact of Google's role in this process cannot be ignored: this is a company that is so strong and so pervasive in our lives that mere nations struggle to protect their own (and their citizens') interests. When Google or Amazon or Facebook steps into your territory, the earth trembles and fear is not an unreasonable response. I worry that idea of digitization itself has been tainted, making it harder for scholars to make their case of the potential benefits of post-digitization research.

Friday, June 01, 2012

Google Books: TBD

The latest

The Google Book Search lawsuit is essentially back to square one. Judge Denny Chin has ruled on an important aspect of the post-(failed)settlement lawsuit of the Author's Guild against Google: Google's objection that the Author's Guild (AG) cannot represent all authors, since copyright must be determined on a case-by-case basis. (It has been widely noted that when it came to declaring the copying to be Fair Use, Google was happy to treat the works en masse, in direct contradiction of their response to this latest suit that a copyright suit claiming infringement would need to be individual.) Chin has ruled that the Author's Guild can proceed as representative of "authors" as a class. The class includes not only those members of the association, but all authors whose books were scanned by Google. This means that the AG suit against Google can move forward, and that sometime in (hopefully) the near future we will have a ruling on whether or not Google's scanning of books is within the guidelines of Fair Use.

Quick update

In about 2004, Google began scanning books in partnership with a handful of major libraries, stating its goal as creating search access to books in the same way that it provides search access to web pages. Google Book Search gave results looking much like those for Google's web search: minimal metadata and about three snippets from the book showing the context for the search terms. Google's claim was that copying the books solely for the purpose of search was a clear case of "fair use."

In 2005, the Author's Guild sued Google for copyright infringement in a class action lawsuit representing "authors" as a class. Shortly thereafter, the Association of American Publishers brought their own lawsuit on the part of publishers.

Scanning of books in the libraries continued through 2008 with no word about the lawsuit. Meanwhile, more libraries had been added to the program, and it is estimated that about 7 million books had been scanned. (The exact number is not known.)

In October of 2008, much to the surprise of nearly everyone, it was announced that Google had arrived at a settlement with the AG and the AAP. The settlement was far-reaching, and created a mechanism that would allow Google to scan out-of-print books and make them available for use or sale, returning revenues to the copyright holders. Copying continued.

The settlement received hundreds of responses, mostly negative, from authors and from publishers, in particular those not in the US. Initially the parties were sent back to revise the settlement to address certain concerns. They did so, but not to the judge's satisfaction: the settlement was rejected in court in March of 2011.

In late 2011, the Author's Guild (without the publishers) updated and reprised its original lawsuit against Google, primarily demanding that all scanning stop. It also brought a suit against Hathitrust, the library-sponsored archival facility that houses many of the library books scanned by Google.  The lawsuit claims that the copies in HathiTrust are not legal copies and demands that they be destroyed.

My opinion

I could imagine getting a fair use ruling based on the original definition of the project, which was the scanning of books solely for the purpose of allowing keyword searching on texts, with minimal metadata and a few short snippets shown to the public, although Google's for-profit status might have nixed such a decision. However, this was complicated by the active participation of the libraries, and by the fact that Google returned a copy of the digital scan (and sometimes also the OCR and the OCR "map" that carries the location of the text on the page) to the library that had offered the book. While the copying for search might be considered fair use, since no actual copies of the books are made available to anyone, the presentation of a copy to the libraries is a pretty clear act of copying.

The terms of the settlement were arrived at through negotiations between the parties, and including input from some of the library partners. It was during this time that library partner U of Michigan began planning the archive now known as HathiTrust. It is undoubtedly not a coincidence that such an archive was described in a fair amount of detail in the settlement document as a requirement for library archiving of their received digital copies. In addition, the settlement allowed for computational research on the corpus, something that would be of great benefit to researchers.

During the time between 2008 and 2011, when the parties presented the first settlement and then the amended settlement, there was a fair amount of optimism that the settlement would be accepted, and plans to engage in the terms of the settlement, including the creation of a special bureau to manage payments to copyright holders, went forward. Google appeared to be all-powerful, able to bend law and legislation in order to create an entirely new view of copyright and digitization.

With the rejection of the settlement, and this latest ruling that allows the AG lawsuit to go forward, the picture has entirely changed, but not necessarily for the better. Many were hoping that we would be able to digitize the entirety of our analog matter, increasing access and preservation capabilities.

Now we are facing the possibility that not only may mass digitization be declared in violation of copyright, at least in this instance, but that the libraries may lose the copies the digital versions of the items in their collection, and researchers will lose the access to these items in HathiTrust and Google Book Search.

At the same time, for Google the lawsuit has become nearly moot. At the moment Google has tens of thousands of publishing partners that allow Google to index digital versions of their books and make them available for sale either in hard copy or as ebooks. Google could lose the out-of-print books in its collection for which it has no publisher agreement, but these books are not providing any revenue for Google.

Resources


the public index - all of the filed documents for the lawsuits, beginning in 2005

HathiTrust Information about the AG lawsuit

James Grimmelman's analysis of yesterday's decision

My posts are found under tag "googlebooks"

Monday, January 02, 2012

Google Book Search Redux

The document I referred to in the previous post would have been so much clearer if I had read the two preceding documents. Now that I have, the story is even more dramatic.

On December 12, 2011, the Author's Guild filed a fourth amended complaint (PDF) against Google. This complaint is nearly identical to the first one, filed on September 20, 2005 (PDF). The two complaints between these (October 28, 2008 and November 16, 2009) included the Association of American Publishers, as did the two attempts at settling the case. (October 28, 2008, and November 13, 2009). The publishers had had their own complaint in 2005 before combining forces with the Authors Guild. Now the Authors Guild is again standing alone against Google's book digitizing efforts.

This fourth amended complaint brings us pretty much back to square one, with the addition of the involvement of more libraries and the creation of HathiTrust as a way for the libraries to store their (allegedly) ill-gotten copies. The library copies are a key element of the suit because they are proof that Google has not only digitized the library books but has made copies (the purview of copyright law) and distributed them to others.

The most interesting document of this latest group, and the one with the greatest detail about Google's actions, is the Memorandum in support of the class certification. This document is the explanation of why the Authors Guild should be considered by the court to be a valid representative of all authors in a class action suit. The document has a number of quotable moments, of which my favorite is the "tell it like it is in plain language" opening:
This litigation arises from Google's business decision to gain a competitive edge over its rivals in the search engine market by making digital copies of millions of "offline" printed materials. ... Rather than obtaining licenses from copyright owners for the digital use of their printed works, Google instead entered into agreements with libraries to gain access to these works. A number of university libraries allowed Google to make digital copies of the books in the libraries' collections, including in-copyright books. In exchange, Google provided digital copies of the books to the libraries. Google refers to this massive copyright infringement as its "Library Project." (p.1)
The assumption on the part of most folks commenting on this latest development in this now 6-year-old case is that the settlement is dead. We are therefore back to the question of whether Google's book scanning is or is not Fair Use. This question, though, is only being asked on the part of authors, not publishers, and if anyone has inside knowledge on what approach the publishers are taking I would love to hear it. It is clear that the position of publishers in relation to Google has changed greatly over these past 5-6 years since the suit was originally filed. There are now reportedly thousands of publishers who are using Google Books to promote and sell their works. It also makes sense that publishers, as corporations, are better able to negotiate with Google than are individual authors. A large publisher with numerous books in print and in its backlist has clout that a single person does not have. In addition, large publishers have lawyers, or access to legal counsel. At least some publishers have made their peace with Google and are seeing the relationship as advantageous.

Looking at this from the library point of view I wonder what will happen to the millions of library books already scanned by Google. I also wonder what this awkward and failed attempt to create the overly broad settlement between Google and the AG/AAP will mean for future digitization projects. There are strong arguments for digitization for scholarly purposes, and the creation of a computational capability over millions of texts could be a positive step for research, especially in the social sciences and humanities. I hope that the botched attempt to commercialize the contents of libraries will not prejudice the future of digital research.

Monday, December 26, 2011

Google files motion to dismiss

"The claims of the associations should be dismissed without leave to amend because they lack standing as a matter of law, since they do not themselves own copyrights and do not meet the test for associational standing set forth in Hunt." p. 19
With that conclusion, Google has filed a motion asserting that the copyright infringement lawsuits filed by the Authors' Guild and the American Society of Media Photographers, Inc. be dismissed. The arguments made in the document are:
  • "Individual copyright owners' participation is necessary to establish a claim for copyright infringement." (p.1)
  • "Plaintiff associations do not own copyrights alleged to have been infringed, and do not have standing to sue for copyright infringement." (p.4)
  • "Every copyright, and every alleged copyright infringement, is different."(p.7)
  • "... a central issue in these cases is whether the conduct alleged in the Complaints constitute fair use under 17 U.S.C. 107. Litigating that issue will require the participation of individual association members, because many of the relevant facts are specific to the particular work in question." (p.11)
All of this sounds plausible to this legal novice, but there are a couple of puzzling issues. First, why did Google not make these arguments in 2005 when the Authors' Guild filed suit? Instead, they negotiated with the association for six years, presumably in good faith, and those negotiations hinged on the acceptance of the AG as a representative of authors and their rights in their works. If Google had thought that the AG did not have standing, none of that negotiation would have made much sense.

Second, Google says in this document that fair use has to be determined on a case-by-case basis. They even quote from Campbell v. Acuff-Rose Music, Inc. that "Fair use must 'be judged case by case, in light of the ends of the copyright law....' It is 'not to be simplified with bright-line rules." (p.11) This seems to undermine Google's original defense that copying for the purposes of creating an index is itself fair use, not something that has to be determined on a case by case basis.

It isn't surprising the Google wants to bring an end to this case. It is now entering its seventh year (the original suit was filed in September of 2005), and has undoubtedly been costly for all parties. Google was moving ahead in putting into place the foundations for the settlement, including the creation of a large database of works and a means for owners to claim the copyrights. They had designated a director for the Book Rights Registry, which would administer the business agreed on in the settlement. The failure of the settlement and the amended settlement to get court approval meant that all of that effort was for naught. Yet it isn't clear to me (and I hope someone can speak to this) what practical outcome Google is seeking for its book digitization effort. A dismissal of this nature would put Google in the rather cynical position of continuing book scanning knowing that few individual authors will have the means to take Google to court, and those individual payments would probably be affordable for this multi-billion dollar company. If dismissal is rejected, then at least that aspect of the suit is clarified, but next steps surely will be that the suit goes forward as first entered.

The one thing that is clear is that negotiations between Google and the AG are no longer on the horizon.

Note, also, that the Authors Guild has filed suit against the HathiTrust for copyright infringement, and the decision here will no doubt reflect on that case as well.

Friday, September 16, 2011

Due diligence do-over

In what I see as both a brave and an appropriate move, the University of Michigan admitted publicly that the Authors Guild had found some serious flaws in its process for identifying orphan works. The statement reaffirms the need to identify orphan works, and promises to revise its procedures.
"Having learned from our mistakes—we are, after all, an educational institution—we have already begun an examination of our procedures to identify the gaps that allowed volumes that are evidently not orphan works to be added to the list. Once we create a more robust, transparent, and fully documented process, we will proceed with the work, because we remain as certain as ever that our proposed uses of orphan works are lawful and important to the future of scholarship and the libraries that support it."
Among other things, what I find interesting in all this is that no one seems to be wondering why our copyright registration process is so broken that sometimes even the rights holders themselves don't know that they are the rights holders. It really shouldn't be this hard to find out if a work is covered by copyright. Larry Lessig covered this in his book "Free Culture," which is available online. The basic process of identifying copyrights is broken, and the burden is being placed on those who wish to make use of works. This is a clear anti-progress, pro-market bias in our copyright system.

Thursday, September 15, 2011

Diligence due

Oooof! Talk about making a BIG, public mistake.

HathiTrust's new Orphan Works Project proposed to do due diligence on works, then post them on the HT site for 90 days, after which those works would be assumed to be orphans and would then be made available (in full text) to members of the HT cooperating institutions. Sounds good, right? (Well, maybe other than the fact of posting the works on a site that few people even know about...)

The Authors Guild blog posted yesterday that it had found the rights holder of one of the books on HTs orphan works list in a little over 2 minutes using Google. (It's hard to believe that they didn't know this when the suit was filed on September 13 -- this is brilliant PR, if I ever saw it.) They then reported finding two others.

James Grimmelman, Associate Professor at New York Law School and someone considered expert on the Google Books case, has titled his blog post on this: "HathiTrust Single-Handedly Sinks Orphan Works Reform," stating that this incident will be brought up whenever anyone claims to have done due diligence on orphan works. I'm not quite as pessimistic as James, but I do believe this will be brought up in court and will work against HT.

Wednesday, September 14, 2011

Authors Guild in Perspective

In its suit against HathiTrust the three authors guilds claim that there are digitized copies of millions of copyrighted books in HathiTrust, and that these should be removed from the database and stored in escrow off-line.

A relevant question is: who do the authors guilds represent, and how many of those books belong to the represented authors?

The combined members of the three authors guilds is about 13,000. That seems like a significant number, but the Library of Congress name authority file has about 8 million names. That file also contains name/title combinations, and I don't have any statistics that tell me how many of those there are. (If anyone out there has a copy of the file and can run some stats on it, I'd greatly appreciate it.) Some of the names are for writers whose works are all in the public domain. Yet no matter how we slice it, the authors guilds of the lawsuit represent a small percentage of authors whose in-copyright works are in the HathiTrust database.

The legal question then is: does this lawsuit pertain to all in-copyright works in HathiTrust, or only those by the represented authors? Could I, for example, sue HathiTrust for violating Fay Weldon's copyright?


Reply to this from James Grimmelman on his blog:
Good question, Karen, and one I plan to address in more detail in a civil procedure post in the next few days. In brief, you couldn’t sue to enforce Fay Weldon’s copyright, as you aren’t an “owner” of any of the rights in it. The Authors Guild and other organizations can sue on behalf of their members, but the details of associational standing are complicated. There is also the question of the scope of a possible injunction (e.g. could Fay Weldon win as to one of her works and obtain an injunction covering others, or works by others), where there are also significant limits on how far the court can go. Again, more soon.

As I suspected, the legal issues are complex. Keep an eye on James' blog for more on this.

Monday, September 12, 2011

Authors Guild Sues HathiTrust

There has been a period of limbo since Judge Chin rejected the proposed settlement between the Author's Guild/Association of American publishers and Google. In fact, a supposedly final meeting between the parties is scheduled for this Thursday, 9/15, in the judge's court.

Monday, 9/12, the Author's Guild (and partners) filed suit against HathiTrust (and partners) for some of the same "crimes" of which it had accused Google: essentially making unauthorized copies of in-copyright texts. In addition, the recent announcement that the libraries would allow their users to access items that had been deemed to be orphan works figures in the suit. That this suit has come over 6 years since the original suit against Google is in itself interesting. Nearly all of the actions of HathiTrust and its member libraries fall within what would have been allowed if the agreement that came out of that suit had been approved by the court. Although we do not know the final outcome of that suit (and anxiously await Thursdays meeting to see if it is revelatory), this suit against the libraries is surely a sign that AG/AAP and Google have not come to a reconciliation.

The Suit

First, the suit establishes that the libraries received copies of Google-digitized items from Google, and have sent copies of these items to HathiTrust, which in turn makes some number of copies as part of its archival function. This is followed by a somewhat short exposition of the areas of copyright law that are pertinent, with an emphasis on section 108, which allows libraries to make limited copies to replace deteriorating works. The suit states that the copying being done is not in accord with section 108. Then it refers to the Orphan Works Project that several libraries are partnering in, and the plan on the part of the libraries to make the full text of orphan works available to institutional users.

Since most of these institutions (if not all of them) are state institutions that have protections against paying out large sums in a lawsuit of this nature, the goal is to regain the control of the works by forcing HathiTrust (and the named libraries) to transfer their digital copies of in-copyright works to a "commercial grade" escrow agency with the files held off network "pending an appropriate act of Congress."

As James Grimmelman comments in his blog post on the suit, there's a lot of mixing up between the orphan works and owned works in the suit. He points out that a group of organizations representing authors could hardly make a case for orphan works since, by definition, the lack of ownership of the orphans means they can't be represented by a guild of people defending their own works.

The Problems

There are numerous problems that I see in the text of this suit. (IANAL, just a Librarian.)
  • The suit mentions large numbers of books that have been copied without permission, but makes no attempt to state how many of those books belong to the members of the plaintiff organizations.
  • The suit throws around large numbers without clearly stating that none of the statements include Public Domain works. It isn't clear, therefore, what the numbers represent: the entire holdings of HathiTrust, or just the in-copyright holdings. Also, in relation to the latter, unless one has done a considerable amount of work there are many works that are post-1923 that are also in the Public Domain. Cutting off at that year does not account for works that were not renewed, or were never copyrighted. I also doubt if anyone has a clear idea how many of the works in question are Public Domain because they are US Federal documents. This imprecision on the copyright status of works is very frustrating, but HathiTrust is not to blame for this state of affairs.
  • Some of their claims do not seem to me to be within legal bounds. For example, in one section they claim that although HathiTrust is not giving users access to in-copyright works, they potentially could. Where does that fit in?
  • They also claim that there is a risk of unauthorized access. However, the security at HathiTrust meets the security standards that the Author's Guild agreed to in the (unapproved) settlement with Google. If it was good enough then, why is it now too risky?
  • They claim that the libraries themselves have been digitizing in-copyright books. I wasn't aware of this, and would like to know if this is the case.
  • They state that the libraries said that before Google it was costing them $100 a book for digitization. Then the plaintiffs say that this means that the value of the digital files is in the hundreds of millions of dollars. First, I have heard figures that are more like $30 a book. Second, I don't see how the cost to digitize can translate into a value that is relevant to the complaint.
  • Although the legislature has failed to pass an orphan works law that would allow the use of these materials and still benefit owners if they do come forth, it seems like a poor strategy to complain about a well-designed program of due diligence and notification, which is what the libraries have designed. Orphan works are the least available works: if you have an owner you can ask permission; if there is no owner you cannot ask permission and therefore there is no way to use the work if your use falls outside of fair use. It's hard to argue for taking these works entirely out of the cultural realm simply because we have a poorly managed copyright ownership record.
  • There are a few odd sections where they make reference to bibliographic data as though that were part of the "unauthorized digitization" rather than data that was created by and belongs to the libraries. There's an odd attempt to make bibliographic data searching seem nefarious.
Parties
Plaintiffs: The Author's Guild, Inc.; The Australian Society of Authors limited ; Union des Erivaines Quebecois; Pat Cummings; Angelo Loukakis; Roxana Robinson; Andre Roy; James Shapiro; Daniele Simpson; T.J. Stiles; and Fay Weldon. (Links are to some sample HathiTrust records.)

Defendants: HathiTrust; The Regents of the University of California, The Board of Regents of the University of Wisconsin System; The Trustees of Indiana University; and Cornell University.

Links
Boing Boing: Authors Guild declares war on university effort to rescue orphaned books
Library Journal: Copyright Clash

Tuesday, March 22, 2011

Judge Chin rejects AAP/Google settlement

I'll say more when I've read it, but I put a copy on the Internet Archive.


After reading:

The judge's decision holds no real surprises. His analysis is fully consistent with the reactions of the interested parties to the case. He rejects the settlement primarily on these grounds:
  • It seems that a significant segment of the class of authors/publishers is not happy with the settlement. "Some 6800 class members opted out." (p.10) Also, a majority of the comments on the proposed settlement were negative, many coming from non-US copyright holders who did not identify with the class.
  • The settlement would make significant alterations to the current copyright regime, which should be a matter for Congress rather than the court.
  • The settlement's conclusion would go beyond the original lawsuit, which was over the digitization of in-copyright works by Google and the presentation of snippets relating to searches. The settlement would allow sales of full text works, which was never an issue at the time of the original lawsuit.
Although he rejects the settlement on numerous grounds, the judge concludes by saying "...many of the concerns raised in the objections would be ameliorated if the ASA were converted from an "out-out" settlement to an "opt-in" settlement." (p. 46) This leaves the door open for yet another settlement attempt between the parties.

It is important to note that the position of digitization and ebooks today are vastly different than they were in 2005 when the authors and publishers first sued Google over its library digitization project. It is possible that if the question of Google's digitizing were to be put forth for the first time today, the actions of the parties and the results would be vastly different. This is clearly a case where technology has moved forward at a rapid pace while the courts were contemplating an agreement that was standing still.

What now?

It's hard to believe that Google and the AAP/AG have not prepared themselves for this possibility. Yet, certain activities have gone forward as if the settlement were already approved.
  • A form of the Book Rights Registry is in place in the sense that there is a database of digitized works and a way to claim them to receive the proposed one-time payment. Presumably that payment is now not going to happen, but meanwhile Google has a large database with copyright holder information (including contact info, if I remember the form correctly).
  • The BRR has a chosen Director (Michael Healey).
  • It isn't clear if Google has continued digitizing books that are under copyright without specific permission. To be sure they have made many deals with publishers and with libraries to digitize works since the 2008 date when the settlement was first proposed, and digitization has gone forward.
  • Some libraries that had partnered with Google prior to the lawsuit have negotiated new contracts that are compatible with some of the conditions contained in the settlement. I don't know if these contracts have been signed or have been awaiting the result of the lawsuit but I do recall that the libraries obtained less rights in relation to retaining copies of their digitized books in the new contracts than they did in the old. The upshot being it isn't clear where this leaves the partner libraries, nor organizations like HathiTrust who are involved in the storage and possible uses of the Google digitized books.
  • For libraries and institutions that were looking forward to subscription access to the books, this access is now a big question mark. It was dependent on conditions in the settlement.
There are undoubtedly many other issues that are now open questions. When the settlement was first announced I began a "question list." It might be a good idea to revive that given this new perspective. And for those wondering "what now?" (that is, all of us) there's a flow chart.

Sunday, July 04, 2010

Catching up: OCLC, GBS, LOD

Some short comments on recurring themes:

OCLC Record Use Policy


OCLC has finalized its record use policy. The content is substantially the same as it was in the previous draft, which I commented on. There is one important improvement, however: the text clarifies OCLC's claims to copyright.
While, on behalf of its members, OCLC claims copyright rights in WorldCat as a compilation, it does not claim copyright ownership of individual records.
Of course, claiming copyright and actually having the right are not the same thing, especially with databases. Here's what BitLaw says:
Databases as Compilations: Databases are generally protected by copyright law as compilations. Under the Copyright Act, a compilation is defined as a "collection and assembling of preexisting materials or of data that are selected in such a way that the resulting work as a whole constitutes an original work of authorship." 17. U.S.C. § 101.
Generally, carefully selected compilations may make the "original work of authorship" cut; I'm not convinced that a union catalog of library holdings does.

Google Books

We are still waiting to hear from the judge in the Google Books case. (Every time I write that I check to see if it hasn't been released in the last hour.) Meanwhile, GBS continues to function in Internet time. Google has many publishers on board with its partners program, enough that GBS is becoming a serious rival to Amazon. It has even announced that it will begin selling e-books. The opening screen is the exact opposite of the Google Search screen -- it loads up many dozens of book covers and requires significant scrolling to browse to the bottom. Google has added personalization options ("my library") and lets you create multiple "shelves" to organize your materials.

Google was first sued in 2005. Five years is a very long time where technology is concerned. In 2005 the ebook was considered dead; now with the Kindle and the iPad, ebooks are alive and well and everyone is trying to get into that game. In that time since 2005, Google has pretty much shown the publishing industry that they can benefit from the online presence that Google is providing. The settlement reads like it was written in another era, trying to solve problems that may not really be considered problems today. The only issue remaining is that of orphan works, and if we could do a decent analysis of copyright holdings, I suspect that the number of orphan works would not be all that large.

Linked Library Data


At ALA there was a one-day preconference on linked data, and a half day un-conference attended by about 50 people. There are notes from the un-conference, which broke out barcamp-style into 6 groups for discussion.

The World Wide Web consortium has an incubator group on linked library data. This group is tasked to spend one year figuring out how to jump-start the creation of linked data in the library world.

There are ongoing efforts at Library of Congress to produce vocabularies, and of course the RDA vocabularies are available (and almost finalized). Ross Singer has announced some of the MARC codes are available (I presume on his own site). FRBR is being defined in linked data form by IFLA.

We've got just about everything but ... linked data. I'm thrilled that things are moving forward, but frustrated that I still can't see usable results. Deep breath; patience.

Sunday, February 21, 2010

Trust and the Settlement

In the week leading up to the hearing (Feb. 18, 2010) in New York in Judge Chin's court on the proposed settlement between the AAP/AG and Google, many parties weighed in with formal documents as well as informal ones. While few if any of these produced new information for the judge, they do reveal the different points of view of the parties involved.

One of these revelatory pieces is a blog post by the University of California's Ivy Anderson. Anderson has been involved in the negotiations with Google probably from the very beginning of UC's involvement. Her post attempts to counter the criticism of the settlement as well as many fears that have been expressed, with an emphasis on academia and academic libraries. For example, Anderson cites checks and balances on pricing that should prevent price gouging, as well as the possibility for the participating libraries to negotiate prices with Google.

I find two fundamental flaws in her arguments. The first is that she speaks from the perspective of a participating library, that is, a library that is able to negotiate directly with Google because of its position as a provider of books to be scanned. I have no doubt that this is a comfortable position for UC and for the other participating libraries, but they are small in number, especially compared to the total number of libraries and institutions that will be affected by the Google Book Search product. And of course their position is diametrically opposed to that of the general public, who have no voice in any of this project.

It doesn't surprise me that Anderson and others in similar positions have positive feelings about the settlement: they have been able to negotiate with Google and to make their needs known. Undoubtedly they have received some concessions. I also have no doubt that Google has been gracious and helpful. For all of the rest of us, however, the entire process has been a black box. We are being asked to trust the participating libraries, and to trust their trust in Google. Even though the needs of the participating libraries, all of whom are large research libraries, are almost certainly not the same as our own.

The second flaw that I see is Anderson's focus on Google as decision-maker. My reading of the composition of the governing body (should the settlement be approved) is that it will solely represent rights holders. It will set prices and even must approve Google's products. I find it interesting that we all (and Anderson included) tend to refer to this as the "Google settlement" -- but Google is the weak party in this particular situation. Remember that Google is the defendant, and that the mere act of settling is an admission of defeat. The libraries have hitched their wagon to the loser in this case. That can't be a good position.

I must say that I am much more afraid, if that's the right word, of the power that could be wielded by the AAP/AG should the settlement be approved. Google has many kind words to say about libraries. The AAP, however, has made it clear that they consider many library uses of materials to be infringements:
We also had significant concerns with respect to the digital copies that
Google was providing to libraries. Libraries might use significant portions, or all, of the contents of books on such copies for a range of purposes that publishers would not regard as permitted by the Copyright Act, including uses in classroom, “e-reserve” access to
students and faculty via institutional servers and lending digital copies to other libraries. Libraries might have raised fair use defenses in an attempt to justify such activities. We might also have been faced with sovereign immunity defenses by state institutions. In
addition, we were concerned about how the libraries could maintain the security of these digital copies. Security breaches might result in broad copying, uploading, downloading, and display of copyrighted works. (Statement of Richard Sarnoff, for the AAP board, p. 3)
The interesting upshot of this entire settlement process is that by digitizing the contents of libraries and managing those digital copies through contracts, the publishers could finally get the kind of control over library uses that they would have liked to have over the paper books held in libraries. They would like to have controls over inter-library loan, classroom use, and reserves, but they cannot exercise such controls in the analog world. Publishers have argued since the very early days of digital documents that all lending of digital documents is the making of a copy, and therefore is not allowed by copyright law.

As a matter of fact, right on page one of the Plaintiff's statement for the judge, among the bullet points describing the main achievements of the settlement, is this one:
Limits library uses of digital copies of Rightsholders’ works.
Perhaps it has been naive of me to see this settlement as being about Google's commercialization of the world of books. It is possible that the more pertinent end result could be a renewed control of books and their uses by the publisher community. Attempts to modify copyright law to cover digital resources have failed, and the rights of the public in relation to those resources are as yet unclear. This has left a gap that the AAP/AG settlement exploits fully.

OK, now I'm afraid!