Tuesday, October 16, 2012

Copyright Victories, Part II

I did a short factual piece for InfoToday on the Authors Guild v. HathiTrust decision that was issued last week. The Authors Guild brought the suit against HathiTrust because HathiTrust is storing copies of books, digitized by Google, that are still under copyright. Fortunately for HathiTrust, its partners, and all of us in libraries, the judge decided:
  • The digitization of books for the purposes of providing a searchable index is transformative, and therefore is a Fair Use under copyright law.
  • The provision of these search capabilities “promotes the Progress of Science and useful Arts” and thus supports the goals of U.S. copyright policy and law.
  • The provision of in-copyright texts for visually impaired students and researchers is in direct support of the Americans With Disabilities Act.
The decision in the case of the Authors Guild v. Hathitrust echoes some of the same thinking as the GSU case, in particular on the educational and research use of intellectual property. This case hinged on the use of the digitized texts for indexing rather than for reading. The judge determined that the books in HathiTrust were not substitutes for the books on the library shelves, since they are not presented to users as texts to be read. The "transformation" of the readable texts to a searchable index that returns only page numbers and the number of times a term appears on the page results in a new product, not an imitation of the hard copy.

The judge decided this for HathiTrust, but this is the same question that is being asked in the Authors Guild lawsuit against Google. There are some obvious differences between the two situations, however. First, unlike HathiTrust Google is a for-profit company, so it loses points on the first factor of the fair use test:

(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;
Because Google is digitizing works primarily from university libraries, both HathiTrust and Google do well on the second factor:

(2) the nature of the copyrighted work;
 Works of a creative nature (defined as "prose fiction, poetry, and drama") are given greater protection than works of fact. HathiTrust reports that only 9% of its digital collection meets the "creative" definition.

The third factor:

(3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole

would seem to go against HathiTrust (and Google), but the judge looked at the two primary uses for the digital texts, keyword indexing and providing digital copies to members of the community with sight disabilities, and determined that they could not be done with anything less than a complete copy. If the argument of transformation is made in the case against Google, this factor should be the same.

Factor four is about the effect on the market:

(4) the effect of the use upon the potential market for or value of the copyrighted work

This is a bit tricky because presumably HathiTrust will point its users to the library-owned hard copies of the book, especially since many of the digitized books will be out of print and not available from publishers. Therefore there isn't much interaction with the market at all. The judge added that, if anything, the greater amount of discovery might lead to sales, but I wouldn't hold out much hope for that. The other use is to provide access to the blind; this is a non-market for print materials if there ever was one. Google, on the other hand, has partnered with publishers to sell digitized books as ebooks, and therefore the positive market force should be stronger in that case if Google can show that previously out-of-print books can be sold through its service.

Not mentioned anywhere that I can find is the question of digital "photographs" of pages vs. OCR'd text. The suit and the decision blend these together as "a digital copy." Having seen some of the results of Google's digitization I can say that the text resulting from the OCR can be quite lossy depending on the page layout (tables of contents in particular come out quite badly) and the quality of the original book. It is also the "transformation" part of the copying, since the photographs of the pages are simply copies of the page and are by their nature human-readable substitutes for the page itself. The judge seems to consider these "transitory" but in fact they are quite solidly real, and are stored in the HathiTrust repository. I suspect it is these pictures of the pages that the Authors Guild fears will be pirated should HathiTrust be hacked, less so the OCR'd pages which are unattractive plain text. However, HathiTrust was able to show the judge that it takes security quite seriously, and the Authors Guild was unable to demonstrate any quantifiable risk. 

What is heartening in this decision is the judge's enthusiasm for the role of libraries in further science and knowledge, and his great admiration of HathiTrust's service to scholars and to the blind. His decision is both factual and moral: he refers to the "invaluable contribution to the progress of science and the cultivation of the arts that at the same time effectuates the ideas espoused by the ADA." We could not have hoped for a better advocate for digital libraries than Judge Harold Baer Jr.

Copyright Victories, Part I

While we are awaiting the results of the long-standing Google Books digitization copyright suit, there have been some important copyright battles libraries have won. The first was the Georgia State University digital reserves. The second is the recent decision regarding HathiTrust. I'll cover GSU in this post, and HathiTrust in a subsequent one, since my comments are long.

Earlier this year the case of publishers vs. Georgia State University regarding their e-reserves program resulted in a win for GSU as 69 out of 74 copyright infringement claims were denied by the judge. The case was about the provision of course readings in digital form. The suit was brought by three academic publishers (Oxford, Cambridge and Sage) but was bankrolled by the Association of American Publishers and the Copyright Clearance Center. Most readings were individual book chapters, and these were digitized by the library. Students in the class were able to access these through a password-protected site. The judge decided that for all but a few of the works the use was a fair use based on the nature of the use (educational) and the amount being used (one chapter). Unfortunately the judge also decided to enforce a "bright line" test of not more than 10% of the work (and the amount used of the works in question averaged 10.1%). As we know, bright lines are not in the spirit of fair use, yet the judge clearly needed some way to make her decisions.

One of the more striking things in the GSU case was that the publishers were unable to prove their ownership of the copyrights for over a third of the original items. They had published the books, but many of them were essentially anthologies and they could not find the appropriate paperwork for all of the individual pieces. This lack of proof in essence revealed these particular documents to be orphan works, even though the publisher of the book itself is known and some of the books may have been in print. I suspect that if you were to require actual proof of rights ownership for books or journal articles that the number of orphans would grow considerably. This would be especially true for articles, at least based on my experience: journal publishers are rather casual about getting signed agreements, and I have often modified agreements through strike-outs which were never contested.

This is just more evidence that our copyright system is a huge mess. Proving ownership requires expensive research (the copyright office charges $165 per hour) and often does not solidly determine who holds the rights at this moment in time. Most of our action around intellectual property rights is based on claims and suppositions, not facts, and we often act as if there were evidence of held rights even though we have no such proof. In contrast, the patent system is fully documented with descriptions and drawings and references to other patents, although by its own admission the patent office has about a 3-year backlog, and filing and researching patents is time-consuming and expensive.

Another interesting aspect of the GSU case is that some of the works being copied were not covered by any available licensing scheme. This is especially interesting since the publisher plaintiffs were backed by the Copyright Clearance Center, presumably the agency that one would turn to when desiring to license a work for use. Licensing is a relatively big business: CCC earns over $200 million per year. The publishers included in this case each earn something shy of $500K per year in fees from CCC licensing. Much of that, however, comes from the commercial printing of course packs, not from direct educational institution use. The judge determined that the percentage of publisher revenue from electronic course content would be .00046 (five one-hundredths of one percent) of the average net revenue for any one of the publishers.

Reading this it becomes rather obvious that the move from traditional course packs, which are produced by commercial copy shops, to digital course readings, which are produced by the library or the professor, would mean a loss of revenue for the academically-oriented publishers. Course packs got slammed by the copyright holders not because copies were being made but because copies were being sold and all profit was going to the copy shops, none to the rights holders. In fact, in this lawsuit there were files on reserve that were never downloaded by students in the class, and the judge removed these titles from the suit because they were not read. This is an interesting answer to the question: "What if you make a copy and no one sees it?" Another way of wording this is: "Is it a copy if it has only been viewed by a computer?" With course packs, every student purchases every item in the course pack and you have no idea if any of those are read. With digital copies, every download can be counted. Although a download does not guarantee that the item has been read by the downloadee, it is a quantifiable use in the same way that the number of course packs printed is quantifiable. A file online does not seem, in and of itself, to be the same as a physical copy. This could have implications for library digitization projects, and relates to the decision in the Authors Guild v. HathiTrust case.

There are some gotchas to the use of a copyright licensing agency because of the inherent nature of the US fair use law. When one approaches CCC to license a work there is no fair use determination that is made as part of that request. It is up to the requestor to decide whether a license is needed or not. There are annual licenses available for educational institutions that cover a set of materials. None of these licenses are needed if the use is a fair use, and for educational institutions many uses are indeed fair uses, in particular classroom use. Therefore the CCC annual educational license may be paying for uses that do not require payment under copyright law. In essence, the license may be seen as a kind of insurance policy against infringement claims, but it may not be money well-spent. As the judge in the GSU case states (p. 66) "In the absence of judicial precedent concerning the limits of fair use for nonprofit educational uses, colleges and universities have been guessing about the permissible extent of fair use."

The decision itself runs to 350 pages, much of which is taken up with the decisions about the 74 documents in question. The judge does a very nice job of defining the nature of a work, and why individual chapters are viable on their own as part of a course syllabus. The decision that 10% of a "work" is permissible makes me wonder, however, if publishers won't see the light and begin digital publishing of individual chapters rather than creating book-length anthologies. 

More legal analysis is available on James Grimmelman's blog.

Saturday, October 06, 2012

Google and publishers settle

Recent news announcements state that the AAP and Google have settled their lawsuit. Essentially this is a formality rather than an actual change in status between Google and the publishers. As you probably recall, the AAP engaged in a lawsuit against Google, in partnership with the Author's Guild, in 2005, and five specific publishers were named in that suit. Since then the lawsuit has gone through two failed attempts to settle and a massive number of pages of legal parrying. Meanwhile, the publishers realized that Google provided them with a new sales opportunity and about 40,000 of them have entered into agreements with Google in which the publisher provides (or allows Google to provide) a digital copy of the book which then can be sold as a Google eBook. Each contract specifies the amount of the book that potential buyers can browse. This browsing is designed to prevent users from satisfying their needs without a purchase: although up to 20% of a digital book may be browsable, Google denies access to sequences of more than 9 pages in a row.

When the Author's Guild revived their suit against Google in late 2011, the AAP was notably absent from the list of plaintiffs. By then the publishers had made a satisfactory agreement with Google and had no interest in suing their business partner. This current settlement appears to be a pro forma legal action to terminate the previous lawsuit, and, as far as I can tell, makes no change to the current status of the business relationship between Google and publishers.

There is one question that remains which is what this means in terms of copyright, if anything. The contractual arrangements between Google and the publishers are standard business agreements which the publishers engage in as representatives of the rights holder through their contract with that rights holder. Sometimes the publisher also holds the copyright, but that isn't the salient point here. So unless I am missing something, this agreement has absolutely no effect on the questions of copyright and digitization that some of us have been so eager to hear about. It's just business as usual.

[After-note: Some authors and journalists question the settlement and want details made public.]

Monday, September 24, 2012

Library signage

After all of the hoopla about libraries converting to BISAC bookstore categories instead of using the Dewey Decimal System, a trip to Barnes and Noble one day last week made me wonder if it's really the categories that matter, or if it's all about the signage.
Here's some recent signage at Barnes and Noble:


Here's what the library signage in my local library looks like:


Which do you think is understood best by the people who step into those institutions?
I've referred to library cataloging as "the secret language of twins," understood by a small in-crowd and completely unknown to others. This library signage is even worse than that; it's as if the library decided to encrypt its subject access, and won't let the users have the key. There is no copy of DDC in the library for users to consult. (I know this because I looked for it.) You can get to a place on the shelf by doing a search in the catalog, but you can't find out what the numbers mean, and there is no natural language translation given in the library, other than "Fiction" and "Non-Fiction" over the doors to the main shelf areas.

How could this possibly be seen as functional?

Tuesday, September 18, 2012

Seven years, and waiting

Just a quick note to say that the Google Book Search lawsuit has entered a new "pause" and will thus be delayed further. Among the salvos that Google and the Author's Guild have fired back and forth are those around the question of whether the Author's Guild can legally represent all authors against Google. As I recall, the Author's Guild has something on the order of 5,000 members, all alive (as far as I know), while Google's digitization now covers somewhere upwards to 10 million books and probably nearly as many authors, from the days of early printed works to the present.

Google challenged the Author's Guild as class representative of all authors, but Judge Denny Chin, the judge who has seen this case through most of its life, allowed the suit to go forward with the Guild as class representative. This has been reversed by an appeals court judge, which means that the question of class representation must be decided before the suit can continue. 

Friday, September 14, 2012

Rich snippets

At the recent Dublin Core annual meeting I heard Dan Brickley talk about Google's use of schema.org for rich snippets. Schema.org is commonly thought of as "search engine optimization" (SEO), which to most people means "how to get onto the first page of a Google results search." But the microdata in web sites can also be used to make the snippets shown more useful by incorporating more information from the web page. The examples above, from the Google rich snippets page, show features like ratings as well as links to actual content within the web page.

Now that WorldCat has schema.org markup, my first thought was: what kind of rich snippet would be good for library data? There is a rich snippet testing tool where you can plug in a URL and see 1) the snippet 2) what microdata is visible to Google. You can plug in a WorldCat permalink and see what the rich snippet result is:

http://www.worldcat.org/oclc/874206  (opens in separate window)

There is no rich snippet displayed here, which tells us that Google hasn't yet developed a rich snippet model for our kind of data. But you can see, in great detail, all of the coded data that is available. (The red warnings indicate that there is data in the OCLC microdata that isn't part of schema.org. OCLC is talking to the schema.org developers to incorporate new elements, some of which show up as warnings here.)

I began to think about how I would like this data used. It could be used to format a more bibliographic-like display, adding author, publisher, pagination. The ISBN could of course link to key online bookstores. (That would also bring in revenue for Google, so might be a popular choice for the search engine.) But what about libraries? How could rich snippets help libraries and library users?

The snippet could  lead back to WorldCat where the user could find a nearby library, but... wait! Google often knows your approximate location, and WorldCat knows whether libraries in your area have the book. AND the library catalog often has information about availability. I don't know how this data would interact with the WorldCat tool, but here's what I would like to see in the snippet:


This definitely goes beyond what "rich snippet" means today, but is not inconsistent with retrievals that pull data from multiple online sales outlets.  In the sales model, Google's assumption is that the searcher wants to obtain (in FRBR-speak) the item, and therefore various outlets that could provide that data are listed. This same logic could apply to libraries, of course. Libraries are a local source of many of the same things that are sold online, so the obtain logic fits.

This analysis of mine obviously ignores the economic incentive for Google to provide library holdings, especially since they would be seen as competing with sales.  I'm just dreaming here, doing the "what if" thing without the practical limitations.






Saturday, August 11, 2012

The success paradox

An article entitled "Study: Public Awareness Gap on Ebooks in Libraries" in the July/August 2012 issue of American Libraries reports that 62% of Americans polled did not know if their library (presumably their local public library) lends ebooks. This statement was followed by a quote from Molly Raphael about how libraries should increase public awareness of their services. I would naturally be inclined to agree with her except for the other statistics that were cited: 56% of those who do borrow ebooks were unable to borrow a particular book they were seeking, and 52% had encountered wait lists for books.

Given that data, you have to wonder what the results would be if libraries did make the public more aware of the availability of ebooks.

An institution with a fixed budget cannot afford to be too successful, or at least not successful in the sense of encouraging more use of its services. The more successful the institution is, the more it will fail its users because demand will overwhelm its ability to serve them. As we see in the ebook example, libraries are failing to serve well even the minority of people who are aware that they have ebooks to lend. What if everyone was made aware of the availability of "free" ebooks from the library? The ebook lending service would be worse from the user's perspective.

Where each purchase of the book makes money for the bookseller, each demand of the library results in a cost rather than a revenue gain. This is because the library is on a fixed income and the basis of the library's budget is only tangentially related to the number of 'customers' it serves. A library with an annual budget of, say, $500,000 has that amount to spend even if use of the library increases greatly during that year. Such an increase of use does not guarantee that the library's budget will be increased when the next fiscal year's budget is decided on by the governing authority (such as city, state, or college). In fact, as we have seen in these hard economic times, increased use and decreased budgets can go hand-in-hand.

Because the library's model is to get more use out of a limited number of books (and DVDs and other items) by lending them sequentially to patrons, the direct result of high demand is an increase in the failure of the library to meet the demand, evidenced by long waiting lists for the book and patrons who are unhappy with the library's service. My local library in the city of Berkeley today has nine copies of "50 Shades of Grey" with 70 holds. This is in a small city with a population of about 120,000. The library of the city of Santa Clara in Silicon Valley, which is similar in size to Berkeley, also has nine copies, but 119 holds. New York Public Library has "Holds: 1657 on 131 copies."

In this sense, success -- that is, many people turning to the library with their desire to read a highly popular book -- is in fact the cause of failure; the failure of the library to meet that demand.

I come around to these thoughts when I find myself frustrated at the reluctance or inability of libraries to promote their services despite obvious opportunities to do so. Then I think about what it looks like inside a branch of my local library, with fewer and fewer staff available and the obvious strain, as evidenced by cart after cart of books that have been checked in but not yet returned to the shelves, by long lines to use the public computers despite the fact that those now take up a significant amount of floor space, and shelf after shelf of sadly worn trade paperbacks on topics that were fleeting fads a decade or so ago.

Really, why would an institution so stretched in its resources want to stimulate more demand?

Libraries, of course, are not the only such institutions; few if any inner-city emergency rooms would consider it a good idea to stimulate the arrival of more patients. I've been in a higher end hospital emergency room in my town (fortunately very seldom) and even that had patients on gurneys in the hallways because every more appropriate space was filled. Again here, high demand promotes failure.

Libraries from their beginnings developed in response to scarcity. It would not be unreasonable to suggest that the library model based on scarcity is not well suited to the current climate of media abundance. Yet, any public institution that would base its services on scarce but rarely sought goods is on a suicide mission, especially in today's economic and social climate.

Is there a solution to this dilemma? In a perfect world (obviously not the one we live in) government and institutions of higher learning would recognize the value of a well-stocked, vibrant community information space. Instead, library budgets are being sharply cut, and library services are not perceived has having high value. If libraries have lost support among their traditional communities, it may have something to do with the current measures of success: number of items checked-out being the primary one for public libraries; number of volumes owned for research libraries.

Whatever success looks like in the future, it simply cannot be based on increasing the number of holds on materials or on providing services that you hope only a few people will discover and make use of. Circulation figures should not be the main measure of the library's value; we need not only new services but new measures of the library's worth to the community. There is a pressing need for actual information services, not just the storage and circulation of items.There is also a desire for participation, as evidenced by social information networking like Wikipedia, LibraryThing, GoodReads, and others. Perhaps the answer lies in asking what the users can contribute to the library so that use adds value as well as incurs costs. Perhaps the library of today should serve its users by giving them a means to crowd-source solutions to their information problems, with the library contributing knowledge organization skills, its awareness of community needs, and a commitment to quality information services. Maybe the library of the future should be less about circulating books and DVDs and more about helping people make sense of the information glut that we live in; less about keeping up with global bestsellers and more about learning with and within the community.

Then maybe success can be success, not failure.