Showing posts with label Digital libraries. Show all posts
Showing posts with label Digital libraries. Show all posts

Monday, April 27, 2020

Ceci n'est pas une Bibliothèque

On March 24, 2020, the Internet Archive announced that it would "suspend waitlists for the 1.4 million (and growing) books in our lending library," a service they then named The National Emergency Library. These books were previously available for lending on a one-to-one basis with the physical book owned by the Archive, and as with physical books users would have to wait for the book to be returned before they could borrow it. Worded as a suspension of waitlists due to the closure of schools and libraries caused by the presence of the coronavirus-19, this announcement essentially eliminated the one-to-one nature of the Archive's Controlled Digital Lending program. Publishers were already making threatening noises about the digital lending when it adhered to lending limitations, and surely will be even more incensed about this unrestricted lending.

I am not going to comment on the legality of the Internet Archive's lending practices. Legal minds, perhaps motivated by future lawsuits, will weigh in on that. I do, however, have much to say on the use of the term "library" for this set of books. It's a topic worthy of a lengthy treatment, but I'll give only a brief account here.

LIBRARY … BIBLIOTHÈQUE … BIBLIOTEK


The roots “LIBR…” and “BIBLIO…” both come down to us from ancient words for trees and tree bark. It is presumed that said bark was the surface for early writings. “LIBR…”, from the Latin word liber meaning “book,” in many languages is a prefix that indicates a bookseller’s shop, while in English it has come to mean a collection of books and from that also the room or building where books are kept. “BIBLIO…” derives instead from the Greek biblion (one book) and biblia (books, plural). We get the word Bible through the Greek root, which leaked into old Latin and meant The Book.

Therefore it is no wonder that in the minds of many people, books = library.  In fact, most libraries are large collections of books, but that does not mean that every large collection of books is a library. Amazon has a large number of books, but is not a library; it is a store where books are sold. Google has quite a few books in its "book search" and even allows you to view portions of the books without payment, but it is also not a library, it's a search engine. The Internet Archive, Amazon, and Google all have catalogs of metadata for the books they are offering, some of it taken from actual library catalogs, but a catalog does not make a quantity of books into a library. After all, Home Depot has a catalog, Walmart has a catalog; in essence, any business with an inventory has a catalog.
"...most libraries are large collections of books, but that does not mean that every large collection of books is a library."

The Library Test

First, I want to note that the Internet Archive has met the State of California test to be defined as a library, and this has made it possible for the Archive to apply for library-related grants for some of its projects. That is a Good Thing because it has surely strengthened the Archive and its activities. However, it must be said that the State of California requirements are pretty minimal, and seem to be limited to a non-profit organization making materials available to the general public without discrimination. There doesn't seem to be a distinction between "library" and "archive" in the state legal code, although librarians and archivists would not generally consider them easily lumped together as equivalent services.

The Collection

The Archive's blog post says "the Internet Archive currently lends about as many as a US library that serves a population of about 30,000." As a comparison, I found in the statistics gathered by the California State Library those of the Benicia Public Library in Benicia California. Benicia is a city with a population of 31,000; the library has about 88,000 books. Well, you might say, that's not as good as over one million books at the Internet Archive. But, here's the thing: those are not 88,000 random books, they are books chosen to be, as far as the librarians could know, the best books for that small city. If Benicia residents were, for example, primarily Chinese-speaking, the library would surely have many books in Chinese. If the city had a large number of young families then the children's section would get particular attention. The users of the Internet Archive's books are a self-selected (and currently un-defined) set of Internet users. Equally difficult to define is the collection that is available to them:
This library brings together all the books from Phillips Academy Andover and Marygrove College, and much of Trent University’s collections, along with over a million other books donated from other libraries to readers worldwide that are locked out of their libraries.
Each of these is (or was, in the case of Marygrove, which has closed) a collection tailored to the didactic needs of that institution. How one translates that, if one can, to the larger Internet population is unknown. That a collection has served a specific set of users does not mean that it can serve all users equally well. Then there is that other million books, which are a complete black box.

Library science

I've argued before against dumping a large and undistinguished set of books on a populace, regardless of the good intentions of those doing so. Why not give the library users of a small city these one million books? The main reason is the ability of the library to fulfill the 5 Laws of Library Science:
  1. Books are for use.
  2. Every reader his or her book.
  3. Every book its reader.
  4. Save the time of the reader.
  5. The library is a growing organism. [0]
The online collection of the Internet Archive nicely fulfills laws 1 and 5: the digital books are designed for use, and the library can grow somewhat indefinitely. The other three laws are unfortunately hindered by the somewhat haphazard nature of the set of books, combined with the lack of user services.

Of the goals of librarianship, matching readers to books is the most difficult. Let's start with law 3, "every book its reader." When you follow the URL to the National Emergency Library, you see something like this:
The lack of cover art is not the problem here. Look at what books you find: two meeting reports, one journal publication, and a book about hand surgery, all from 1925. Scroll down for a bit and you will find it hard to locate items that are less obscure than this, although undoubtedly there are some good reads in this collection. These are not the books whose readers will likely be found in our hypothetical small city. These are books that even some higher education institutions would probably choose not to have in their collections. While these make the total number of available books large, they may not make the total number of useful books large. Winnowing this set to one or more (probably more) wheat-filled collections could greatly increase the usability of this set of books.

"While these make the total number of available books large, they may not make the total number of useful books large."

A large "anything goes" set of documents is a real challenge for laws 2 and 4: every reader his or her book, and save the time of the reader. The more chaff you have the harder it is for a library user to find the wheat they are seeking. The larger the collection the more of the burden is placed on the user to formulate a targeted search query and to have the background to know which items to skip over. The larger the retrieved set, the less likely that any user will scroll through the entire display to find the best book for their purposes. This is the case for any large library catalog, but these libraries have built their collection around a particular set of goals. Those goals matter. Goals are developed to address a number of factors, like:
  • What are the topics of interest to my readers and my institution?
  • How representative must my collection be in each topic area?
  • What are the essential works in each topic area?
  • What depth of coverage is needed for each topic? [1]
If we assume (and we absolutely must assume this) that the user entering the library is seeking information that he or she lacks, then we cannot expect users to approach the library as an expert in the topic being researched. Although anyone can type in a simple query, fewer can assess the validity and the scope of the results. A search on "California history" in the National Emergency Library yields some interesting-looking books, but are these the best books on the topic? Are any key titles missing? These are the questions that librarians answer when developing collections.

The creation of a well-rounded collection is a difficult task. There are actual measurements that can be run against library collections to determine if they have the coverage that can be expected compared to similar libraries. I don't know if any such statistical packages can look beyond quantitative measures to judge the quality of the collection; the ones I'm aware of look at call number ranges, not individual titles.  There

Library Service


The Archive's own documentation states that "The Internet Archive focuses on preservation and providing access to digital cultural artifacts. For assistance with research or appraisal, you are bound to find the information you seek elsewhere on the internet." After which it advises people to get help through their local public library. Helping users find materials suited to their need is a key service provided by libraries. When I began working in libraries in the dark ages of the 1960's, users generally entered the library and went directly to the reference desk to state the question that brought them to the institution. This changed when catalogs went online and were searchable by keyword, but prior to then the catalog in a public library was primarily a tool for librarians to use when helping patrons. Still, libraries have real or virtual reference desks because users are not expected to have the knowledge of libraries or of topics that would allow them to function entirely on their own. And while this is true for libraries it is also true, perhaps even more so, for archives whose collections can be difficult to navigate without specialized information. Admitting that you give no help to users seeking materials makes the use of the term "library" ... unfortunate.

What is to be done?


There are undoubtedly a lot of useful materials among the digital books at the Internet Archive. However, someone needing materials has no idea whether they can expect to find what they need in this amalgamation. The burden of determining whether the Archive's collection might suit their needs is left entirely up to the members of this very fuzzy set called "Internet users." That the collection lends at the rate of a public library serving a population of 30,000 shows that it is most likely under-utilized. Because the nature of the collection is unknown one can't approach, say, a teacher of middle-school biology and say: "they've got what you need." Yet the Archive cannot implement a policy to complete areas of the collection unless it knows what it has as compared to known needs.

"... these warehouses of potentially readable text will remain under-utilized until we can discover a way to make them useful in the ways that libraries have proved to be useful."

I wish I could say that a solution would be simple - but it would not. For example, it would be great to extract from this collection works that are commonly held in specific topic areas in small, medium and large libraries. The statistical packages that analyze library holdings all are, AFAIK, proprietary. (If anyone knows of an open source package that does this, please shout it out!) If would also be great to be able to connect library collections of analog books to their digital equivalents. That too is more complex than one would expect, and would have to be much simpler to be offered openly. [2]

While some organizations move forward with digitizing books and other hard copy materials, these warehouses of potentially readable text will remain under-utilized until we can discover a way to make them useful in the ways that libraries have proved to be useful. This will mean taking seriously what modern librarianship has developed over its circa 2 centuries, and in particular those 5 laws that give us a philosophy to guide our vision of service to the users of libraries.

-----

[0] Even if you are familiar with the 5 laws you may not know that Ranganathan was not as succinct as this short list may imply. The book in which he introduces these concepts is over 450 pages long, with extended definitions and many homey anecdotes and stories.

[1] A search on "collection development policy" will yield many pages of policies that you can peruse. To make this a "one click" here are a few *non-representative* policies that you can take a peek at:
[2] Dan Scott and I did a project of this nature with a Bay Area public library and it took a huge amount of human intervention to determine whether the items matched were really "equivalent". That's a discussion for another time, but, man, books are more complicated than they appear.

Tuesday, October 10, 2017

Google Books and Mein Kampf

I hadn't look at Google Books in a while, or at least not carefully, so I was surprised to find that Google had added blurbs to most of the books. Even more surprising (although perhaps I should say "troubling") is that no source is given for the book blurbs. Some at least come from publisher sites, which means that they are promotional in nature. For example, here's a mildly promotional text about a literary work, from a literary publisher:



This gives a synopsis of the book, starting with:

"Throughout a single day in 1892, John Shawnessy recalls the great moments of his life..." 

It ends by letting the reader know that this was a bestseller when published in 1948, and calls it a "powerful novel."

The blurb on a 1909 version of Darwin's The Origin of Species is mysterious because the book isn't a recent publication with an online site providing the text. I do not know where this description comes from, but because the  entire thrust of this blurb is about the controversy of evolution versus the Bible (even though Darwin did not press this point himself) I'm guessing that the blurb post-dates this particular publication.


"First published in 1859, this landmark book on evolutionary biology was not the first to deal with the subject, but it went on to become a sensation -- and a controversial one for many religious people who could not reconcile Darwin's science with their faith."
That's a reasonable view to take of Darwin's "landmark" book but it isn't what I would consider to be faithful to the full import of this tome.

The blurb on Hitler's Mein Kampf is particularly troubling. If you look at different versions of the book you get both pro- and anti- Nazi sentiments, neither of which really belong  on a site that claims to be a catalog of books. Also note that because each book entry has only one blurb, the tone changes considerably depending on which publication you happen to pick from the list.


First on the list:
"Settling Accounts became Mein Kampf, an unparalleled example of muddled economics and history, appalling bigotry, and an intense self-glorification of Adolf Hitler as the true founder and builder of the National Socialist movement. It was written in hate and it contained a blueprint for violent bloodshed."

Second on the list:
"This book has set a path toward a much higher understanding of the self and of our magnificent destiny as living beings part of this Race on our planet. It shows us that we must not look at nature in terms of good or bad, but in an unfiltered manner. It describes what we must do if we want to survive as a people and as a Race."
That's horrifying. Note that both books are self-published, and the blurbs are the ones that I find on those books in Amazon, perhaps indicating that Google is sucking up books from the Amazon site. There is, or at least at one point there once was, a difference between Amazon and Google Books. Google, after all, scanned books in libraries and presented itself as a search engine for published texts; Amazon will sell you Trump's tweets on toilet paper. The only text on the Google Books page still claims that Google Books is about  search: "Search the world's most comprehensive index of full-text books." Libraries partnered with Google with lofty promises of gains in scholarship:
"Our participation in the Google Books Library Project will add significantly to the extensive digital resources the Libraries already deliver. It will enable the Libraries to make available more significant portions of its extraordinary archival and special collections to scholars and researchers worldwide in ways that will ultimately change the nature of scholarship." Jim Neal, Columbia University
I don't know how these folks now feel about having their texts intermingled with publications they would never buy and described by texts that may come from shady and unreliable sources.

Even leaving aside the grossest aspects of the blurbs and Google's hypocrisy about its commercialization of its books project, adding blurbs to the book entries with no attribution and clearly not vetting the sources is extremely irresponsible. It's also very Google to create sloppy algorithms that illustrate their basic ignorance of the content their are working with -- in this case, the world's books.

Thursday, November 14, 2013

It's FAIR!

"In my view, Google Books provides significant public benefits. It advances the progress of the arts and sciences, while maintaining respectful consideration for the rights of authors and other creative individuals, and without adversely impacting the rights of copyright holders. It has become an invaluable research tool that permits students, teachers, librarians, and others to more efficiently identify and locate books. It has given scholars the ability, for the first time, to conduct full-text searches of tens of millions of books. It preserves books, in particular out-of-print and old books that have been forgotten in the bowels of libraries, and it gives them new life. It facilitates access to books for print-disabled and remote or underserved populations. It generates new audiences and creates new sources of income for authors and publishers. Indeed, all society benefits." p. 26
With that statement, Judge Denny Chin has ruled (PDF) that Google's digitization of books from libraries is a fair use.  And a very long saga ends.

Google was first brought to court in 2005 by the Author's Guild in a copyright infringement suit for its mass digitization of library holdings. Since then the matter has gone back to the court a number of times. Most significantly, Google, authors, and publishers developed two complex proposed settlements that were, however, so fraught with problems that the Department of Justice weighed in. Finally, the publishers bowed out and the original Author's Guild suit was revived. At that point, the question became: Is Google's digitization of books for the purposes of indexing (and showing snippets as search results) fair use?

Of course, much happened between 2005 and 2013. One important thing that happened was the development of HathiTrust, the digital repository where libraries can store the digital copies that they received from Google of their own books. The same Authors Guild sued HathiTrust for copyright infringement, but Judge Baer in that case decided for fair use.

I cannot over-emphasize either the role of libraries in this case nor the support that both judges expressed for libraries and for their promotion of "progress and the useful arts." Chin refers frequently to the amicus brief (PDF) presented by the American Library Association, as well as the conclusions in the HathiTrust case. Both judges clearly admire the mission of libraries, and it seems clear to me that the educational use of the materials by libraries was seen to offset the for-profit use by Google. In fact, Judge Chin reverses the roles of Google and the libraries when he says:
"Google provides the libraries with the technological means to make digital copies of books that they already own. The purpose of the library copies is to advance the libraries' lawful uses of the digitized books consistent with the copyright law." p. 26
In those terms, Google has simply helped libraries do what they do, better. Google's digitization of the library books is thus a public service.
"Google Books helps to preserve books and give them new life. Older books, many of which are out-of-print books that are falling apart buried in library stacks, are being scanned and saved." p. 12
Note that Google and the libraries (in HathiTrust) are exceedingly careful to stay within the letter of the law. Google's snippet display algorithm is rococo in design, making it literally impossible to reconstruct a book from the snippets it displays. So much so that it would probably take less time to re-scan the book at home on your page-at-a-time desktop scanner.

The full impact of this ruling is impossible (for me) to predict, but there are many among us who are breathing a great sigh of relief today. This opens the door for us to rethink digital scholarship based on materials produced before information was in digital form. 

I do have a wishlist, however, and at the top of that is for us to turn our attention to making the digitized texts even more useful by turning that uncorrected OCR into a more faithful reproduction of the original book. While large-scale linguistic studies may be valid in spite of a small percentage of errors, the use of the digitized materials for reading, in the case of those works in the public domain, and for listening, in the case of works made available to VIPs (visually impaired persons), is greatly hampered by the number and kinds of errors that result. In a future post I will give the results of a short study that I have done in that area.

See all my posts on Google Books

Monday, September 12, 2011

Authors Guild Sues HathiTrust

There has been a period of limbo since Judge Chin rejected the proposed settlement between the Author's Guild/Association of American publishers and Google. In fact, a supposedly final meeting between the parties is scheduled for this Thursday, 9/15, in the judge's court.

Monday, 9/12, the Author's Guild (and partners) filed suit against HathiTrust (and partners) for some of the same "crimes" of which it had accused Google: essentially making unauthorized copies of in-copyright texts. In addition, the recent announcement that the libraries would allow their users to access items that had been deemed to be orphan works figures in the suit. That this suit has come over 6 years since the original suit against Google is in itself interesting. Nearly all of the actions of HathiTrust and its member libraries fall within what would have been allowed if the agreement that came out of that suit had been approved by the court. Although we do not know the final outcome of that suit (and anxiously await Thursdays meeting to see if it is revelatory), this suit against the libraries is surely a sign that AG/AAP and Google have not come to a reconciliation.

The Suit

First, the suit establishes that the libraries received copies of Google-digitized items from Google, and have sent copies of these items to HathiTrust, which in turn makes some number of copies as part of its archival function. This is followed by a somewhat short exposition of the areas of copyright law that are pertinent, with an emphasis on section 108, which allows libraries to make limited copies to replace deteriorating works. The suit states that the copying being done is not in accord with section 108. Then it refers to the Orphan Works Project that several libraries are partnering in, and the plan on the part of the libraries to make the full text of orphan works available to institutional users.

Since most of these institutions (if not all of them) are state institutions that have protections against paying out large sums in a lawsuit of this nature, the goal is to regain the control of the works by forcing HathiTrust (and the named libraries) to transfer their digital copies of in-copyright works to a "commercial grade" escrow agency with the files held off network "pending an appropriate act of Congress."

As James Grimmelman comments in his blog post on the suit, there's a lot of mixing up between the orphan works and owned works in the suit. He points out that a group of organizations representing authors could hardly make a case for orphan works since, by definition, the lack of ownership of the orphans means they can't be represented by a guild of people defending their own works.

The Problems

There are numerous problems that I see in the text of this suit. (IANAL, just a Librarian.)
  • The suit mentions large numbers of books that have been copied without permission, but makes no attempt to state how many of those books belong to the members of the plaintiff organizations.
  • The suit throws around large numbers without clearly stating that none of the statements include Public Domain works. It isn't clear, therefore, what the numbers represent: the entire holdings of HathiTrust, or just the in-copyright holdings. Also, in relation to the latter, unless one has done a considerable amount of work there are many works that are post-1923 that are also in the Public Domain. Cutting off at that year does not account for works that were not renewed, or were never copyrighted. I also doubt if anyone has a clear idea how many of the works in question are Public Domain because they are US Federal documents. This imprecision on the copyright status of works is very frustrating, but HathiTrust is not to blame for this state of affairs.
  • Some of their claims do not seem to me to be within legal bounds. For example, in one section they claim that although HathiTrust is not giving users access to in-copyright works, they potentially could. Where does that fit in?
  • They also claim that there is a risk of unauthorized access. However, the security at HathiTrust meets the security standards that the Author's Guild agreed to in the (unapproved) settlement with Google. If it was good enough then, why is it now too risky?
  • They claim that the libraries themselves have been digitizing in-copyright books. I wasn't aware of this, and would like to know if this is the case.
  • They state that the libraries said that before Google it was costing them $100 a book for digitization. Then the plaintiffs say that this means that the value of the digital files is in the hundreds of millions of dollars. First, I have heard figures that are more like $30 a book. Second, I don't see how the cost to digitize can translate into a value that is relevant to the complaint.
  • Although the legislature has failed to pass an orphan works law that would allow the use of these materials and still benefit owners if they do come forth, it seems like a poor strategy to complain about a well-designed program of due diligence and notification, which is what the libraries have designed. Orphan works are the least available works: if you have an owner you can ask permission; if there is no owner you cannot ask permission and therefore there is no way to use the work if your use falls outside of fair use. It's hard to argue for taking these works entirely out of the cultural realm simply because we have a poorly managed copyright ownership record.
  • There are a few odd sections where they make reference to bibliographic data as though that were part of the "unauthorized digitization" rather than data that was created by and belongs to the libraries. There's an odd attempt to make bibliographic data searching seem nefarious.
Parties
Plaintiffs: The Author's Guild, Inc.; The Australian Society of Authors limited ; Union des Erivaines Quebecois; Pat Cummings; Angelo Loukakis; Roxana Robinson; Andre Roy; James Shapiro; Daniele Simpson; T.J. Stiles; and Fay Weldon. (Links are to some sample HathiTrust records.)

Defendants: HathiTrust; The Regents of the University of California, The Board of Regents of the University of Wisconsin System; The Trustees of Indiana University; and Cornell University.

Links
Boing Boing: Authors Guild declares war on university effort to rescue orphaned books
Library Journal: Copyright Clash

Monday, August 01, 2011

Suggestions for HathiTrust UI

Here are my concrete suggestions for improvements to the HathiTrust user interface. This is based on my own experience and should not be considered to be complete or universal. These are simply the things that would have made my experience better:

On the home page, there should be two links:
  • member login
  • guest login
By each there should be a link to help (one of those question mark circles, for example).

Member help will explain: that you must be someone associated with one of these institutions (link) with an institutional id. Members can: [whatever they can do - view everything, download all PD materials, create bibliographies...]

Guest help will explain: that HT is a member-sponsored db. Guests can search and can view the full text some materials. A guest account allows you to create a persistent bibliography.

On the page for a work, do NOT say: Public domain, Google-digitized. Instead, say what the user needs to know:
Public domain; member-only download.
Public domain; anyone can download.

If you ask for a login at the time of download, ONLY ask for a member login since a guest login does not provide access at this point. The message ("member-only download") may be enough, but the login request could read: requires member login.

This was as far as I got in HT, and I'm not going to be spending much more time there, since as a non-member I am actually served better on other sites. It's a superficial look from a first-time, non-member user.

Sunday, July 31, 2011

User-friendliness, a lesson

I was looking for Melvil Dewey's first published version of his classification system. My first instinct was to head to Google Book Search but I decided instead to use HathiTrust as a kind of gesture to non-commercial access. I did find what I was looking for, his 1876 pamphlet, opened it up in their reader and looked through it. I knew I'd want a copy, so I found the "download as PDF" link. That popped up a box telling me to "Login to determine whether you can download this book." The copyright is listed as "public domain in the United States." I don't see why I need to log in, and I downloaded it from GBS instead, without logging in but adding to the slime trail of my life that Google owns. The added step of logging in (to be started by creating yet another login on a system I will use only occasionally), for all that it may be no more or even less invasive of my privacy, is not user-friendly. It also didn't make sense to me at the time, and I was given nothing to convince me that logging in was beneficial ... to me.

Yes, it's all about ME, me the user, me the person at the other end of the connection. I'm also not just any user, I am an advocate of libraries, a librarian, and I made the effort to go to HathiTrust -- a site that has not shown up for me in search engines.

This seems to be such a basic lesson that I do not understand why libraries can't learn it. User-friendliness.


Ooof! It just gets worse. I decided to see what login is about. To get to login, you have to search, select a book, and click on login. On the book page, you may see that a book is "Public Domain" or it may say "Public Domain, Google-digitized". When you log in, you log in either as someone from a member institution or a guest. The guest log in form states:
Does NOT provide access to full PDF downloads of public domain & open access items where not publicly available
However, it turns out that it DOES provide access to PD books (see comment by anonymous) if the book is not digitized by Google -- but that isn't what you've been told. "... not publicly available" isn't what you see on the book page, you see "Google-digitized." The page on policies has two different categories, "Open Access" and "Open Access, Google-digitized." Nothing in the definitions of those categories mentions member and guest downloading.

Basically, HathiTrust turns out to be a tiered system with member and non-member access. You don't encounter this until you try to download something that is PD but not "publicly available." Nothing on the home page mentions that this is a member-based service, therefore you don't know that as a non-member you will encounter walls.

OK, it is resolved, that from now on I will always go first to the Open Library, a site where Open means what I think it should.

Wednesday, July 20, 2011

Unequal Access

With the recent indictment of an advocate for open information access who had set up a way to download about 4 million JSTOR articles, presumably with the intent to liberate them from their native closed access, we need to step back and look at how unequal information access is in this world. In major universities in the US, academics and students log on to their computers in their offices or at home and a whole world opens up to them. That's not some kind of accident. The prime goal of university libraries is to make good on "seek and ye shall find." The proof of the success of these libraries is that researchers are oblivious to the complexity of the system that serves them. I would guess that many members of the US university community have no idea how their access to journals is managed and controlled. They don't see the contract negotiations with information providers, the continual development of software that makes single-point searching possible, the multi-faceted delivery systems that blend (or attempt to) digital and paper resources into a single stream. And they don't think about how different it would be if they weren't members of that privileged community.

Contrast that to the access available to a member of the US public who is not part of this academic sector. Like myself. Like the majority of people in this country. There is no access to JSTOR. No openURL server gives me multiple access options. The local public library does have some electronic materials, but these are much less extensive (and less expensive) than the ones in academic libraries. I may have to wait weeks to get a book that isn't in my local library's collection, if I can get it at all. I am often in the embarrassing position of not being able to access articles that I would like to read or quote from, including ones that I myself have authored.

In spite of this, I know that my information access, as a mere member of the US public, is far superior to that found in other countries; countries where serious researchers struggle to participate in research because they do not have the access that many academics here take for granted. Two anecdotes:

-- When I lived in Italy in the 1970's my friends were mainly college students or recent graduates. University education was free, but it was generally accepted that the only way to complete ones final thesis was to be able to afford to go abroad for two or three months. The purpose of this trip was to spend time in a country with a good library system, since libraries in Italy were limited. This was not just for students studying foreign literatures, but even those studying sciences, history, and art. These kids were essentially "library tourists." I don't know if this continues today.

-- During the time I worked at UC I was in a conversation with someone involved in the licensing of databases. For some reason we got talking about enforcement of contractual clauses having to do with excessive downloading and/or piracy. This person told me that all access to one of the UC campuses had been cut off recently for a few days because it was discovered that someone was systematically downloading entire journal runs. When they found the student it turned out that it was a foreign graduate student who would soon be returning home. Knowing that leaving the UC system would mean losing access to the journals he would need to continue his research, he was making himself a copy to take home.

It occurs to me as I write this that the "Digital Public Library of America" could create an information revolution in this country by upgrading the access of the general public to that of an academic or student in a large college or university, without ever digitizing a single page. What makes Stanford "Stanford" or Harvard "Harvard" is not just its famed faculty but the full range of information that is shared by that community. Everything they do, every bit of research, every new idea, is facilitated by the library and its services.

The information access gap between a university researcher and the average person on the street is immense. We have an information elite that, like most elites, considers its position to be earned, just, and reasonable. Few in academia worry that the access they have isn't widely shared. If they did, they would hopefully decide that something should be done.

Tuesday, May 31, 2011

All the ____ in the world

"All the ___ in the world"
"Every ____ ever created"
"World's largest ____ "
"Repository of all knowledge in ____"

There's something compelling about completeness, about the idea that you could gather ALL of something, anything, together into a single system or database or even, as in the ancient library of Alexandria, physical space. Perhaps it's because we want the satisfaction of being finished. Perhaps it's something primitive in our brain stems that has the evolutionary advantage of keeping us from declaring victory with a job half done. (Well, at least some of us.) To be sure, setting your goal to gather all of something means you don't have to make awkward choices about what to gather/keep and what to discard. The indiscriminate everything may be the easier target.

Worldcat has 229,322,364 bibliographic records.
OpenLibrary has over 20 million records and 1.7 million fulltext books.
LibraryThing has records for 6,102,788 unique works.
If you read one book a week for 60 years, you will have read 3,120 books. If you read one book a day for that same length of time, you will have read 21,360 (not counting leap years).
The trick, obviously, is to discover the set of books, articles, etc., that will enhance your brief time on this planet. To do this, we search in these large databases. By having such large databases to search we are increasing our odds of finding everything in the world about our topic. Of course, we probably do not want everything in the world about our topic, we want the right books (articles, etc.) for us.

There are some down sides to this everything approach, not surprisingly. The first is that any search in a large database retrieves an unwieldy, if not unusable, large set of stuff. For this reason, many user interfaces give us ways to reduce the set using additional searches, often in the form of facets. Yet even then one is likely to be overwhelmed.

Everything includes key works and the odd bits and pieces of dubious repute and utility. Retrieving everything places a great burden on the user to sort out the wheat from the chaff. This is especially difficult when you are investigating an area where you are not an expert. Ranking may highlight the most popular items but those may not be what you are seeking. In fact, they may be items that you have retrieved before, even multiple times, because every search begins with a tabula rasa.

Another down side is that although computers are more powerful than ever and storage space is inexpensive, these large databases tend to collapse under the demands of just a few complex queries. Because of this, what users can and cannot do is controlled by the user interface which serves to protect the system by steering users to safe functions. Users often can create their own lists, can add tags, can make changes to the underlying data, but they cannot reorder the retrieved set by an arbitrary data element, they can't compare their retrieved set against items they have already saved or seen previously, they can't run analyses like topic maps on their retrieved set to better understand what is there.

I conclude, therefore, that what would be useful would be to treat these large databases as warehouses or raw materials, and provide software that allow users to select from these to create a personal database. This personal database software would resemble, ta da!, Vannevar Bush's Memex, a combination database and information use system. I can see it having components that are analogous to some systems we already have:
The personal database would be able to interact with the world of raw material and with other databases. I can imagine functions like: "get me all of the books and articles from this item's bibliography." Or: "compare my library to The Definitive Bibliography of [some topic]." Or: "Check my library and tell me if there are new editions to any of my books." In other words, it's not enough to search and get; in fact, searching and getting should be the least of what we are able to do.

There are a whole lot of resource management functions that a student or researcher could find useful because within a selected set there is still much to discover. These smaller, personal databases should also be able to interact with each other, doing comparisons and cross-database queries. We should be able to make notes and create relationships and share them (a Memex feature). The personal database should be associated with person, not a particular library or institution, and must work across institutions and services. I can't imagine what it must be like today to graduate and to lose not only the privileged access that members of institutions enjoy but also the entire personal space that one has created while attached to that institution.

In short, it's not about the STUFF, it's about the services. It doesn't matter how much STUFF you have it's what people can DO with it. Verb, not noun. Quality not quantity.

Friday, May 13, 2011

Dystopias

In the 1990's I wrote often about information dystopias. In 1994 I said:

It's clear to me that the information highway isn't much about information. It's about trying to find a new basis for our economy. I'm pretty sure I'm not going to like the way information is treated in that economy. We know what kind of information sells, and what doesn't.

In 1995 I painted a surprisingly accurate picture of 2015 that included:
Big boys, like Disney and Time/Warner/Turner put out snippets of their films and have enticed viewers to upgrade their connection to digital movie quality. News programs have truly found their place on the Net, offering up-to-the second views of events happening all over the world, perfectly selected for your interests....Online shopping allows 3-D views of products and virtual walk-throughs of vacation paradises.

If there were a stock market for cynical investments, I'd be sitting pretty right now. But wait... there's more! Because there's always a future, and therefore more dystopia to predict.

My latest is concern is about searching and finding. And of course that means that I am concerned about Google, but this is in a new context. I have spent the last five years trying to convince libraries that we need to be of the web -- not only on the web but truly web resources. I strongly believe this is the only possible way to keep libraries relevant to new generations of information seekers. This has been interpreted by many as a digitization project that will result in getting the stuff of libraries (books mainly) onto the web, and getting the metadata about that stuff out of library catalogs and onto the web. Hathitrust, for example, is a massive undertaking that will store and preserve huge amounts of digitized books. The Digital Public Library of America (DPLA), just in its early planning stages today, wants to make all books available to everyone for "free."

All of these are highly commendable projects, but there is a reality that we don't seem to be have embraced, and that is that searching and finding are as important to the information seeking process as the actual underlying materials. As we can easily see with Google, the search engine is the gate-keeper to content. If content cannot be found then it does not exist. And determining what content will be accessed is real power in the information cloud. [cf. Siva Vaidhyanathan, Googlization of Everything.]

There is a danger that when this mass of library materials becomes of the web that we could entirely lose control of its discovery. But it isn't just a question of library materials, this is true for the entire linked data cloud: who will create the search engine that makes all of that data findable? With its purchase of freebase.com, it is clear that Google has at least an eye on LD space. And of course Google has the money, the servers, the technology to do this. We know, however, from our experience with the current Google search engine that the application of Google's values to search produces a particular result. We also know that Google's main business model is based on making a connection between searchers and advertisers. [cf. Ken Auletta, Googled] .

It's not enough for libraries to gather, store and preserve huge masses of information resources. We have to be actively engaged with users and potential users, and that engagement includes providing ways for them to find and to use the resources libraries have. We must provide the entry point that brings users to information materials without that access being mediated through a commercial revenue model. So for every HathiTrust or DPLA that focuses on the resources we need a related project -- equally well-funded -- that focuses on users and access. Not just creating a traditional library-type catalog but providing a whole host of services that will help uses find and explore the digital library. This interface needs to be part search engine, part individual work space, and part social networking. Users should be able to do their research, store their personal library (getting into Memex territory here), share their work with others, engage in conversations, and perhaps even manage complex research projects. It could be like a combination of Zotero, VIVO, Zoho, Yahoo pipes, Dabble, and MIT's OpenCourseWare.

Really, if we don't do this, the future of libraries and research will be decided by Google. There, I said it.

Tuesday, November 18, 2008

Google Giveth ... and Taketh Away

Some additions, amendments.

The agreement between Google and the AAP is of great significance for libraries. It is also very long, written in "legalese", and contains conclusions of a lengthy negotiation without revealing the nature of the discussion. Given that many lawyers were involved, we may never get the back story of this historic settlement, yet it has the potential to change the landscape on rights, digitization, and libraries.

I am basing much of my analysis on the summary of the agreement produced by ARL. This unfortunately means that some errors may be introduced between their summary and my interpretation. I have gone to the original document to check some particulars, such as definitions, but much of that document goes unread for now.

Key Points

(... or, a summary of the summary)

  • The agreement is primarily about books that are presumed to be in copyright but which are no longer in print. In-print books continue to be managed directly by the rights holders, who can make agreements with Google (or anyone else) for uses of those items.

  • The agreement has some odd limitations that baffle me: it only covers books published in the US that have been registered with the Copyright Office. It does not include any books published after January 5, 2009 .The settlement does cover non-US books (e.g. Berne countries); I'm still unclear on the statement about registration for US books, but it was cited in the ARL document.

  • The agreement trades off Google's liability with payment to rights holders. That is, as long as Google requires payment from users to displays and copies, and passes 2/3 of those monies to the rights holder, Google is exempt from copyright infringement claims by rights owners. So users of the digital files will pay to keep Google legal.

  • The agreement does not answer the all-important question of whether scanning for the purposes of searching is an allowed use under copyright law.

  • The agreement flaunts the concept of Fair Use by quantifying the amount of an in-copyright book that users can view for free ("20% of the text," "five adjacent pages," but not the final 5% of a fiction book, to keep the endings a surprise.) The ARL document has Google saying that it will not interfere with fair use. I can't find that statement in the actual settlement. These quantities are contractual, and I'm assuming that that technology will not allow users to exert fair use rights, only the contractual agreement.

  • Google will sell digital copies of in-copyright books to users, who will have perpetual access to the book online. Some printing will be allowed but all printed pages will have a watermark that identifies the user. (I'm calling this "ratwear," software that rats you out.) Users will be able to make notes on the book's pages, but they will only be able to share those notes with other purchasers of the book. (Thus buying a Google book is like joining a secret reading club.) The settle states that the watermark will identifier either the user, or other information "which could be used to identify the authorized user that printed the material or the access point from which the material was printed." Agreement, p. 47

Key Points Relating to Libraries

This is the hard part for me. Hard in that it really hurts.

  • After digitizing books held in libraries, Google will then turn around and become a library vendor, supplying those same books back to libraries under Google's control. Each public library in the US will get a single "terminal" provided (and presumably controlled) by Google that allows users to view (but not copy and paste from) books in the Google database. Some printing is allowed, but there will be a per-page fee charged.

  • Libraries and institutions can also subscribe to all or part of the database of out of print books. Access is not perpetual, but limited to the life of the subscription.

  • There is verbiage about how users in these institutions can share their "annotations." In other words, if you take notes on your own, obviously those are yours. But if you use the capabilities of the system to make your notes in the system, you cannot share your own notes freely.

Now for the Clincher


... this is the pact with the devil.

  • A library can partner with Google for digitization of its collection and get the same release from liability that Google has. The library can keep copies of these digitized books, however, it must follow security standards set by Google and the AAP and must submit its security plan for review and allow yearly auditing. (The security measures are formidable and quite possibly not affordable for all but the wealthiest institutions. There are huge penalties up to millions of dollars for not getting security right.)

  • Libraries that make this pact with the devil are thereby allowed to preserve the files, print replacement copies for deteriorating books, and provide access for people with disabilities. Note that all of these uses by libraries are already allowed by copyright law.

  • The libraries that make this pact with the devil cannot let their users read the digitized books. Well, they can let them read up to five (5!) pages in any digitized book. Presumably if the library wants to provide other uses it must subscribe to Google's service. Libraries are expressly forbidden from using their copies of the books for interlibrary loan, e-reserves, or in course management systems.

... and if you refuse to negotiate with the devil...

  • Current Google library partners who do not choose to become party to this must delete all copies of digitizations of in-copyright works made by the Google project in order to obtain a release from liability. If they choose not to delete the copies, they are on their own in terms of liability for the in-copyright books that Google did digitize (and Google knows exactly which books are involved.)

  • Even if the library was only allowing Google to digitize public domain works, those libraries must destroy all of their copies to get release from liability in case they mis-judged the copyright status of one of the those books.
In other words, this agreement is making the assumption that if anyone sues Google for copyright infringement, the library will be a party to that suit.

They say that "the devil is in the details." In this case that is not true: the devil is right up front, in the main message. That message is that Google has agreed with the publishers, and is selling out the libraries that is has been working with. The deal that Google and the libraries had was that in exchange for working with Google to digitize books in their collections, the libraries received a copy of the digital file. After that, it was up to the libraries to do the right thing based on their understanding of copyright law. Participating with Google has been an expensive proposition for the libraries in terms of their own staff time and in the development of digital storage facilities. Part of the appeal of working with Google was the assumption that partnering with the search giant gve the entire project clout and provided some protection for the libraries. With Google and the AAP now in cahoots, the libraries must join them or try to stand alone in an unclear legal situation; an unclear situation that Google invited the libraries into in the first place.

This is classic bait and switch. And it is bait and switch with powerful commercial interests against public institutions. There is no question about it...

THIS IS EVIL

Note: I've added more comment and info in the comments area as things pop up. So read on....

Monday, May 26, 2008

Amputation

OK, maybe I'm just in a particularly bad mood, but I guess I've just about had it with libraries shooting themselves in the foot. Then letting gangrene set in and going for amputation.

I'm talking about not sharing the data we have. We've all heard the complaints that the Web community is going forward "reinventing the wheel" in terms of bibliographic data. But then when anyone shows an interest in our bibliographic data, we withhold like the anal retentives we are. (Wow, I must be in an extraordinarily bad mood!)

We've just learned that OCLC is sharing its data with Google, and that libraries can download records for Google books from OCLC -- although there was no mention of the "usual fee," which I'm sure applies. I happen to know that Google receives a full bibliographic record with each book that it digitizes. I have no idea what they do with that data because it doesn't appear on the screen in Google Book Search. What is significant is that the quality data that has been created by libraries is still essentially invisible to most users of the Internet. Do you still wonder why we are overlooked by most of the information seeking population? How can they possibly know what we've done to organize the bibliographic world if we won't let them see?

So the latest thing that got my goat came across on the NGC4LIB list. First, it turns out that University of Michigan has made available a file of record representing the books of their digitized by Google that are in the public domain via OAI-PMH. (The OAI Identify command.) This is a good thing, obviously. But in this post, we learn that the records have been "truncated" to meet the requirements of OCLC for record sharing. So I decided to see what "truncated" means. Here are some examples. The bolded text is what you get when you retrieve the records from Michigan. The text in italics is what shows in the Michigan catalog for the same item.

1.
LDR nam 22003251i 4500
005 19901105000000.0
006 m d
007 cr bn ---auaua
008 901105s1980 dcu b f00010 eng d
020 |b pbk.
035 |a (OCoLC)ocm06624048
040 |a GPO |c GPO |d m.c. |d m/c |d EYM
074 |a 968-H-1
0860 |a J 26.2:C 73/10
1001 |a Villano, Clair E.
24510 |a Complaint and referral handling / |c by Clair E. Villano, Metropolitan Denver District Attorneys' Office of Consumer Fraud and Economic Crime.
260 |a [Washington] : |b Dept. of Justice, Law Enforcement Assistance Administration, |c 1980.
300 |a v, 25 p. ; |c 28 cm.
440 0 |a Operational guide to white-collar crime enforcement
500 |a At head of title: The National Center on White-Collar Crime.
500 |a Project supported by Grant No. 77-TA-99-0008, awarded to the Battelle Memorial Institute Law and Justice Study Center.
500 |a May 1980.
504 |a Includes bibliographical references.
533 |a Electronic text and image data |b Ann Arbor, Mich. : |c University of Michigan Library |d 2007 |e Includes both image files and keyword searchable text. |f [Michigan Digitization Project]
538 |a Mode of access: Internet.
650 0 |a Complaints (Criminal procedure) |z United States.
650 0 |a White collar crimes |z United States.
7101 |a United States. |b Law Enforcement Assistance Administration.
7102 |a National Center on White-Collar Crime (U.S.)
7102 |a Metropolitan Denver District Attorneys' Office of Consumer Fraud and Economic Crime.
7102 |a Battelle Law and Justice Study Center.
856 4 |uhttp://hdl.handle.net/2027/mdp.39015034803505
|wmdp.39015034803505 |xeContent

2.
LDR nam 2200301 a 4500
005 19901105000000.0
006 m d
007 cr bn ---auaua
008 901105s1980 dcu bc f00010 eng c
010 |a 81600912 //r90
035 |a (OCoLC)ocm06999418
040 |a DGPO/DLC |c DLC |d GPO |d EYM
043 |a n-us---
05000 |a Z6616.B3215 |b B37 |a E746.B37
074 |a 383-B
08200 |a 016.355/0092/4 |2 19
0860 |a D 214.13:B 26/859-930
1001 |a Bartlett, Merrill L.
24510 |a George Barnett, 1859-1930 : |b register of his personal papers / |c compiled by Merrill L. Bartlett.
260 |a Washington, D.C. : |b History and Museums Division, Headquarters, U.S. Marine Corps, |c 1980.
300 |a vii, 18 p. ; |c 27 cm.
504 |a Bibliography: p. 17-18.
533 |a Electronic text and image data |b Ann Arbor, Mich. : |c University of Michigan Library |d 2007 |e Includes both image files and keyword searchable text. |f [Michigan Digitization Project]
538 |a Mode of access: Internet.
60010 |a Barnett, George, |d 1850-1930 |x Manuscripts |x Catalogs.
61010 |a United States. |b Marine Corps. |x History |x Sources |x Manuscripts |x Catalogs.
61010 |a United States. |b Marine Corps. |b History and Museums Division |x Catalogs.
650 0 |a Manuscripts, American |z Washington (D.C.) |x Catalogs.
7101 |a United States. |b Marine Corps. |b History and Museums Division.
856 4\|uhttp://hdl.handle.net/2027/mdp.39015035037590
|wmdp.39015035037590|xeContent


There are obviously many other examples, but the trend is clear (well, partly so). The lack of a 245 $c (statement of responsibility) and all of the 7XX's (added entries) means that the record is incomplete in terms of authorship. You won't be able to see or search on anything but the main author. I'm baffled by the removal of the place of publication, since it's not used for retrieval (the coded place is in the 008 field). Ditto the 300 $c with the size in centimeters. The subject headings have been rendered entirely useless. As we know, the 6XX $a is not the top of some logical hierarchy, but is idiosyncratically the first term based on some rather complex rules. So in the first record we lose "United States" because it is the second term, but in the second record we get only "United States" and lose all references to "Marine Corps." which is the actual topic of the item.

I note also that the output records do not have the required 001/003/005 fields. (And not having required fields is a problem in itself.) These are needed to identify the source of the record (003) and provide a unique identity for the record in that source system (001 + 003). The 005 would give the date of the most recent update before the record was exported. The combination of these three would allow a receiving system to accept periodic updates to these records. I suspect (and this is just me) that the existence of the 001 would also allow one to retrieve the un-truncated record from Michigan's catalog using something like Z39.50. I admit, however, that the lack of these fields could be a simple error.

I'm not going to say that these records are completely useless, and I am talking to the Open Library folks about adding them to their database. But I do consider the maiming of these records to be an embarrassment, a kind of self-mutilation done by a profession with amazing low self-esteem. Please, folks, let's find our power and stand up for what we know is right!

Sunday, February 03, 2008

The ILS minus the catalog

The greatest amount of action happening today regarding library user services is the separation of the user interface from the integrated library system (ILS). This seems odd, perhaps, since only two decades ago the integration of all of the functions of library systems was seen as a real step forward. Until then, one system had handled acquisitions, another circulation, and another cataloging. Many functions, such as serials check-in and bindery management, were not managed through automation. This situation had a number of problems: different data about the same book were stored in multiple databases or in card files, leading to inconsistencies throughout the system; the data had to be keyed or copied multiple times; system-wide updates were nearly impossible. The "integration" of the integrated library system was the creation of a single database for bibliographic and management data, where all of the information about an item would be stored once and only once. This also was the first time that the full bibliographic record was linked to the library management functions. Independent systems like acquisitions and circulation systems primarily used brief records only. At best, these brief records contained an identifier (such as the item's barcode or call number) that could connect it to other records in other systems. Sometimes even that wasn't possible.

In theory, this database integration is a dandy way to organize your data and the activities that use your data. In reality, the user interface suffered in this design. Not that anyone purposely shorted the user interface, but in a world of scarcity, there are things that just have to get done; and then there are other things. In the have to category, libraries have to make purchases, manage accounts, receive and check-in serials, perform interlibrary loans, and check items out to borrowers. These are clear, quantitative, auditable library functions, and ones that library administrators focus on. These are the functions that can have dollar amounts attached to them in terms of staff time. These functions are the inside view of the library, the library being a library.

User success, on the other hand, is qualitative, hard to define, and does not have a direct effect on the library's bottom line. We count the countables, like numbers of bibliographic records, items circulated, and online database accesses, but there appears to be no penalty for a lousy user interface and no premium for the creation of a good one. If at any point there is a conflict between quantitative library management and qualitative user service, my gut feeling is that the latter loses out.

Users have everything to gain from the separation of the user interface from the library management system. Libraries, however, are in a bit of a bind. The new "user interface on top of the ILS" adds features for users but it doesn't result in any less work in the ILS. Libraries are still hanging all of their management functions off of full bibliographic records in the catalog (which the users no longer see). Librarians still see the data creation functions in the early management steps of acquisitions and receipt to be a direct line to the standard bibliographic record that in the end will appear on the users' screens. They are still storing the full bibliographic records in a local database, although these records are siphoned off nightly to the "real" user interface.

Much of the objection to using more EDI (electronic data interchange) functions with our vendors is that their data doesn't conform to library cataloging. Yet our library management systems are getting further from the user interface. We may need to rethink the library management workflow as well as the basis of our cataloging activity. What could we achieve if we move cataloging and catalogs out of our individual library databases to the network level? Could this provide the basis for increased sharing of the cataloging effort? Are there other efficiencies that could be gained in the "back room" functions of purchasing and managing the library inventory?

Monday, March 19, 2007

Unintended consequences

I was working at the University of California in the early 1980's when the UC union catalog, MELVYL, was developed. Shortly after MELVYL became available as one of the early online public access catalogs, we obtained a copy of NLM's Medline database, consisting of articles and books in a wide variety of fields related to medicine. This was the first article database that we were making available. At that time, the only people who had access to Medline were medical researchers in the four or five medical centers at the university, and they accessed it via the Dialog search service. Dialog charged by the minute and was relatively pricey. Few members of the university community had access to Dialog's databases in any subject area because of the cost.

I don't know how many minutes or hours of searching were done monthly on Medline before we added the database to the university's library system, but within a few years the number of searches on Medline were rivaling the number of searches in the entire union catalog of the 9 university campuses. It was heavily used even on those campuses that had no medical school. Had everyone developed a sudden interest in the bio-medical sciences? Perhaps, but not to that degree. I think that we had created a monster of "availability." As the only freely available online database of articles, Medline became the one everyone searched.

(If anyone is looking for a master's thesis, try looking at the citations in dissertations granted at the University of California for the period 1985-1995, and compare it to the previous decade. Count how many of the cited articles come from journals that are indexed in NLM's database. I have a feeling that it will be possible to see the "Medline-ization" of the knowledge produced by that generation of scholars, from architecture to zoology. )

When we make materials available, or when we make them more available than they have been in the past, we aren't just providing more service -- we are actually changing what knowledge will be created. We tend to pay attention to what we are making available, but not to think about how differing availability affects the work of our users. We are very aware that many of us are searching online and not looking beyond the first two screens, which produces an idiosyncratic view of the information universe. But we don't see when libraries create a similar situation by making certain materials more available than others (for example scanning all of their out-of-copyright works and making them available freely as digital texts, while the in-copyright books remain on the shelves of the library).

There's a discussion going on at the NGC4LIB list about the meaning of "collection." What is a library collection today? Is it just what the library owns? Is it what the library owns and also what it licenses? Does it include some carefully selected Internet resources? Some have offered that the collection is whatever the library users can access through the library's interface. I am beginning to think that access is a tricky concept and it is inevitably tied up with the realities of a library collection. Users will view the library's collection through the principle of least effort. In the user view, ease of access trumps all -- it trumps quality, it trumps collection, it trumps organization. So we can't just look at what we have -- we have to look at what the user will perceive as what we have, and that perception will necessarily be tempered by effort and attention. To our users, what the library has will be what is easiest to locate and fastest to arrive.

In other words, our collection is not a quantity of materials. The collection is a set of services built around a widely divergent set of resources. To the user, the services are the library, especially because any one user will see only a tiny portion of what the library has to offer. The actual collection -- those thousands or millions of library-owned items -- is not what the user sees. The user sees the first two screens of any search.

Hopefully, they are not in main-entry alphabetical order.

Tuesday, February 27, 2007

Ebooks in XML - The IDPF/OEB standards

I received a notice today about a conference being held by O'Reilly on digital publishing. The conference has some tutorial sessions on using XML to create digital books, but I fear that these will not include the work being done to create an XML standard for ebooks. Once again proof that the East coast (where "traditional" publishing takes place) and the West coast (where technology happens) are very far apart.

The International Digital Publishing Forum (IDPF - once known as the Open eBook Forum) has recently announce a beta version of its e-book coding standard. I've been watching, and sometimes participating in, this group for a while, and I really think they deserve our attention and support.

To begin with, the IDPF publication structure standard (still termed OEBPS - Open Ebook Publication Structure) is designed to be used by publishers in the preparation of files that will be sent to the technology companies that transform the raw files into actual ebooks. As you know, there are dozens of e-book formats (PDF, Microsoft Reader, Mobipocket, Palm reader... etc.). The publishers need to create a single file that can be transformed into all of those formats, and the OEBPS standard is designed to meet that need. It is also designed to be an ebook format in its own right, and the upcoming Adobe ebook reader, "Adobe Digital Editions," based on Adobe's flash technology, will be able to display books in the OEBPS format.

The standard will seem overly simple to many people. It is that way on purpose. The original OEBPS standard used HTML, based on the assumption that even the publishers, who are notoriously lacking in technology chops, would have someone on board who knows HTML. The second version of the standard, the one out for comment, uses XHTML and CSS. I think this is brilliant. It means that 1) anyone can create a book and 2) anyone can display it, even in a simple browser. The KISS principle is essential for industry acceptance of the standard.

Another key thing to mention is that the OEBPS has been greatly influenced by members of the accessibility community who participate in the IDPF. The Digital Talking Book standard, which was first developed by the DAISY consortium and is now NISO standard Z39.86, uses an earlier version of the OEBPS as its book structure. This is the format that allows synchronization between a text and a reading of the text by a human reader, making it ideal for sighted and non-sighted readers alike (read it in bed, then continue listening in the car).

There is a DTD for the publication structure, although I am currently unable to get it to validate and behave. I have a question out to the authors of the DTD and will post here when I get an answer. Meanwhile, you can comment out the offending entity definition and play with the DTD.

A companion to the ebook standard is the Open Packaging Format (OPF). The easiest way to understand this is to take a look at it. Download Thoughts.epub.
Now open it in Winzip -- yes, it is a simple zipped file. In it you can find the raw xhtml of the publication; an OPF file that is the manifest for the package, and contains Dublin Core metadata for the item; a file that contains the mimetype; any images or other files that are required by the document; and an XML document that defines the overall container. Note that this is a very simple publication. The examples in the documentation show how you would create a document with multiple chapters, cover art, and illustrations. It also covers the areas of encryption and keys, for files that will be transmitted in protected formats. There is a nifty tutorial that steps you through the creation of an OCF file using Winzip.

If you have comments, suggestions, questions, or whatever, go to the discussion area of the IDPF web site and say your piece. And let me know if you have any thoughts on these standards, especially as to how they might be applicable to digital libraries.

Sunday, December 17, 2006

Digitization and the Catalog

I have just posted the preprint of my current column for the Journal of Academic Librarianship, titled "Mass Digitization of Books." It takes about 4-6 months for the columns to be published, and as I read over this one I can see that things have already changed. For example, when I wrote the column, Google was not yet allowing the download of its public domain books.

However, I should have included one more very important issue in the article, but it hadn't occurred to me at the time: the effect of this mass digitization on our catalogs. The cataloging rules require that the digital copy be represented in the catalog with its own record. This means that a library that undergoes a mass digitization project on its book collection faces doubling the number of book records in its catalog. Leaving aside the issues of user display for now, and assuming that the creation of the records requires very little human intervention, we can probably still calculate a significant cost in storage space (albeit cheap these days), the size of backups, the time to load and index all of those records, and a general overhead in the underlying database.

This brings up the issue of creating catalog entries that represent "multiple versions," that is, having a single record that contains the information for all of the different formats in which the book is available -- regular print, e-book version, digitized copy, large print. There are good arguments both for and against, and it's a complex discussion, but I'll just say that I am convinced that we could structure our catalog records in a way that would make this work.

Wednesday, November 01, 2006

Cat-a-log(gue)

I was taken to task in the FRBR Blog for saying that FRBR is not about catalogs.
I don’t understand the statement that FRBR isn’t about catalogues. It’s Functional Requirements for Bibliographic Records, and when bibliographic records are shown to users, that’s called a catalogue.
I realized from this comment that I am using the wrong term when I say catalog. A catalog is defined as a list of items, coming from word roots that mean "list, to count up." Examples are a library catalog, a list of items offered for sale, or the catalog of a museum exhibit that lists the items in the exhibit. In the library, the catalog is an inventory of items owned or held by the library. The library catalog is an inventory of items held in the library, and it was once equal to what users could access in the library. From the early or mid-1800's, the physical library catalog also served the purpose of the user interface to the library's holdings. (Prior to that time, catalogs were mainly used by librarians, and not members of the library's public.)

The as-yet-unnamed thing that I wish to define (and which I mistakenly called a catalog) serves these functions (please add on what I have forgotten):
  1. A list of items owned by the library. This list is at a macro level (e.g. serial titles but not the articles in the serial). Often that level is determined by the purchase unit, since this list interacts with the library's acquisitions function.
  2. Serial issues received. This is usually found in a separate module called a serials check-in system (which replaced the old Kardex)
  3. Licenced resources. These may be listed in the catalog, but they may either/also be found in a database used by an OpenURL resolver or in an ERM system (which is not accessible to users). In some cases, these are listed on a web site managed by the library.
  4. Journal article indexes. These used to be hard-copy reference books. They are now often electronic databases. User interface to these varies.
  5. Items available via ILL. This could be a union catalog of libraries in a borrowing unit. It also is a function that interacts with OCLC's ILL system. This latter usually isn't visible in the user view of the library.
  6. Links and connections from information systems not hosted by the library, such as the ability to link from an article in a licensed database to the full text of the article from another source; or a link from an Internet search engine to library-managed resources through a browser plug-in or a web service.
  7. Location and circulation status information, plus the ability for users to place holds on items or to request delivery of items.
  8. Interaction with institutional services such as courseware.
  9. One or more user interfaces. Many of these services above will have a separate interface just for that service, but there are also meta-interfaces that will combine services.
This is a first stab. I'll post this on the futurelib wiki where it can be easily modified.

Thursday, September 21, 2006

Description and Access?

The revision of the library cataloging rules that is underway is being called "Resource Description and Access" or RDA. Although it is undoubtedly an unpopular view point, I would like to suggest that description and access are two very different functions and that they should not be covered by a single set of rules, nor should they necessarily be performed by a single metadata record.

The pairing of description and access is functionality based on card catalog technology. A main purpose of the 19th and 20th century card catalog was access. Indeed, great discussion took place in the late 19th century about the provision of a public access point for library users: access through authors, titles, and subjects. The descriptive element, the main body of the card, was essentially a bibliographic surrogate helping users make their decision on whether to go over to the shelf to look for the book. Before easy reproduction of cards, that is, before LoC began selling card sets early in the 20th century, access cards did not carry the full description of the book. Instead, all catalog cards except the main entry card had a brief entry that would allow the user to find the main entry card which had the full description.

The combination of description and access is a habit that has carried over from the card catalog and has left a legacy that to many of us is so natural that we have trouble seeing it for what it is. For example, it is because of this combination that we create artificial "headings" that cause us to display author names in the famed "last name first" order. The heading is designed for access in a system where the means of finding items is through a linear alphabetical order, which even in library systems is no longer the predominant finding method. These headings set library systems apart from popular information systems such as Amazon, Barnes & Noble, Google Books. As a matter of fact, you can find examples of library catalogs that attempt a popular display by displaying the title and statement of responsibility as the main display, hiding the now odd-looking headings. What these headings say to anyone who is tech-savvy is that libraries are hindered by an obsolete technology. Libraries still create these contorted headings when markup of data can make display and ordering of data flexible and user friendly.

Not only does our use of arcane headings set libraries apart from more popular information resources, our concepts of "description" and "access" are not serving our users. The description provided by libraries might serve to identify the work bibliographically (something that matters to libraries for collection development purposes but is not of great interest to library users), however it doesn't describe the work to users in a way that can help them make a selection. We need at least reviews, thumbnails of images, sample chapters, and even local commentary ("Required reading for Professor Smith's class in European History"). And as for access, we know that the library-assigned subject headings are woefully inadequate discovery tools.

RDA claims that its purpose in the area of description "should enable the user to: a) identify the resource described...." Yet today we are in dire need of machine-to-machine identification, which RDA does not address. Increasingly our catalogs are interacting with other sources of discovery, such as web sites, search engines, and courseware. "Identification" that must be interpreted by a human being is going to be less and less useful as we go forward into an increasingly digital and networked information environment.

We are also greatly in need of an ability to share our data with systems that are not based on library cataloging. Each rule that varies from what would be common practice moves libraries further from the information world that our users occupy in their daily life. It is somewhat ironic that many pages of rules instruct catalogers on the choice of the "title proper," which is then marred by the addition of the statement of responsibility, a bit of library arcana that no one else considers to be part of the title of the work. And who else would create a title heading "I [heart symbol] New York"?

All this to say that the next generation library catalog cannot succeed if it is to be based on a set of rules that still carry artifacts from the days of the physical card catalog. It's time to get over the concepts of description and access that were developed in the 19th century. Let's move on, for goodness sake.

Tuesday, August 29, 2006

The dotted line

The University of California has released its agreement with Google ("uc:" in quotes below). As a public institution, all such contracts must be made publicly available on request. We similarly have access to the University of Michigan agreement ("um:" in quotes below), which gives us the ability to do some comparison.

As is always the case, the language of the contract does not entirely reveal the intentions of the Parties. It is instead a strange almost-Shakespearean courtship where neither Party wishes to say what they really want, and everyone pretends like it's all so wonderfully fine, while at the same time each player is hoping to pull the wool over the eyes of the other. So my interpretation here may reveal more about my own assumptions than any truth about the contracts.

[Note: I apologize for any typos, but I had to transcribe much of this from the PDF files, which did not allow for text copy. If you see errors, let me know and I'll fix them.]

Quality Control

The Michican Contract gives the library the quality control review function, and allows the library to actual hold up the digitizing process if its quality requirements are not met:
um: 2.4 Digitizing the Selected Content. Google will be responsible for Digitizing the Selected Content. Subject to handling constraints or procedures specified in the Project Plan, Google shall at its sole discretion determine how best to Digitize the Selected Content, so long as the resulting digital files meet the benchmarking guidelines agreed to by Google and U of M, and the U of M Digital Copy can be provided to U of M in a format agreed to by Google an U of M. U of M will engage in ongoing review (thorough sampling) of the resulting digital files, and shall inform Google of files that do not meet benchmarking guidelines or do not comply with the agreed-upon format, U of M may stop new work until this failure can be rectified.
Perhaps Google has learned its lesson about trying to meet the standards of libraries, because UC's contract is notably silent on the QC topic. The paragraph that begins ...
uc: 2.4 Digitizing the Selected Content. Subject to handling constraints or procedures specified in the Project Plan, Google shall in its sole discretion determine how best to Digitize the Selected Content.
... then goes on to talk about Google's responsibility in taking care of UC's books, and its promise to replace any that get damaged. Nothing more about the digital files that result.

[Note: Peter Brantley points out section 4.7.1 which contains a reference to image standards and lets UC QA up to 250 books a month to "assess quality." However, there is no stated recourse so UC and Google are relying on each other's good intentions here.]
note 4.7.1, which refers to image standards in line with
those established as a community of library partners.

What the Libraries Get

In this case, UC seems to have learned from the past experience of others, and negotiated for more from Google:
um: 2.5 ... the U of M Digital Copy will consist of a set of image and OCR files and associated information indicating at a minimum (1) bibliographic information consisting of the title and author of each Digitized work, (2) which image files correspond to that Digitized work, and (3) the logical order of these image files.
uc: 4.7 University Digital Copy. Unless otherwise agreed by the Parties in writing, the "University Digital Copy" means the digital copy of the selected content that is Digitized by Google consisting of (a) a set of image and OCR files, (b) associated meta-information about the files including bibliographic information consisting of title and author of each Digitized work and technical information consisting of the date of scanning the work, information about which image files correspond to what digitized work, and information pertaining to the logical order of image files that make up a Digitized work, (c) a list of works that are supplied for Digitization but not actually Digitized, and (d) the image coordinates for each Digitized Work ("Image Coordinates"); provided that Image Coordinates will only be provided (i) so long as University complies with the volume commitments set forth in Section 2.2 and (ii) pursuant to the restrictions on University's use and distribution of such Image Coordinates set for the Section 4.10.
The "Image Coordinates" are what make it possible to locate a word in a page image, for example for the purpose of highlighting the query word on the screen. Michigan didn't get these coordinates with its files, and possibly the other four original Google library partners did not, either. We'll look at Section 4.10 and its limitations in a moment, but the "volume commitments" in 2.2 say that:
uc: University will use reasonable efforts to provide or provide Google with access to no less than three thousand (3,000) books (or such amount that is mutually agreed to by the Parties) of Selected Content per day to Digitize commencing on the sixty-first (61st) day after the Effective Date...
So it sounds like the University really, really, really wanted the coordinates and Google really, really, really wanted to make sure that the University would not drag its heels in terms of providing the books to Google. So these two inherently unrelated desires became bargaining chips.

Using the Files

In both contracts it is stated that Google owns the Digital Copy, and makes it clear that neither Google nor the library are claiming any ownership of the underlying texts that have been digitized. This seems to be in keeping with US copyright law, although there is the inherent difficulty that occurs when you display a digital version of a public domain resource on a screen. At that point, any controls desired by the owner of the digital file are hard to enforce. In the Michigan contract, Google basically states that the university will not allow wholesale downloading of the files, and will attempt to prevent any downloading for commercial purposes (as if they could tell that from a download action):
um: 4.4.1 Use of U of M Digital Copy on U of M Website. U of M shall have the right to use the U of M Digital Copy, in whole or in part at U of M's sole discretion, as part of services offered on U of M's website. U of M shall implement technological measures (e.g., through use of the robots.txt protocol) to restrict automated access to any portion of the U of M Digital Copy or the portions of the U of M website on which any portion of the U of M Digital Copy is available. U of M shall also make reasonable efforts (including but not limited to restrictions placed in Terms of Use for the U of M website) to prevent third parties from (a) downloading or otherwise obtaining any portion of the U of M Digital Copy for commercial purposes, (b) redistributing any portions of the U of M Digital Copy, or (c) automated and systematic downloading from its website image files from the U of M Digital Copy. U of M shall restrict access to the U of M Digital Copy to those persons having a need to access such materials and shall also cooperate in good faith with Google to mutually develop methods and systems for ensuring that the substantial portions of the U of M Digital Copy are not downloaded from the services offered on U of M's website or otherwise disseminated to the public at large.

There are a number of interesting items in the above paragraph. First, that a robots.txt file is considered a "technological measure." In fact, it is at best a gentleman's agreement; there is nothing that forces you to abide by the instructions in the robots.txt file so the less gentlemanly are not stopped from accessing the items that it calls "disallowed." It also is theoretically only a message center for web crawlers, and not a way to limit access to sections of ones web site. More robust technology must be employed for that. The next is that access is to be restricted to "those persons having a need to access such materials" which is about the vaguest access condition that I can imagine. How could any of us show that we have such a need in relation to an information resource? Well, if nothing else, that language doesn't appear in the UC contract, so maybe they've re-thought that particular requirement.

Next, the contract allows Michigan to use the digital files to provide services to its consortial partners, but basically leaves it up to Michigan to get the proper agreements from those libraries:
um: 4.4.2 Use of U of M Digital Copy in Cooperative Web Services. Subject to the restrictions set forth in this section, U of M shall have the right to use the U of M Digital Copy, in whole or in part at U of M's sole discretion, as part of services offered in cooperation with partner research libraries such as the institutions in the Digital Library Federation. Before making any such distribution, U of M shall enter into a written agreement with the partner research library and shall provide a copy of such agreement to Google, which agreement shall: (a) contain limitations on the partner research library's use of the materials that correspond to and are at least as restrictive as the limitations placed on U of M's use of the U of M Digital Copy in section 4.4.1; and (b) shall expressly name Google as a third party beneficiary of that agreement, including the ability for Google to enforce the restrictions against the partner research library.
Another possible learning experience, or perhaps just a result of the particular negotiations between UC and Google, but the UC/Google contract is much more specific about uses of the files and the agreements that are required for UC to exchange part of all of the files with other parties. It does limit access to the digital files to UC Library patrons, which means that these will probably be treated similarly to licensed resources in the Library, which require a user ID login for access. Although the UC contract also contains the "robots.txt" language, it also contains some stronger wording about creating technological protection measures for the files:
uc: 4.9 Use of University Digital Copy. University shall have the right to use the University Digital Copy, in whole or in part at University's sole discretion, subject to copyright law, as part of services offered to the University Library Patrons. University may not charge, receive payment or other consideration for the use of the University Digital Copy except that University may charge of use of any services supplemental to the original work that the University supplies that add value to the University Digital Copy (for example, University may charge University Library Patrons for access to annotations to works from professors and scholars but the original work will always be accessible without a fee), and to recover copying costs actually incurred. University agrees that to the extent it makes any portion of the University Digital Copy publicly available, that it will identify the works, in a statement on a web page or other access point to be mutually agreed to by the Parties, as "Digitized by Google" or in a substantially similar manner. University shall implement technological measures (e.g., through use of the robots.txt protocol) to restrict automated access to any portion of the University Digital Copy, or the portions of the University website on which any portion of the University Digital Copy is available. University shall also prevent third parties from (a) downloading or otherwise obtaining any portion of the University Digital Copy for commercial purposes, (b) redistributing any portions of the University Digital Copy, or (c) automated and systematic downloading from its website image files from the University Digital Copy. University shall develop methods and systems for ensuring that substantial portions of the University Digital Copy are not downloaded from the services offered on University's website or otherwise disseminated to the public at large. University shall also implement security and handling procedures for the University Digital Copy which procedures shall be mutually agreed by the Parties. Except as expressly allowed herein, University will not share, provide, license, or sell the University Digital Copy to any third party.

The image coordinates, which UC seems to have "won" as a special deal, cannot be shared at all:
4.10: (a) University shall not share, provide, license, distribute or sell the Image Coordinates to any entity in any manner. University may use the Image Coordinates only as part of the University Digital Copy for the services provided to University Library Patrons set forth in Section 4.9 above.
What is particularly odd below in 4.10 (b) is that Google states that UC can distribute no more than 10% of the digital copy (which is, by definition, owned by Google), but it can distribute 100% of the digital copies of public domain works. I can imagine that UC insisted on this, but it seems to contradict the distinction that Google is making between the rights in the digital files created by Google and the rights in the underlying works.
(b) Subject to the restrictions contained herein, University shall have the right to distribute (1) no more than ten percent (10%) of the University Digital Copy (but not any portion of the Image Coordinates) to (i) other libraries and (ii) educational institutions, in each case for non-commercial research, scholarly or academic purposes and (2) all or any portion of public domain works contained in the University Digital Copy (but not any portion of the Image Coordinates) to research libraries for research, scholarly and academic purposes by those libraries and the faculty, students, scholars and staff authorized by said libraries to access their commercially licensed electronic information products. Any recipient of the University Digital Copy under this Section 4.10 is referred to herein as a "Recipient Institution." Prior to any distribution by University to a Recipient Institution, Google and the Recipient Institution must have entered into a written agreement on terms acceptable to Google governing the use of the University Digital Copy and that, among other things, provide an indemnity to Google. In addition, any distribution by University to a Recipient Institution is subject to a written agreement that (A) prohibits that Recipient Institution from redistributing without first obtaining the prior written consent of Google, (B) makes Google an express third party beneficiary of such agreement, (C) provides an indemnity to Google from the Recipient Institution for the Recipient Institutions's use of the Selected Content, (D) contains limitations at least as restrictive as the restrictions on University set forth in Section 4.9, (E) contains limitations on the use of the University Digital Copy consistent with copyright law and the limitations set for in clauses (1) and (2) above, and (E) requires each Recipient Institution, to the extent it makes any portion of the University Digital Copy publicly available, to identify the works, in a statement on the applicable web page or other access point, as "digitized by Google" or in a substantially similar manner.
Here I notice especially "(E) contains limitations on the use of the University Digital Copy consistent with copyright law" and I'm wondering to what this refers. It seems it either means that Google is asserting some intellectual property rights in the digital copies, or that they are reminding the University that it cannot re-distribute the digital copies beyond that allowed by fair use. Since the latter is a given, and not a matter of contract, it would appear that the first interpretation is correct. Yet I don't see a clear statement of Google's IP rights in the contract.

My final comment has to do with the fact that the licenses are for limited times. Michigan's extends until 2009, and UC's is stated as being for six years from the signing. Someone with more expertise in contract law will need to help me understand what this means for the restrictions given above. This may be clearer through a reading of the contracts, and I encourage anyone with the necessary skills to read them and let the rest of us know what some of this language means. Naturally, our concerns are about ownership and use, and getting a fair shake for library users.