Wednesday, April 25, 2012

Digital Urtext

As we reach a point where many of the classic books of literature and science published before the magical date of 1923 have been digitized, it is time to consider the quality of those copies and the issue of redundancy.

A serious concern in the times before printing was that copying -- and it was hand-copying in those times -- introduced errors into the text. When you received a copy of a Greek or Latin work you might be reading a text with key words missing or mis-represented. In our digitizing efforts we have reproduced this problem, and are in a similar situation as that of the Aldine Press when it set out to reproduce the classics for the first time in printed form: we need to carry the older texts into the new technology as accurately as possible.

While the digitized images of pages may be relatively accurate, the underlying (and uncorrected, for the most part) OCR introduces errors into the text. The amount of error is often determined by quality of the original or the vagaries of older fonts.If your OCR is 99.9% accurate, you still have one error for every 1,000 characters. A modern book has about 1500 characters on a page, so that means one error for every page. Also, there are particular problems in book scanning, especially where text doesn't flow easily on the page. Tables of contents seem to be full of errors:
IX. Tragedy in the Gra\'eyard 80 

X. Dire Prophecy of the Howling Dog .... 89 
XL Conscience Racks Tom 98 
 
In addition, older books have a tendency to use hyphenated line breaks a great deal:

and declined. At last the enemy's mother ap-
peared, and called Tom a bad, vicious, vulgar child, 

These remain on separate lines in the OCR'd text, which is accurate to the original but which causes problems for searching and any word analysis.

The other issue is that for many classic works we have multiple digital copies. Some of these are different editions, some are digitizations (and OCR-ing) of the same edition. Each has different errors.

For the purposes of study, and for the use of these texts for study, it would be useful to have a certified "Urtext" version, a quality digitization with corrected OCR that scholars agree represents the text as closely and accurately as possible. This might be a digital copy of the first edition, or it might be a digital copy of an agreed "definitive" edition.

We have a notion of "best edition" (or "editions") for many ancient texts. Determining one or a small number of best editions for modern texts should not be nearly as difficult. Having a certified version of such texts must be superior to having students and scholars reading from and studying a wide variety of flawed versions. Professors could assign the Urtext version to their classes, knowing that every one of the students was encountering the same high quality text.




(I realize that Project Gutenberg may be an example for a quality control effort -- unfortunately those texts are not coordinated with the digital images, and often do not have page numbers or information about the edition represented. But they are to be praised for thinking about quality.)

Thursday, April 19, 2012

Clarification from Sweden on OCLC negotiations

The National Library of Sweden has issued a short blog post clarifying their objections to the WorldCat Rights and Responsibilities (WCRR) policy. The inability of the two parties to reconcile these issues led the library to break off contract negotiations with OCLC. I find the Library's objections to be logical and undeniable:

1. The relationship with OCLC around record use is asymmetrical, with OCLC having the right to do whatever it wishes with the records while library use is restricted by the policy.

2. The policy actually requires libraries to favor WorldCat over other services, and thus hinder competition, which is not appropriate for a national library. [kc: This may even be illegal for publicly funded libraries in the US.]

3.  Open data is of strategic importance for libraries.

They conclude with:

To this end we urge OCLC to allow members to treat downloaded records as their own, including releasing them under any open license such as CC0. We feel that this would strengthen rather than diminish OCLCs strong status as a service provider to the library community.

Sunday, April 08, 2012

Content and carrier

In the midst of a discussion regarding the description of extents in RDA, I came to a realization that I might have noticed sooner if I did cataloging. As it is, I am probably coming to this a bit late.

RDA chapter three describes carriers. This is where you find all of the terms of measurement that appear in library data, things like:

12 slides
1 audiocassette
1 map
box 16 × 30 × 20 cm

There is a controlled vocabulary in RDA for carriers. It has 54 entries that are in 8 categories:
audio carriers
computer carriers
microform carriers
microscopic carriers
projected image carriers
stereographic carriers
unmediated carriers
video carriers

Note that one of the examples above, "map," is not included in the list of carriers. Nor is the most common extent used, "pages."* These are described in their own lists, "Extent of cartographic resource" and "Extent of text."** Why are these separate from other carriers? The answer is: Because they are not carriers, they are types of content. The carrier of a map is either a globe or a sheet, but map is not a carrier, it is a type of Expression, as is text.

It turns out that cataloging has been mixing content and carrier descriptions in the extent area for ... well, perhaps forever.
1 map on 4 sheets
1 atlas (xvii, 37 pages, 74 leaves of plates)
1 vocal score (x, 190 pages)

In addition, when describing books the carrier isn't mentioned at all, just the content:
xvii, 323 pages
unless there is no extent of the content to record, at which point the book is called a "volume:"
1 volume (unpaged)
I have no doubt that there are clear rules that cover all of this, telling catalogers how to formulate these statements. Yet I am totally perplexed about how to turn this into a coherent data format. In FRBR, there is something called "extent of content" as an attribute of the Expression entity:

4.3.8 Extent of the Expression
The extent of an expression is a quantification of the intellectual content of the expression (e.g., number of words in a text, statements in a computer program, images in a comic strip, etc.). For works expressed as sound and/or motion the extent may be a measure of duration (e.g., playing time).

while "extent of carrier" is an attribute of the Manifestation entity:

4.4.10 Extent of the Carrier
The extent of the carrier is a quantification of the number of physical units making up the carrier (e.g., number of sheets, discs, reels, etc.).

RDA does not have "extent of content," in part (I am told) because it would have separated the instructions for formulating the extent of content and carrier between chapters 7 and 3, respectively, and thus made it difficult for catalogers to create this mixed statement. Of course, one possible response might be that we shouldn't be creating a mixed statement, but two separate statements that could be displayed together as desired. These statements should probably also be linked to the content or carrier vocabulary term that is now carried in MARC 336, 337, or 338.

I looked at ONIX to see how this might have been handled by another bibliographic schema, and it appears that ONIX has two different measures: extent, which is used for extent of the content, and measure, which measures the physical item.

We have to clear up inconsistencies of this nature if we hope to produce a rational format or framework for bibliographic data. Dragging along practices from the past will result in poor quality data that cannot interact well with data from any other sources.

I will add this to the analysis of MARC on the futurelib wiki.

* I can't find "box" anywhere in any list, but perhaps I am missing something.

** Extent of Text is even more complex than I thought. Here is the list of terms to be used:
approximately
case
column
folded
in various foliations
in various numberings
in various pagings
incomplete
leaf
page
portfolio
sheet
unnumbered pages
unnumbered sequence of pages
volume
volume (loose-leaf)
These seem to not be extent of the text itself but the gathering of paper that something, mostly text possibly, is printed on. Volume is a carrier, as is leaf or page or case. However, approximately is totally out of place. Incomplete seems to be a statement about the content, although I suppose you could say that the carrier is incomplete when pages are missing. Note that sheet is here, but not in the list for cartographic resources, so it seems that in describing the carrier for a map one would use sheet from Extent of Text.

Friday, April 06, 2012

If not RDF, then what?

There's no question that the data format known as RDF is darned difficult. Let's suppose that we in the library world decide not to hitch our wagon to RDF, but would still like to create a new bibliographic framework. After all, if MARC simply won't work for the creation of RDA records, we still need something besides MARC that we can use to create data. And even if (although this is unlikely) we should decide not to move to RDA, our records still need some upgrading to fit better into current data processing models. We still need to:
  • define our entities
  • use data wherever possible, not text
  • use identifiers for things
  • relate attributes to entities (that is, say things about some thing)
  • use a mainstream serialization

Should we do this, the mainstream serialization could be anything from JSON to XML to RDF. In fact, it could be all of those if we play our cards right and define our data in a format neutral way. RDA does some of this for us, but not all. In particular, RDA does not distinguish between data and text, and although it allows for the use of identifiers it doesn't give any guidance on how to use them. RDA is probably fine as guidance rules for decision-making, but it needs the corresponding data definition before it becomes useful. Having that data definition could help to clarify some ambiguities in RDA. We have to expect that there will need to be some iteration between RDA and a data definition. (I will post shortly on a problem that I have run into.)

It also seems to me that we have everything to gain by beginning our work on a data format with no particular serialization in mind. We could go from RDA to RDA-as-data and then on to RDA-as-RDF. I see some dangers in skipping the middle step, mainly that we could end up making some decisions that fit RDA into RDF but that are problematic for other serializations.

VIAF gets serious

There has been an announcement by OCLC that the Virtual International Authority File (VIAF) is "transitioning" to OCLC. Since it was already being run by OCLC, this may seem like no news, but in fact it is a sign of commitment on OCLC's part both to VIAF and to linked data. The announcement states:
The change has been made to assure that VIAF will be well-positioned to scale efficiently as a long-term, cooperative activity. The transition also assures that http://viaf.org will continue to have appropriate infrastructure to respond to rising levels of traffic as VIAF gains momentum and popularity as a resource for library authority work and linked data activities.
 One of the many advantages of linked data is that you can role out your implementation of linked data gradually. This announcement about VIAF leads me to wonder if library linked data can't begin with names, followed perhaps closely by subject lists. The question then becomes how we can link from name data on the web to library catalogs. There are already links from VIAF to Wikipedia, Wikipedia to VIAF. This means that there are also links from DBPedia to VIAF, since DBPedia is a linked data form of Wikipedia. DBPedia is the center point of the linked data cloud, thus assuring maximum linking. After that, we reach a dead end, at least as far as library data is concerned, because many library systems do not make use of the authority file identifiers for names, and none, as far as I know, support the use of URIs as identifiers. We would need to pass through VIAF to get the appropriate text string to match against library data.

As for subject access, there are a number of library subject heading lists that have been coded in SKOS and that have some interlinking. The ones I know about are:
These can be found at the Data Hub under the "Bibliographic" group.

(I should mention that VIAF is being licensed as ODC-BY, meaning even commercial use is allowed. What isn't immediately clear is whether the 'BY' -- that is, attribution -- applies to the service as a whole or to individual records or even to individual headings. Attribution adds a small but significant burden on downstream users, so it would be great if the only requirement were that one acknowledge VIAF in appropriate documentation.)

Ideas on how to proceed from this point are very welcome.

Saturday, March 31, 2012

Can libraries change?



The declaration by Library of Congress that the time has come to make the long overdue change to a new data format has rocked the library world. A common reaction is: "How can we do that, when we have 1) thousands of library systems that are designed primarily to work with MARC records 2) no money to pay for a major change, and 3) have no clear idea what we should be heading towards." This fear is increased by the fact that so far there is little public evidence of activity on this project.

A change of this nature is huge. It's not quite Europe converting to the Euro, but within the library world it is a change of the magnitude of converting the Internet from IPv4 to IPv6. We've made other big changes in the past, in particular the change from the card catalog to the OPAC. That effort required us to purchase new systems and to convert the whole of the printed card catalog to the MARC format. Amazingly, it took only about decade to complete that transition.

However, here is the key difference:  that change was entirely internal to the library community. This next one has an additional complication brought on by the fact that the target environment for the future of library data is the Web. This new framework will need to be integrated with that massively complex environment of networked information. This adds unknowns ("How will Web users interact with library data?"), but it also affords some possibilities that we didn't have with the change to MARC. Mainly, it allows us to make use of existing Web technology and the Web community for help in both designing and implementing the change.

What this means to me is that this is not a "library-only" activity that we are embarking on. Unknown numbers of users and systems will want or need to make use of library data on the Web. At least we hope so. Right now, Web services in need of bibliographic data often point to Amazon. Others rely on "crowd-sourced" solutions like Mendeley or Zotero. What will make library data most useful and usable to the larger community? This isn't a question we should be asking of libraries, but of potential users of the data.

There are other important questions we should be thinking about. How will we test whether the new framework is well designed for system functionality and efficiency? How will we convert from what we have to this new framework? What structures must we put in place to maintain and extend the framework over time?

It seems very unlikely that the Library of Congress can address all of this in the 18-month period that is allotted for this work (of which perhaps 12 months remain) because: a) their focus is understandably primarily on the needs of their organization and b) this effort, to be successful, must have input from organizations that are external to the library community. That’s a very tall order for short time span.

None of this should be taken to imply that LC doesn't have smart, skilled staff to work on this -- they do. But if you've ever taken on a large project in your institution you know that the staff working on the project is also doing much of the day-to-day work that occupied them before the project was begun. Few are able to dedicate 100% of their time to a new effort. The question therefore becomes: How can a larger community help LC with this project, taking on appropriate task areas?

I have in mind a set of tasks that could be worked on in parallel, by a number of different interested constituencies and with some good coordination. More details are needed, but the big picture is something like this:




The Web track is an obvious one, especially given that the W3C has already shown an interest in facilitating the entry of library data onto the Semantic Web. There is also a growing realization in the library community that we need to fairly quickly begin to build on the foundation standards developed by the W3C Semantic Web activity. There appears to be a similar awareness in the Semantic Web community that library data presents interesting challenges. For example, library data has revealed that an approach to authority data is needed that cannot currently be provided by SKOS (Simple Knowledge Organization System). Discussions on lists that focus on the Semantic Web make it clear that our early efforts in defining library data in RDF are helping to inform the thinking of the experts in linked data creation.

The bibliographic description track is of course the key one for libraries. This to me is the solid ground of LC, along with its community partners: to determine the semantics of the data that libraries will use to describe their resources and to provide access for users. RDA already does a great deal of this but the task ahead is to make sure that one can express those concepts in a new data format. There should also be an analysis of the cataloging workflow and even of the expected functionality of a cataloging interface. The requirements arising out of this track will inform the work of the Web and IT tracks as they help the Library convert these requirements to implementable structures and applications.

The IT track is absolutely essential: How do we assure that we have data structures that work well with the entire gamut of library systems functions, from acquisitions through circulation? One question I have in particular is about the efficiency of a large bibliographic database structured around the FRBR entities. Efficiency must consider more than just the creation of the bibliographic data, it also must be efficient for the retrieval and display of that data. The report on the Future of Bibliographic Control recommended testing of FRBR. RDA has served to test many of the FRBR concepts, but as yet there is no proof of concept of a data structure that uses the FRBR entities. This track, as I see it, would involve library systems vendors as well as some computer professionals who work with "big data" and semantic web technologies.

The management track is very important but in LC's plan it might be relegated to a later phase since it appears to deal mainly with future activities such as maintenance and modification of the standard. This, however, would be a mistake because the standards for maintenance and extension must be in place from the very beginning of the new framework. I would even say that the new framework should be developed from the beginning with a core and extensions. This eliminates the need to have on opening day a standard that is "everything for everybody," and could allow for a phased implementation of the framework. Note that there are some immediate issues in RDA that require a maintenance standard, such as how to handle open-ended controlled lists in a way that would be compatible with Semantic Web standards.

A critical part of this is the coordination between all of these activities. Such a role, however, is not unusual in a large IT project where work is spread across groups with intersecting milestones.

It seems to me that a division of this nature (and not necessarily exactly how I have described it here) would relieve LC of some of the work that it is undoubtedly currently considering taking on; it could increase the speed with which the full design could be completed; and in my opinion it has the possibility to produce a higher quality solution than could be achieved by a single organization. Logical participants include NISO (both in its role as the standards body that manages the MARC standard and in its role as a focus for the library technology community), W3C's Semantic Web community, Dublin Core Metadata Initiative (which is working on standards for application profiles in RDF), and IFLA (which now has a Semantic Web interest group). There is also some possible synergy with projects like the Internet Archive, the Digital Public Library of America, schema. org, and the Zotero community.  Clearly funding would be needed, and that's also not a simple task.

My concern is that if we don't organize ourselves in this way, that come January, 2013 we will not be anywhere near having the ability to create bibliographic data in a new framework. RDA will be implemented inadequately in MARC and, as that solution is the path of least resistance, work to create a new framework will slow to a crawl. If we don't step up to this task, for many years to come we will continue to see library data housed in frameworks and silos that are invisible to most information seekers.  That would indeed be very unfortunate.

Note: Planned session for ELAG2012 to be led by Lukas Koster with a very similar approach, and with the intention of delivering ideas to LC for the new framework.

Monday, February 27, 2012

What's the question?

I've been meaning to comment on this for a while... If you receive the New York Times in hard copy, and if, like some of us, you turn perhaps too quickly to the page with the famed "Crossword, edited by Will Shortz," for quite a while now you have seen Google's addition to the "puzzle page."

First, let me describe the page, in case you are not an aficionado. Along side the remainder of one or more articles begun on an earlier page, the page contains the aforesaid famed crossword puzzle, two "KenKen" math puzzles, and two "adverpuzzles": the Jeopardy "Clue of the day" and the "Google a day."

The interesting thing is the difference between the two "adverpuzzles." The Jeopardy one gives you one of the answers that will be used on that evening's Jeopardy show. (In Jeopardy, for those who are living in a different culture to mine, you are given an answer, and you must come up with the question.) The Jeopardy adverpuzzle is one column wide (about 2 inches) and about 5 inches high.

The Google one is more than one quarter of the page. It's about 5 inches wide by 11 inches high. Much of that is blank space. And nothing says "We've got more money than we know what to do with" than a daily purchase of blank space in the New York Times.

The other interesting difference is that the Jeopardy puzzle tests your knowledge. It gives you a difficult topic and you are supposed to come up with the answer. For example, today's Jeopardy answer is:
"No day shall erase you from the memory of time," from Virgil's Aeneid, is inscribed on a wall at this memorial."
The Google puzzle invites you to look up the answer on Google. It even provides a specific site for you to use, one that won't be tainted by the other users looking up the same answer.  There are no points for knowing the answer.

The third difference is that to find out if you got the right answer on the Jeopardy question you have to watch that evening's show. To get the answer from Google a Day you check the next day. But, presumably, you've already spent some time at http://agoogleaday.com/ looking for the answer. Here's today's answer to the previous question:

Yesterday's A Google a Day: If you compare the half-lives of cesium-137 and uranium-238, which one outlives the other?
How to find the answer: Search [half-life cesium-137] to find that it's 30 years. Search [half-life uranium-238] to learn that it's 4.5 billion years, which is just a bit longer.
 Maybe I'm making too much of this, but I see two conflicting cultures here: the one of knowing things, and the one of looking things up. It makes me wonder if in a few years there will be a hit TV show where contestants vie to see who can look it up the fastest. Heck, I don't know why we don't have such a show already. Knowing is definitely "old school," and as a librarian I am firmly ensconced in the "look it up" culture. But I have a strong gut reaction, a negative one, to becoming totally dependent on a network connection for knowledge. It could just happen that I could find myself out in some wilderness area with no satellite signal and a life-or-death need to know the half-lives of certain elements on the periodic table. And then what would I do?