Showing posts with label linked data. Show all posts
Showing posts with label linked data. Show all posts

Wednesday, April 12, 2017

If It Ain't Broke

For the first time in over forty years there is serious talk of a new metadata format for library bibliographic data. This is an important moment.

There is not, however, a consensus within the profession on the need to replace the long-standing MARC record format with something different. A common reply to the suggestion that library data creation needs a new data schema is the phrase: "If it ain't broke, don't fix it." This is more likely to be uttered by members of the cataloging community - those who create the bibliographic data that makes up library catalogs - than by those whose jobs entail systems design and maintenance. It is worth taking a good look at the relationship that catalogers have with the MARC format, since their view is informed by decades of daily encounters with a screen of MARC encoding.

Why This Matters

When the MARC format was developed, its purpose was clear: it needed to provide the data that would be printed on catalog cards produced by the Library of Congress. Those cards had been printed for over six decades, so there was no lack of examples to use to define the desired outcome. In ways unimagined at the time, MARC would change, nay, expand the role of shared cataloging, and would provide the first online template for cataloging.

Today work is being done on the post-MARC data schema. However, how the proposed new schema might change the daily work of catalogers is unclear. There is some anxiety in the cataloging community about this, and it is understandable. What I unfortunately see is a growing distrust of this development on the part of the data creators in our profession. It has not been made clear what their role is in the development of the next "MARC," not even whether their needs are a driving force in that development. Surely a new model cannot be successful without the consideration (or even better, the participation) of the people who will spend their days using the new data model to create the library's data.

(An even larger question is the future of the catalog itself, but I hardly know where to begin on that one.)


If it Ain't Broke...

The push-back against proposed post-MARC data formats is often seen as a blanket rejection of change. Undoubtedly this is at times the case. However, given that there have now been multiple generations of catalogers who worked and continue to work with the MARC record, we must assume that the members of the cataloging community have in-depth knowledge of how that format serves the cataloging function. We should tap that knowledge as a way to understand the functionality in MARC that has had a positive impact on cataloging for four decades, and should study how that functionality could be carried forward into the future bibliographic metadata schema.

I asked on Twitter for input on what catalogers like about MARC, and received some replies. I also viewed a small number of presentations by catalogers, primarily those about proposed replacements for MARC. From these I gathered the following list of "what catalogers like about MARC." I present these without comment or debate. I do not agree with all of the statements here, but that is no matter; the purpose here is to reflect cataloger perspectives.

(Note: This list is undoubtedly incomplete and I welcome comments or emails with your suggestions for additions or changes.)


What Catalogers Like/Love About MARC



There is resistance to moving away from using the MARC record for cataloging among some in the Anglo-American cataloging community. That community has been creating cataloging data in the MARC formats for forty years. For these librarians, MARC has many positive qualities, and these are qualities that are not perceived to exist in the proposals for linked data. (Throughout the sections below, read "library cataloging" and variants as referring to the Anglo-American cataloging tradition that uses the MARC format and the Anglo-American Cataloging Rules and its newer forms.)

MARC is Familiar

Library cataloging makes use of a very complex set of rules that determine how a resource is described. Once the decisions are made regarding the content of the description, those results are coded in MARC. Because the creation of the catalog record has been done in the MARC format since the late 1970's, working catalogers today have known only MARC as the bibliographic record format and the cataloging interface. Catalogers speak in "MARC" - using the tags to name data elements - e.g. "245" instead of "title proper".

MARC is WYSIWYG

Those who work with MARC consider it to be "human readable." Most of the description is text, therefore what the cataloger creates is exactly what will appear on the screen in the library catalog. If a cataloger types "ill." that is what will display; if the cataloger instead types "illustrations" then that is what will display. In terms of viewing a MARC record on a screen, some cataloger displays show the tags and codes to one side, and the text of those elements is clearly readable as text.

MARC Gives Catalogers Control

The coding is visible, and therefore what the cataloger creates on the screen is virtually identical to the machine-readable record that is being created. Everything that will be shown in the catalog is in the record (with the exception of cover art, at least in some catalogs). The MARC rules say that the order of fields and subfields in the record are the order in which that information should be displayed in the catalog. Some systems violate this by putting the fields in numeric order, but the order of subfields is generally maintained. Catalogers wish to control the order of display and are frustrated when they cannot. In general, changing anything about the record with automated procedures can un-do the decisions made by catalogers as part of their work, and is a cause of frustration for catalogers.

MARC is International

MARC is used internationally, and because the record uses numerics and alphanumeric codes, a record created in another country is readable to other MARC users. Note that this was also the purpose of the International Standard Bibliographic Description (ISBD), which instead of tags uses punctuation marks to delimit elements of the bibliographic description. If a cataloger sees this, but cannot read the text:

  245 02   |a לטוס עם עין אחת / |c דני בז.

it is still clear that this is a title field with a main title (no subtitle), followed by a statement of the author's name as provided on the title page of the book.

MARC is the Lingua Franca of Cataloging

This is probably the key point that comprises all of the above, but it is important to state it as such. This means that the entire workflow, the training materials, the documentation - all use MARC. Catalogers today think in MARC and communicate in MARC. This also means that MARC defines the library cataloging community in the way that a dialect defines the local residents of a region. There is pride in its "library-ness". It is also seen as expressing the Anglo-American cataloging tradition.

MARC is Concise

MARC is concise as a physical format (something that is less important today than it was in the 1960s when MARC was developed), and it is also concise on the screen. "245" represents "title proper"; "240" represents "uniform title"; "130" represents "uniform title main entry". Often an entire record can be viewed on a single screen, and the tags and subfield codes take up very little display space.

MARC is Very Detailed

MARC21 has about 200 tags currently defined, and each of these can have up to 36 subfields. There are about 2000 subfields defined in MARC21, although the distribution is uneven and depends on the semantics of the field; some fields have only a handful of subfields, and in others there are few codes remaining that could be assigned.

MARC is Flat

The MARC record is fairly flat, with only two levels of coding: field and subfield. This is a simple model that is easy to understand and easy to visualize.

MARC is Extensible

Throughout its history, the MARC record has been extended by adding new fields and subfields. There are about 200 defined fields which means that there is room to add approximately 600 more.

MARC has Mnemonics

Some coding is either consistent or mnemonic, which makes it easier for catalogers to remember the meaning of the codes. There are code blocks that refer to cataloging categories, such as the title block (2XX), the notes block (5XX) and the subject block (6XX). Some subfields have been reserved for particular functions, such as the use of the numeric subfields in 0-8. In other cases, the mnemonic is used in certain contexts, such as the use of subfield "v" for the volume information of series. In other fields, the "v" may be used for something else, such as the "form" subfield in subject fields, but the context makes it clear.

There are also field mnemonics. For example, all tagged fields that have "00" in the second and third places are personal name fields. All fields and subfields that use the number 9 are locally defined (with a few well-known exceptions).

MARC is Finite and Authoritative

MARC defines a record that is bounded. What you see in the record is all of the information that is being provided about the item being described. The concept of "infinite graphs" is hard to grasp, and hard to display on a screen. This also means that MARC is an authoritative statement of the library bibliographic description, whereas graphs may lead users to sources that are not approved by or compatible with the library view.

Friday, January 16, 2015

Real World Objects

I was asked a question about the meaning and import of the RDF concept of "Real World Object" (RWO) and didn't give a very good answer off the cuff. I'll try to make up for that here.

The concept of RWO comes out of the artificial intelligence (AI) community. Imagine that you are developing robots and other machines that must operate within the same world that you and I occupy. You have to find a way to "explain," in a machine-operational way, everything in our world: stairs and ramps, chairs and tables, the effect of gravity on a cup when you miss placing it on the table, the stars, love and loyalty (concepts are also objects in this view). The AI folks have actually set a goal to create such descriptions, which they call ontologies, for everything in the world; for every RWO.

You might consider this a conceit, or a folly, but that's the task they have set for themselves.

The original Scientific American article that described the semantic web used as its example intelligent 'bots that would manage your daily calendar and make appointments for you. This was far short of the AI "ontology of everything" but the result that matters to us now is that there have been AI principles baked into the development of RDF, including the concept of the RWO.

RWO isn't as mysterious as it may seem, and I can provide a simple example from our world. The MARC record for a book has the book as its RWO, and most of its data elements "speak" about the book. At the same time, we can say things about the MARC record, such as who originally created it, and who edited it last, and when. The book and the record are different things, different RWO's in an RDF view. That's not controversial, I would assume.

Our difficulties arise because in the past we didn't have a machine-actionable way to distinguish between those two "things": the book and the record. Each MARC record got an identifier, which identified the record. We've never had identifiers for the thing the record describes (although the ISBN sometimes works this way). It has always been safe to assume that the record was about the book, and what identified the book was the information in the record. So we obviously have a real world object, but we didn't give it its own identifier - because humans could read the text of the record and understand what it meant (most of the time or some of the time).

I'm not fully convinced that everything can be reduced to RWO/not-RWO, and so I'm not buying that is the only way to talk about our world and our data. It should be relatively easy, though, without getting into grand philosophical debates, to determine the difference between our metadata and the thing it describes. That "thing it describes" can be fuzzy in terms of the real world, such as when the spirit of Edgar Cayce speaks through a medium and writes a book. I don't want to have to discuss whether the spirit of Edgar Cayce is real or not. We can just say that "whoever authors the book is as real as it gets." So if we forget RWO in the RDF sense and just look sensibly at our data, I'm sure we can come to a practical agreement that allows both the metadata and the real world object to exist.

That doesn't resolve the problem of identifiers, however, and for machine-processing purposes we do need separate identifiers for our descriptions and what we are describing.* That's the problem we need to solve, and while we may go back and forth a bit on the best solution, the problem is tractable without resorting to philosophical confabulations.

* I think that the multi-level bibliographic descriptions like FRBR and BIBFRAME make this a bit more complex, but I haven't finished thinking about that, so will return if I have a clearer idea.

Monday, November 24, 2014

Multi-Entity Models.... Baker, Coyle, Petiya

Multi-Entity Models of Resource Description in the Semantic Web: A comparison of FRBR, RDA, and BIBFRAME
by Tom Baker, Karen Coyle, Sean Petiya
Published in: Library Hi Tech, v. 32, n. 4, 2014 pp 562-582 DOI:10.1108/LHT-08-2014-0081
Open Access Preprint

The above article was just published in Library hi Tech. However, because the article is a bit dense, as journal articles tend to be, here is a short description of the topic covered, plus a chance to reply to the article.

We now have a number of multi-level views of bibliographic data. There is the traditional "unit card" view, reflected in MARC, that treats all bibliographic data as a single unit. There is the FRBR four-level model that describes a single "real" item, and three levels of abstraction: manifestation, expression, and work. This is also the view taken by RDA, although employing a different set of properties to define instances of the FRBR classes. Then there is the BIBFRAME model, which has two bibliographic levels, work and instance, with the physical item as an annotation on the instance.

In support of these views we have three RDF-based vocabularies:

FRBRer (using OWL)
RDA (using RDFS)
BIBFRAME (using RDFS)

The vocabularies use a varying degree of specification. FRBRer is the most detailed and strict, using OWL to define cardinality, domains and ranges, and disjointness between classes and between properties. There are, however, no sub-classes or sub-properties. BIBFRAME properties all are defined in terms of domains (classes), and there are some sub-class and sub-property relationships. RDA has a single set of classes that are derived from the FRBR entities, and each property has the domain of a single class. RDA also has a parallel vocabulary that defines no class relationships; thus, no properties in that vocabulary result in a class entailment. [1]

As I talked about in the previous blog post on classes, the meaning of classes in RDF is often misunderstood, and that is just the beginning of the confusion that surrounds these new technologies. Recently, Bernard Vatant, who is a creator of the Linked Open Vocabularies site that does a statistical analysis of the existing linked open data vocabularies and how they relate to each other, said this on the LOV Google+ group:
"...it seems that many vocabularies in LOV are either built or used (or both) as constraint and validation vocabularies in closed worlds. Which means often in radical contradiction with their declared semantics."
What Vatant is saying here is that many vocabularies that he observes use RDF in the "wrong way." One of the common "wrong ways" is to interpret the axioms that you can define in RDFS or OWL the same way you would interpret them in, say, XSD, or in a relational database design. In fact, the action of the OWL rules (originally called "constraints," which seems to have contributed to the confusion, now called "axioms") can be entirely counter-intuitive to anyone whose view of data is not formed by something called "description logic (DL)."

A simple demonstration of this, which we use in the article, is the OWL axiom for "maximum cardinality." In a non-DL programming world, you often state that a certain element in your data is limited to the number of times it can be used, such as saying that in a MARC record you can have only one 100 (main author) field. The maximum cardinality of that field is therefore "1". In your non-DL environment, a data creation application will not let you create more than one 100 field; if an application receiving data encounters a record with more than one 100 field, it will signal an error.

The semantic web, in its DL mode, draws an entirely different conclusion. The semantic web has two key principles: open world, and non-unique name. Open world means that whatever the state of the data on the web today, it may be incomplete; there can be unknowns. Therefore, you may say that you MUST have a title for every book, but if a look at your data reveals a book without a title, then your book still has a title, it is just an unknown title. That's pretty startling, but what about that 100 field? You've said that there can only be one, so what happens if there are 2 or 3 or more of them for a book? That's no problem, says OWL: the rule is that there is only one, but the non-unique name rule says that for any "thing" there can be more than one name for it. So when an OWL program [2] encounters multiple author 100 fields, it concludes that these are all different names for the same one thing, as defined by the combination of the non-unique name assumption and the maximum cardinality rule: "There can only be one, so these three must really be different names for that one." It's a bit like Alice in Wonderland, but there's science behind it.

What you have in your database today is a closed world, where you define what is right and wrong; where you can enforce the rule that required elements absolutely HAVE TO be there; where the forbidden is not allowed to happen. The semantic web standards are designed for the open world of the web where no one has that kind of control. Think of it this way: what if you put a document onto the open web for anyone to read, but wanted to prevent anyone from linking to it? You can't. The links that others create are beyond your control. The semantic web was developed around the idea of a web (aka a giant graph) of data. You can put your data up there or not, but once it's there it is subject to the open functionality of the web. And the standards of RDFS and OWL, which are the current standards that one uses to define semantic web data, are designed specifically for that rather chaotic information ecosystem, where, as the third main principle of the semantic web states, "anyone can say anything about anything."

I have a lot of thoughts about this conflict between the open world of the semantic web and the needs for closed world controls over data; in particular whether it really makes sense to use the same technology for both, since there is such a strong incompatibility in underlying logic of these two premises. As Vatant implies, many people creating RDF data are doing so with their minds firmly set in closed world rules, such that the actual result of applying the axioms of OWL and RDF on this data on the open web will not yield the expected closed world results.

This is what Baker, Petiya and I address in our paper, as we create examples from FRBRer, RDA in RDF, and BIBFRAME. Some of the results there will probably surprise you. If you doubt our conclusions, visit the site http://lod-lam.slis.kent.edu/wemi-rdf/ that gives more information about the tests, the data and the test results.

[1] "Entailment" means that the property does not carry with it any "classness" that would thus indicate that the resource is an instance of that class.

[2] Programs that interpret the OWL axioms are called "reasoners". There are a number of different reasoner programs available that you can call from your software, such as Pellet, Hermit, and others built into software packages like TopBraid.

Tuesday, November 18, 2014

Classes in RDF

RDF allows one to define class relationships for things and concepts. The RDFS1.1 primer describes classes succinctly as:
Resources may be divided into groups called classes. The members of a class are known as instances of the class. Classes are themselves resources. They are often identified by IRIs and may be described using RDF properties. The rdf:type property may be used to state that a resource is an instance of a class.
This seems simple, but it is in fact one of the primary areas of confusion about RDF.

If you are not a programmer, you probably think of classes in terms of taxonomies -- genus, species, sub-species, etc. If you are a librarian you might think of classes in terms of classification, like Library of Congress or the Dewey Decimal System. In these, the class defines certain characteristics of the members of the class. Thus, with two classes, Pets and Veterinary science, you can have:
Pets
- dogs
- cats

Veterinary science
- dogs
- cats
In each of those, dogs and cats have different meaning because the class provides a context: either as pets, or information about them as treated in veterinary science.

For those familiar with XML, it has similar functionality because it makes use of nesting of data elements. In XML you can create something like this:
<drink>
    <lemonade>
        <price>$2.50</price>
        <amount>20</amount>
    </lemonade>
    <pop>
        <price>$1.50</price>
        <amount>10</amount>
    </pop>
</drink>
and it is clear which price goes with which type of drink, and that the bits directly under the <drink> level are all drinks, because that's what <drink> tells you.

Now you have to forget all of this in order to understand RDF, because RDF classes do not work like this at all. In RDF, the "classness" is not expressed hierarchically, with a class defining the elements that are subordinate to it. Instead it works in the opposite way: the descriptive elements in RDF (called "properties") are the ones that define the class of the thing being described. Properties carry the class information through a characteristic called the "domain" of the property. The domain of the property is a class, and when you use that property to describe something, you are saying that the "something" is an instance of that class. It's like building the taxonomy from the bottom up.

This only makes sense through examples. Here are a few:
1. "has child" is of domain "Parent".

If I say "X - has child - 'Fred'" then I have also said that X is a Parent because every thing that has a child is a Parent.

2. "has Worktitle" is of domain "Work"

If I say "Y - has Worktitle - 'Der Zauberberg'" then I have also said that Y is a Work because every thing that has a Worktitle is a Work.

In essence, X or Y is an identifier for something that is of unknown characteristics until it is described. What you say about X or Y is what defines it, and the classes put it in context. This may seem odd, but if you think of it in terms of descriptive metadata, your metadata describes the "thing in hand"; the "thing in hand" doesn't describe your metadata. 

Like in real life, any "thing" can have more than one context and therefore more than one class. X, the Parent, can also be an Employee (in the context of her work), a Driver (to the Department of Motor Vehicles), a Patient (to her doctor's office). The same identified entity can be an instance of any number of classes.
"has child" has domain "Parent"
"has licence" has domain "Driver"
"has doctor" has domain "Patient"

X - has child - "Fred"  = X is a Parent 
X - has license - "234566"  = X is a Driver
X - has doctor - URI:765876 = X is a Patient
Classes are defined in your RDF vocabulary, as as the domains of properties. The above statements require an application to look at the definition of the property in the vocabulary to determine whether it has a domain, and then to treat the subject, X, as an instance of the class described as the domain of the property. There is another way to provide the class as context in RDF - you can declare it explicitly in your instance data, rather than, or in addition to, having the class characteristics inherent in your descriptive properties when you create your metadata. The term used for this, based on the RDF standard, is "type," in that you are assigning a type to the "thing." For example, you could say:
X - is type - Parent
X - has child - "Fred"
This can be the same class as you would discern from the properties, or it could be an additional class. It is often used to simplify the programming needs of those working in RDF because it means the program does not have to query the vocabulary to determine the class of X. You see this, for example, in BIBFRAME data. The second line in this example gives two classes for this entity:
<http://bibframe.org/resources/FkP1398705387/8929207instance22>
a bf:Instance, bf:Monograph .

One thing that classes do not do, however, is to prevent your "thing" from being assigned the "wrong class." You can, however, define your vocabulary to make "wrong classes" apparent. To do this you define certain classes as disjoint, for example a class of "dead" would logically be disjoint from a class of "alive." Disjoint means that the same thing cannot be of both classes, either through the direct declaration of "type" or through the assignment of properties. Let's do an example:
"residence" has domain "Alive"
"cemetery plot location" has domain "Dead"
"Alive" is disjoint "Dead" (you can't be both alive and dead)

X - is type - "Alive"                                         (X is of class "Alive")
X - cemetery plot location - URI:9494747      (X is of class "Dead")
Nothing stops you from creating this contradiction, but some applications that try to use the data will be stumped because you've created something that, in RDF-speak, is logically inconsistent. What happens next is determined by how your application has been programmed to deal with such things. In some cases, the inconsistency will mean that you cannot fulfill the task the application was attempting. If you reach a decision point where "if Alive do A, if Dead do B" then your application may be stumped and unable to go on.

All of this is to be kept in mind for the next blog post, which talks about the effect of class definitions on bibliographic data in RDF.


Note: Multiple domains are treated in RDF as an AND (an intersection). Using a library-ish example, let's assume you want to define a note field that you can use for any of your bibliographic entities. For this example, we'll define entities Work, Person, Manifestation for ex:note. You define your note property something like:

ex:note
    a rdf:Property ;
    rdfs:label "Note"@en ;
    rdfs:domain ex:Work ;
    rdfs:domain ex:Person ;
    rdfs:domain ex:Manifestation ;
    rdfs:range rdfs:literal .

Any subject on which you use ex:note would be inferred to be, at the same time, Work and Person and Manifestation - which is manifestly illogical. There is not a way to express the rule: "This property CAN be used with these classes" in RDF. For that, you will need something that does not yet exist in RDF, but is being worked on in the W3C community, which is a set of rules that would allow you to validate property usage. You might also want see what schema.org has done for domain and range.

Saturday, October 25, 2014

Citations get HOT

The Public Library of Science research section, PLOSLabs (ploslabs.org) has announced some very interesting news about the work that they are doing on citations, which they are calling "Rich Citations".

Citations are the ultimate "linked data" of academia, linking new work with related works. The problem is that the link is human-readable only and has to be interpreted by a person to understand what the link means. PLOS Labs have been working to make those citations machine-expressive, even though they don't natively provide the information needed for a full computational analysis.

Given what one does have in a normal machine-readable document with citations, they are able to pull out an impressive amount of information:
  • What section the citation is found in. There is some difference in meaning whether a citation is found in the "Background" section of an article, or in the "Methodology" section. This gives only a hint to the meaning of the citation, but it's more than no information at all.
  • How often a resource is cited in the article. This could give some weight to its importance to the topic of the article.
  • What resources are cited together. Whenever a sentence ends with "[3][7][9]", you at least know that those three resources equally support what is being affirmed. That creates a bond between those resources.
  • ... and more
As an open access publisher, they also want to be able to take users as directly as possible to the cited resources. For PLOS publications, they can create a direct link. For other resources, they make use of the DOI to provide links. Where possible, they reveal the license of cited resources, so that readers can know which resources are open access and which are pay-walled.

This is just a beginning, and their demo site, appropriately named "alpha," uses their rich citations on a segment of the PLOS papers. They also have an API that developers can experiment with.

I was fortunate to be able to spend a day recently at their Citation Hackathon where groups hacked on ongoing aspects of this work. Lots of ideas floated around, including adding abstracts to the citations so a reader could learn more about a resource before retrieving it. Abstracts also would add search terms for those resources not held in the PLOS database. I participated in a discussion about coordinating Wikidata citations and bibliographies with the PLOS data.

Being able to datamine the relationships inherent in the act of citation is a way to help make visible and actionable what has long been the rule in academic research, which is to clearly indicate upon whose shoulders you are standing. This research is very exciting, and although the PLOS resources will primarily be journal articles, there are also books in their collection of citations. The idea of connecting those to libraries, and eventually connecting books to each other through citations and bibliographies, opens up some interesting research possibilities.

Saturday, June 28, 2014

Linking, really linking

Some time ago I posted a bit of wishful thinking about linking search engine results to library holdings. Now Overdrive has made this a reality, at least in Bing:

This appears in the "extended information" area of a Bing search for the Girl with a Dragon Tattoo trilogy. This is based on you having an Overdrive account and I believe you may also have to have given the browser permission to use your location.




I don't usually use Bing, and so I was unaware that Bing has made much better use of linked data (in part promoted by the use of schema.org standards) than Google. Here is the Google extended sidebar for the same book:
Google results
Now look at what Bing provides:
Bing results
 
If you can't see that there are advantages to linked data after looking at these examples, then, like the global warming deniers, you just don't want to be confused by the facts. Now to get on to how we can make library offerings as rich as this.

**HT to Eric Hellman for blogging this from ALA.

Saturday, May 10, 2014

Let's link!

Tom Johnson responded to a statement of mine in which I said that we need non-programmer tools for linked data work. He asked for case studies, and so I'm going to do some "off the top of my head" riffing here, just to see what might come of it.

First, let me say a little something about some analogous technology. Let's take HTML. I'm one of those folks who learned HTML many decades ago when all you needed was <p>, <i>, <b> and maybe <hr>. With these, you could create a web page. Web pages in those days didn't have banners, sidebars, tables ... adding an image was going all fancy. There were no WYSIWYG tools because it was very simple to create such a web page. Then we got more sophisticated formatting (server-side includes, side bars, tables(!)), and now there is a whole programming language called CSS to handle page creation (and destruction, since CSS is very complex.) It doesn't take much looking around to understand that most people today need TOOLS to create a web page and lots of tools exist, tools that cost little or nothing (WordPress, Drupal) and can be used by folks who've never written 50 lines of working code.

Essentially, on the web today, a few people are providing the structure and the tools, but most folks are providing content without knowing the guts of the technology. Content is, in my mind, the actual goal of the web; technology is the means.

I've spent a lot of time talking to folks about linked data, but linked data is not itself a goal. I've started trying to move the conversation from the underlying technology to what I see as the real goal: making connections -- connections between concepts, ideas, statements about things. This is inherently more social than technical, but of course it needs the technology behind it in order to work. The easier it is for people to make connections, the more connections they will make.

The problem that I'm seeing today in the linked data space is that we don't have an idea of what kinds of connections people will want to make, or what they will do with them. I don't think we're going to know until we apply some scientific method to the problem -- that is, try, fail, try again, rinse, repeat.

I created a really funky web page with one idea. The page looks like this:



It talks about having the ability to profile types of data that can then be selected in the content of the page. Each of these will then pull additional content (maps, term definitions, biographical information)  from available linkable data into the page. This could begin as a very simple application with only a small number of options, as a proof of concept. (WordPress?) I'm sure that others can greatly improve on it.  Take a look at the display in the FAO AGRIS catalog to get an idea of how this might look with a better display.

Have at it, please.

Monday, October 14, 2013

Who uses Dublin Core? - the original 15

The original 15 Dublin Core elements are included in the Dublin Core Metadata Terms using the namespace http://purl.org/dc/elements/1.1/. There is an "updated" version of each of the original terms in the namespace http://purl.org/dc/terms (dcterms). The difference is that the /dc/terms includes formal domains and ranges, in conformance with linked data standards; the original 15 elements in the /dc/elements/1.1/ namespace have no domain or range constraints defined. This means that the original 15, often given the namespace prefix of "dc:" or "dce:", are compatible with legacy uses of the Dublin Core elements.

In the first post of this series, I showed that the most used terms are from the dcterms vocabulary, followed immediately by a cluster of terms from the dce namespace. In addition, the majority of the top dcterms are the linked data equivalents of the dce terms, thus confirming the "coreness" of the original Dublin Core 15.

From this explanation one might expect that the uses of dce in the wilds linked data would be limited to legacy data. That does not, however, seem to be the case. Out of a total of 125 datasets from the Linked Open Vocabularies,  nearly half (60) use both the linked data vocabulary (dcterms) and the dce terms. Of the top five datasets with the greatest number of uses of dce, only one, "Wikipedia 3," does not also use the dcterms.

Europeana Linked Open Data 
Wikipedia 3 
Linked Open Data Camera dei deputati 
B3Kat - Library Union Catalogues of Bavaria, Berlin and Brandenburg 
Yovisto - academic video search 
 
There are reasons why datasets may use both "generations" of the Dublin Core vocabulary. One is that their data contains a mix of legacy metadata and linked data, either because the dataset has grown over time, or because the set combines data from different sources. Another is that there may be situations in which the dcterms use of domains and ranges is too restrictive for the needs of the data creators.

The LOV dataset of dce usage has over 24 million uses (compared to 192 million uses of dcterms). Library and bibliographic data is again by far the majority of the use, although it is rivaled by government data, in part because of the over 4 million uses contributed by the Italian Camera dei deputati, which also uses dcterms but to a lesser extent. In fact, government data is overall a strong contender in the dce space.

My overall conclusion from looking at this data is that Dublin Core is used widely for bibliographic and non-bibliographic data; that there is a new "core" based on usage that overlaps greatly with the old core; some dcterms elements are hardly used at all in these datasets; and finally that both the linked data dcterms and the legacy dce elements show themselves to be useful, even in the linked data environment.


Related posts:

Friday, October 11, 2013

Who uses Dublin Core - dcterms?

In my previous post I gave some data on Dublin Core field use. Today I look at who is using Dublin Core's dcterms vocabulary.

The LOV statistics show 212 datasets that use the vocabulary at http://purl.org/dc/terms/, and the number of instances of usage. I did some "back of the envelope" counts on what types of organizations or projects use the terms, and also the type of use. By these calculations, the highest use was from libraries (> 60%). The next highest use was in a single language study called Semantic Quran (~10%). Third was the use in government data at less than 1%.

If one looks at the type of data, bibliographic data makes up nearly 90% of the usage. In this category I included archives, eprint repositories, and a few databases of videos and teaching materials.

From this one might conclude that dcterms isn't used much outside of the bibliographic world, but in fact traditional libraries provide only 28 of the 212 datasets on this list. The range of users and uses is impressive. Here are a few to peak your interest:
  • Southampton University has a number of datasets of civic information, including a list of bus stops.
  • There is a biomedical data service called eagle-i used by 24 universities or departments that provides information on specimens, reagents and services. This contributes nearly 500,000 instance of dcterms usage.
  • The New York Times linked data service uses dcterms. This service consists of topics (persons, organizations, locations, topics) covered by the newspaper.
  • I've mentioned the Semantic Quran. This is a linguistic database consisting of 43 translations of the Quran. It contributes over 6 million instances of dcterms use.
  • There is government data covering a wide range of topic areas. By my estimate there are at least 70 sets of government data in this compilation (including international), with everything from the aforementioned bus stops to election data, patents, economic indicators and scientific information.
If one is to make conclusions from this evidence, it could be said that the dcterms vocabulary is a core vocabulary for the description of intellectual resources, such as the holdings of libraries and archives, but that it also provides functionality for a wide range of data types. 

There are also users of the original Dublin Core vocabulary, now referred to as "1.1". I will cover that usage next.

Wednesday, October 09, 2013

Dublin Core usage in LOD

Thanks to some projects that gather statistics on the growth of linked data, we can find out various interesting things about the vocabularies being used and the degree of linking between data sets from different communities. The data I report here comes from LODstats via the Linked Open Vocabularies (LOV) project.

The LOV project looks particularly at the interrelations between vocabularies. For example, it can show which vocabularies use terms from other vocabularies. This crossover of terms is one of the things that makes links between datasets possible. For example, this shows that the geoSpecies vocabulary is not itself referenced by other vocabularies, but can link through its use of vocabularies like FOAF and Dublin Core. You can watch the visualization grow here.
In contrast, this is what Dublin Core terms looks like at LOV:

With the animated visualization here.

Dublin Core does seem to have fulfilled its role as a core vocabulary that many different communities have found useful, at least in part. The set of terms often abbreviated as "dcterms" (or sometimes "dct") and whose namespace is http://purl.org/dc/terms/ has been used approximately 192 million times as reported in the LOD statistics. This is only the usage in the 2289 linked data datasets used by that project. The earlier set of Dublin Core terms, the original fifteen terms, whose namespace is http://purl.org/dc/elements/1.1/, has been used 24.2 million times. This gives us a total of 216 million uses of Dublin Core in this particular count.

The interesting question, then, is what parts of DC are heavily used? I have a sorted list, from most to least, of all terms in the http://purl.org/dc/ namespace. The top fifteen terms are all from the "dcterms" namespace:

count          term
24147876    subject
22575133    identifier
17120343    title
17065873    issued
14459601    publisher
11605978    language
9930733    medium
9795117    format
9792064    BibliographicResource
7700745    isPartOf
7371553    creator
7241777    contributor
6590791    description
6184994    type
5983236    extent

Of this list, only four were not part of the original "Dublin Core 15" vocabulary: issued, medium, BibliographicResource, and isPartOf. The terms of that original vocabulary cluster together beginning right after the last term in the above list. I believe this provides an interesting affirmation that the original fifteen terms were a fair definition of "core." 

However, these terms, in the "dcterms" namespace got less than ten uses, and some were even zero:

accrualPeriodicity
Frequency
AgentClass
dateSubmitted
isRequiredBy
Jurisdiction
LicenseDocument
LinguisticSystem
MediaType
MediaTypeOrExtent
PeriodOfTime
PhysicalResource
RightsStatement

The last term, which got zero in the LOD calculations, is particularly interesting because the element "rights" in the original "DC 15" got 398,361 uses, and is ranked 39th in the list of elements the overall http://purl.org/dc namespace.

Next, I'll take a quick look at which datasets are contributing to the use of Dublin Core terms, and who is creating those datasets.



Tuesday, July 23, 2013

Linked Data First Steps & Catch-21

Often when I am with groups of librarians talking about linked data, this question comes up:
"What can we do TODAY to get ready for linked data?"
It's not really a hard question, because, at least in my mind, there is an obvious starting point: identifiers. We can begin today to connect the textual data in our bibliographic records with identifiers for the same thing or concept.

What identifiers exist? Thanks to the Library of Congress we have identifiers for all of our authority controlled elements: names and subjects. (And if you are outside of the US, look to your national library for their work in this area, or connect to the Virtual International Authority File where you can.) LoC also provides identifiers for a number of the controlled lists used in MARC21 data.

The linked data standards require that identifiers be in the form of an HTTP-based URI. What this means is that your identifier looks like a URL. The identifier for me in the LC name authority file is:
http://id.loc.gov/authorities/names/n89613425
Any bibliographic data with my name in a field should also contain this identifier. (OK, admittedly that's not a lot of bib data.) That brings us to "Catch-21" -- the MARC21 record. Although a control subfield was added to MARC21 for identifiers ($0), that subfield requires the identifier to be in a MARC21-specific format:
The control number or identifier is preceded by the appropriate MARC Organization code (for a related authority record) or the Standard Identifier source code (for a standard identifier scheme), enclosed in parentheses.
The example in the MARC21 documentation is:
100 1#$aBach, Johann Sebastian.$4aut$0(DE-101c)310008891
Modified to use LC name authorities that would be:
 100 1#$aBach, Johann Sebastian,$d1685-1750$0(LoC)n79021425
The contents of the $0 therefore is not a linked data identifier even in those instances where we have a proper linked data identifier for the name.  Catch-21. I therefore suggest that, as an act of Catch-21 disobedience, we all declare that we will ignore the absurdity of having recently added an anti-linked data identifier subfield to our standard, and use it instead for standard HTTP URIs:
100 1#$aBach, Johann Sebastian,$d1685-1750$0http://id.loc.gov/authorities/names/n79021425
Once we've gotten over this hurdle, we can begin to fill in identifiers for authority-controlled elements. Obviously we won't be doing this by hand, one record at a time. This should be part of a normal authority update service, or it may be feasible within systems that store and link national authority data to bibliographic records.

We should also insist that cataloging services that use the national authority files begin to include these subfields in bibliographic data as it is created/downloaded.

Note that because the linked data standard identifiers are HTTP URIs, aka URLs, by including these identifiers in your bibliographic data you have already created a link -- a link to the web page for that person or subject, and a link to a machine-readable form of the authority data in a variety of formats (MARCXML, JSON, RDF, and more). In the LC identifier service, the name authority data includes a link to the VIAF identifier for the person; the VIAF identifier for some persons is included in the Wikipedia page about the person; the Wikipedia identifier links you to DBpedia and the DBpedia identifier is used by Freebase ...

That's how it all gets started, with one identifier that becomes a kind of identifier snowball rolling down hill, collecting more and more links as it goes along.

Pretty easy, eh?

Thursday, February 21, 2013

Open to Creativity

The brilliance of Google's PageRank is not the computational methods behind it but the target of those methods: links created by people in the course of making something meaningful on the Web. Without that human input, Google (and Bing and Yahoo) would be simply counting up term frequencies and perhaps analyzing linguistic characteristics, but missing most of what makes Web searching work. The results would be at least as bad as the ranking on Google Books because it would be devoid of the human commitment of significant connections between the pages.

Although Google's mission statement is "... to organize the world's information and make it universally accessible," Google is not really organizing anything. It is reading the organization provided by the Web's population. Similarly, Facebook is reading the relationships between people that are made in the course of using its software, information that it would not have otherwise.

What has made the web the rich environment that it is is that anyone can link anything to anything else. That linking is an expression; and even though we might not be able to characterize it in a few words (what does it mean that page A links to page B?) Google has shown that we can make use of those patterns of linking to help people find stuff on the Web that meets their needs.

In this regard, the Semantic Web has a serious problem. Much of the focus of Semantic Web work is the creation of vocabularies of defined relationships between things, with the intention that these relationships can be traversed and manipulated by algorithms. That is fine in itself, but the Semantic Web enthusiasts are primarily creating pre-determined, fixed relationships between things, mainly based on defining each thing in a class/sub-class relationship with other things. The Semantic Web vocabularies also define requirements called "range" which limit what you can link to from a specific element of your vocabulary. [1]

The tendency to pre-define vocabularies with strict rules is a carry-over from the Artificial Intelligence (AI) environment from which the Semantic Web arose. If you expect to work with machines, and only machines, then you have to define for them exactly what they can and cannot do, and you must present all decisions as formulas that can be calculated.

For reasons that is unclear to me, the Semantic Web work seems to be unaware of the great success that Google has had in using human information sharing activity to creating a meaningful web of links. There is little attention to the fact that establishment of relationships through linking can be done by humans, much less that the best and most useful linking will be done by humans. Human information linking is not definable a priori -- in fact, an a priori definition of allowed links essentially limits the future to re-running past concepts in new variations. It is an absolute barrier to creativity if you can only act on what has been pre-defined.

It's the difference between a prefab house, which can only become what it was designed to become (with some minor modifications that don't change its essence) and a box of blocks that can be used to create anything within the realm of physics (although superglue could extend the possibilities).

In part this is why I tend to speak of "linked data" rather than "Semantic Web." Linked data, as it has evolved as the description of metadata activity, carries less of the AI baggage than Semantic Web does.

To me what will be exciting about linked data is what people will do with it; what they will create, what they will experiment with, and both successes and failures. However, for people to do something with linked data they need tools -- tools that are as easy to use as those used to create Web pages and the links between them.

They also need a box of blocks to work with. These blocks need to be as free of predetermined rules as possible. [2] The terms need to be defined just enough to make them usable. This is also true of the relationships or links. Using our box of blocks metaphor, we want to be able to put the blocks in relationships like "above" "below" "near" each other. If the square blocks are defined as always "below" the rectangular blocks, then that limits what we can create.

The Semantic Web as a machine environment guided with AI formalities appeals to some because it promises to be neat and unambiguous. It will, however, foster only a very constrained amount of creativity, and will not be able to satisfy the full range of human curiosity.

It is a shame that many Semantic Web enthusiasts have little faith (or little interest) in the human potential that linking and openness can unleash. I, for one, am looking for partners in the development of a messy, intelligent, quirky, technology that can produce surprising results, created by people using linked data as a tool of expression. I am particularly exciting by the fact that we don't know what forms that expression will take.



[1] For example, you can define a creator has having to be type "Person" and that Person must be expressed as a URI, like http://kcoyle.net/kcoyle.rdf. In that case, you can't have a creator that is "Karen Coyle" because "Karen Coyle" is just a string of characters, not an identified entity. This means that if you don't have identifiers for your creators, you can't create data about what they created.]

[2] This is referred to as "minimum ontological commitement," introduced in Toward Princiciples for the Design of Ontologies Used for Knowledge Sharing, by T Gruber [http://www2.iath.virginia.edu/time/readings/ontology-semantics-metaphor/designing-ontologies.pdf]

Thursday, October 18, 2012

Is Linked Data the Answer?

I recently gave keynote talks at Dublin Core 2012 and Emtacl12 with the title "Think 'Different'." Since the slides of my talks don't generally have much text on them, I wrote up the talk as a document. The document has a kind of appendix covering the point in my presentation where I took advantage of my position on stage to ask and answer what I think is a common question: Is linked data the answer?

Many would expect me to answer "yes" to this question, but my answer is a bit more complex. Linked data is a technology that I believe we will make use of to connect library data to other information resources. That's what the "linked" in linked data is all about -- creating a web of information by connecting bits of data in different documents and datasets. However, we have to be very cautious about having "an answer." When you have an answer you tend to stop looking at the questions that arise, and you also tend to ignore questions that aren't going to be solved by the answer you have chosen. There is no technology that will do everything that we need, so while linked data can be useful for some things we may need to do, it cannot be the answer to all of our technical requirements.

Note that I describe linked data as "connecting bits of data." The origin of the semantic web is in the need and desire to make actionable data that today is essentially hidden within the text of documents. For example, if I say:

"My name is Karen. I will be holding a webinar on June 4 at 3:00 Pacific time for anyone who wants to learn about my paperweight collection."

That's text. There is interesting information in there, but it isn't available for any computational uses. The Semantic Web, as implemented through linked data, would make that information actionable. There are various ways to do this, and one is through the use of microformats which mark up data within a document. This could look something like:

<p>My name is <span class="name">Karen</span>. I will be holding a <span class="event">webinar</span> on <span class="datetime" title="2012-06-04T03:00-09:0000">June 4 at 3:00 Pacific time</span> for anyone who wants to learn about my <span topic="paperweights">paperweight</span> collection.</p>

This text now also has bits of data that can be used for various purposes, including linking. The linking capabilities in this particular example are low, but some additional information, like standard identifiers for the person and for the topic, would then increase the linkability of this data.

<p>My name is <span class="name" id="http://viaf.org/viaf/48369992/">Karen</span>. I will be holding a <span class="event">webinar</span> on <span class="datetime" title="2012-06-04T03:00-09:0000">June 4 at 3:00 Pacific time</span> for anyone who wants to learn about my <span topic="paperweights" id="http://id.loc.gov/authorities/subjects/sh85097666.html">paperweight </span> collection.</p>

This isn't a perfect example, but I wouldn't claim that we're heading toward perfect data. What we need is to get more out of the information we have. 

I perceive an assumption in the library linked data movement that what the Web needs (because linked data is data on the Web) is our bibliographic data. I disagree. The Web is awash in bibliographic data - from Amazon to Google Books, from fan sites like IMDB or MusicBrainz, and from sharing sites like LibraryThing and GoodReads. Libraries may have some unique bibliographic data, but most of what we have would duplicate what is already there, many times over.

There's also the fact that much of bibliographic data isn't DATA in the linked data sense. It isn't actionable data elements for the most part. In fact, bibliographic data is more like a structured document: it mainly has text, and that text is to be displayed to humans. It is possible to extract actual data (dates of publication, numbers of pages, various identifiers), but the text itself is a large part of the point about bibliographic data.

What this means for us in libraries is that we shouldn't be thinking that linked data will replace bibliographic data. It will encode the aspects of bibliographic data that will give us the most and the best links.

Then we need to ask: why are we linking? What will we get? Well, we can get connections between books and maps, between books and documents, and between search retrievals and libraries. This latter interests me especially. Google is experimenting with using microformat data, in particular the schema.org data that it is fostering along with Yahoo!, Bing, and Yandex (the Russian search engine).  Schema.org microformat data allows a search engine like Google to enrich the snippets with more than just a block of text from the page. This is an example from the Google Webmaster pages on Rich Snippets:
Below is my conceptualization of what we could do with library data. The bibliographic data, as I've said before, often already exists on the Web and we may not be helping things by adding many more duplicate copies of that data. But what we have in libraries that no one else has is library holdings data. We know where Web users can find "stuff" in their local community. If that could be linked to the Web, a future rich snippet might look like:

Obviously there are steps to be taken to make this possible, but if you want to think about how library data might fit into the Web of data that information seekers make use of millions or billions of times a day, this is one option. It's a start, and it uses data we already have.

You can take a look at the schema.org data that is created for WorldCat records simply by doing a WorldCat search and scrolling down to the section called "Linked data." The number of holdings is included (and this in itself is something that might interest Google as a measure of popularity). Making the link to the holdings of an actual library, and making that possible for all libraries, not just OCLC member libraries, is something I consider a worthy experiment for linking library data.

Friday, September 14, 2012

Rich snippets

At the recent Dublin Core annual meeting I heard Dan Brickley talk about Google's use of schema.org for rich snippets. Schema.org is commonly thought of as "search engine optimization" (SEO), which to most people means "how to get onto the first page of a Google results search." But the microdata in web sites can also be used to make the snippets shown more useful by incorporating more information from the web page. The examples above, from the Google rich snippets page, show features like ratings as well as links to actual content within the web page.

Now that WorldCat has schema.org markup, my first thought was: what kind of rich snippet would be good for library data? There is a rich snippet testing tool where you can plug in a URL and see 1) the snippet 2) what microdata is visible to Google. You can plug in a WorldCat permalink and see what the rich snippet result is:

http://www.worldcat.org/oclc/874206  (opens in separate window)

There is no rich snippet displayed here, which tells us that Google hasn't yet developed a rich snippet model for our kind of data. But you can see, in great detail, all of the coded data that is available. (The red warnings indicate that there is data in the OCLC microdata that isn't part of schema.org. OCLC is talking to the schema.org developers to incorporate new elements, some of which show up as warnings here.)

I began to think about how I would like this data used. It could be used to format a more bibliographic-like display, adding author, publisher, pagination. The ISBN could of course link to key online bookstores. (That would also bring in revenue for Google, so might be a popular choice for the search engine.) But what about libraries? How could rich snippets help libraries and library users?

The snippet could  lead back to WorldCat where the user could find a nearby library, but... wait! Google often knows your approximate location, and WorldCat knows whether libraries in your area have the book. AND the library catalog often has information about availability. I don't know how this data would interact with the WorldCat tool, but here's what I would like to see in the snippet:


This definitely goes beyond what "rich snippet" means today, but is not inconsistent with retrievals that pull data from multiple online sales outlets.  In the sales model, Google's assumption is that the searcher wants to obtain (in FRBR-speak) the item, and therefore various outlets that could provide that data are listed. This same logic could apply to libraries, of course. Libraries are a local source of many of the same things that are sold online, so the obtain logic fits.

This analysis of mine obviously ignores the economic incentive for Google to provide library holdings, especially since they would be seen as competing with sales.  I'm just dreaming here, doing the "what if" thing without the practical limitations.






Monday, July 23, 2012

Futures and Options

No, I'm not talking about the stock market, but about the options that we have for moving beyond the MARC format for library data. You undoubtedly know that the Library of Congress has its Bibliographic Framework Transition Initiative that will consider these options. In an ALA Webinar last week I proposed my own set of options -- undoubtedly not as well-studied as LC's will be, but I offer them as one person's ideas.

It helps to remember the three database scenarios of RDA. These show a progressive view of moving from the flat record format of MARC to a relational database. The three RDA scenarios (which should be read from the bottom up) are

  1. Relational database model -- In this model, data is stored as separate entities, presumably following the entities defined in FRBR. Each entity has a defined set of data elements and the bibliographic description is spread across these entities which are then linked together using FRBR-like relationships.
  2. Linked authority files -- The database has bibliographic records and has authority records, and there are machine-actionable links between them. These links should allow certain strings, like name headings, to be stored only once, and should reflect changes to the authority file in the related bibliographic records.
  3. Flat file model -- The database has bibliographic records and it has authority records, but there is no machine-actionable linking between the two. This is the design used by some library systems, but it is also a description of the situation that existed with the card catalog.

These move from #3, being the least desirable, to #1, being the intended format of RDA data. I imagine that the JSC may not precisely subscribe to these descriptions today because of course in the few years since the document was created the technology environment has changed, and linked data now appears to be the goal. The models are still interesting in the way that they show a progression.

I also have in mind something of a progression, or at least a set of three options that move from least to most desirable. To fully explain each of these in sufficient detail will require a significant document, and I will attempt to write up such an explanation for the Futurelib wiki site. Meanwhile, here are the options that I see, with their advantages and disadvantages. The order, in this case, is from what I see as least desirable (#3, in keeping with the RDA numbering) to most desirable (#1).

#3 Serialization of MARC in RDF

Advantages

  • mechanical - requires no change to the data
  • would be round-tripable, similar to MARCXML
  • requires no system changes, since it would just be an export format

Disadvantages

  • does not change the data at all - all of the data remains as text strings, which do not link
  • keeps library data in a library-only silo
  • library data will not link to any non-library sources, and even linking to library sources will be limited because of the profusion of text strings

#2 Extraction of linked data from MARC records

Advantages

  • does not require library major system changes because it extracts data from current MARC format
  • some things (e.g. "persons") can be given linkable identifiers that will link to other  Web resources
  • the linked data can be re-extracted as we learn more, so we don't have to get it right or complete the first time
  • does not change the work of catalogers

Disadvantages

  • probably not round-trippable with MARC
  • the linked data is entirely created by programs and algorithms, so it doesn't get any human quality control (think: union catalog de-duping algorithms)
  • capabilities of the extracted data are limited by what we have in records today, similar to the limitations of attempting to create RDA in MARC

#1 Linked data "all the way down", e.g. working in linked data natively

Advantages

  • gives us the greatest amount of interoperability with web resources and the most integration with the information in that space
  • allows us to implement the intent of RDA
  • allows us to create interesting relationships between resources and possibly serve users better

Disadvantages

  • requires new library systems
  • will probably instigate changes in cataloging practice
  • presumably entails significant costs, but we have little ability to develop a cost/benefit analysis

There is a lot behind these three options that isn't explained here, and I am also interested in hearing other options that you see. I don't think that our options are only three -- there could be many points between them -- but this is intended to be succinct.

To conclude, I don't see much, if any, value in my option #3; #2 is already being done by the British Library, OCLC, and the National Library of Spain; I have no idea how far in our future #1 is, nor even if we'll get there before the next major technology change. If we can't get there in practice, we should at least explore it in theory because I believe that only #1 will give us a taste of a truly new bibliographic data model.

Thursday, July 05, 2012

ISBN as URI

One thing that is annoying in this early stage of attempting to create linked data is that we lack URI forms for some key elements of our data. One of those is the ISBN. At the linked data conference in Florence, I asked a representative of the international ISBN agency about creating an "official" URI form for ISBNs and was told that it already exists: ISBN-A, which uses the DOI namespace.

The format of the ISBN URI is:
  • Handle System DOI name prefix = "10."
  • ISBN (GS1) Bookland prefix = "978." or "979."
  • ISBN registration group element and publisher prefix = variable length numeric string of 2 to 8 digits
  • Prefix/suffix divider = "/"
  • ISBN Title enumerator and checkdigit = maximum 6 digit title enumerator and 1 digit check digit
I was thrilled to find this, but as I look at it more closely I wonder how easy it will be to divide the ISBN at the right point between the publisher prefix and the title enumerator. In the case where an ISBN is recorded as a single string (which is true in many instances in library data):

9781400096237
9788804598770

there is nothing in the string to indicate where the divider  should be, which in these two cases is:

978.14000/96237
978.8804/598770

I have two questions for anyone out there who wants to think about this:

1) Is this division into prefix/suffix practical for our purposes?
2) Would having a standard ISBN-A URI format relieve us of the need to have a separate property (such as bibo:ISBN) for these identifiers? In other words, is the URI format identification enough that this could be used directly in a situation like:

<bookURI> <RDVocab:hasManifestationIdentifier><http://10.978.14000/96237>


Saturday, June 23, 2012

Europe and Library Linked Data

I have just returned from a conference on library linked data in Florence, Global Interoperability and Linked Data in Libraries.  The talks were fascinating and there is only a small amount that I can convey in a blog post, but I thought that a good way to start would be to highlight some differences I see in the European approach.

Cultural Heritage

The cultural timeline in Europe is on an entirely different scale from what we are used to in the US. Library of Congress's "Historical Newspapers" collection covers 1836-1922. At the museum of the synagog of Florence, the docent referred to an event in 1571 as "the first in modern times." We have history, but Europe has History. This means that there is a great emphasis on archives, manuscripts, and museums in all work done by cultural heritage organizations. At no time during the conference was there discussion of "STM" materials (that's Science, Technology, Medicine) other than a talk on Renaissance science, or scholarly communication, both of which are often on programs in the US. (See talks on linked data and the Vatican library and Linked Heritage.) Note also that the shared European culture database, Europeana, uses linked data, which encourages all contributors to also move in that direction.

National Libraries

Closely connected to the view of cultural heritage is the role of the national library. In many countries, if not most, the national library's primary role is to conserve the written heritage of the nation. This fact could be used for better data sharing; for example, each national library could have the responsibility for its subset of name authorities, and the name file could be 'in the cloud.' (See talk by Malmsten.) Ditto for the cataloging of modern publications.

Government Data

Led by various European Union initiatives there is currently a strong movement to make government data available. (Note: there is also an open government movement in the US, but it has less emphasis, as I see it, on the sharing aspects of releasing the data.) Government data from the member countries is needed to make possible the analyses needed for Europe-wide programs. Since this data must be shared and linked, providing it as linked data makes perfect sense.(See: datacatalogs.org, a catalog of government data catalogs, talks by Morando, Moriando, Menduni. )

Rights and licenses

As countries decide to make their government data open, they must decide about rights, in particular data ownership. The European Union has recommendations for data rights, and individual countries are developing licenses that will be used for their released data. Because libraries are often government-funded institutions, these licenses will also apply to library data. This is not always a good thing. Some countries have declared government data on an open license, but others, like UK, Italy, and France, are using a license similar to CC-BY. The reasons for this have to do with the need to maintain provenance of government data since that data often has an official role in decision-making. (I wonder if the addition of provenance to linked data will help, and that makes me wonder about provenance "spoofing" and how much of a problem that will be.) (See talk by Morando.)

International Standards

People talk about how insular countries like China and North Korea are, but I become deeply aware of how insular we are in the US when I attend meetings outside of this country. We are beginning here (finally) to pay more attention to the global network of libraries, but Europe follows IFLA standards "religiously," as well as EU standards for data sharing. The Web of Data as a global resource makes even more sense in the European context. This makes me wish I knew more about the remainder of the globe. 

My talk at the conference is here: English, italiano.


Wednesday, May 30, 2012

FRBR, FRAD, ISBD in LD by BNE

To translate the title of this post: the National Library of Spain (BNE) has developed a linked data set for its bibliographic data that makes use of FRBRer, FRAD and ISBD in RDF. There are a number of interesting aspects to their decisions, and we can now compare this solution to the solution used by the British Library for their National Bibliography (diagram in PDF).

There is an English-language article with a short explanation of the project and a high-level design of the data structure. The diagram there didn't print out terribly well, so you might want to consult the Spanish-language version from their ontology documentation page:
What I find particularly interesting about this interpretation of FRBR is that the Work ("obra") and Expression ("expresión") are considered authority entities, while the manifestation, in the darker area below, is a bibliographic entity. (If this is common knowledge, I must admit that I hadn't seen this expressed so clearly before.) The Work and Expression are coded using RDF properties (e.g. "data elements") from FRBRer, FRAD and FRSAD; the Manifestation uses RDF properties from ISBD. These are the same element sets that Pat Riva recently announced on various lists.

There is a fascinating visualization of the entities and links for de Cervantes and his works.  Note that you can zoom in and see more detail in the many nodes there. The visualization project is "in progress" and at the moment you cannot move directly from the nodes on the screen to the actual data (except by plugging in the URIs that you see). The resulting displays are similar to those of DBPedia, which means that you basically have to know what you are looking at in order to follow from one entity to another.

To view some actual entity records, Daniel Vila of the OEG-UPM (Ontology Engineering Group) provided these examples that are linked together, a Person, a Work, and an Expression:

- Miguel de Cervantes: http://datos.bne.es/resource/XX1718747

- Don Quijote work: http://datos.bne.es/resource/XX3383563

- Don Quijote de la Mancha Spanish (original) Expression: http://datos.bne.es/resource/XX3383563spa

The Expression leads to many dozens of manifestations, but here is one example:
Obviously this isn't easy to read since the display does not yet substitute the names for the ID numbers of the ISBD and FRBR entities (e.g. isbd:P1004 is "has title proper").

If you look at the Work and Expression records you will see that they consist primarily of links -- the Work links to creators, subjects, and Expressions. The Work also provides some alternate display text and links to some outside sources:


The Expression relates to the Work and the many manifestations. The only thing we would think of as a data element in the Expression in this example is a coded version of the language:

I find it particularly interesting to see this example of library data in linked data form because the only thing new here is the use of linking -- the data appears to me to be basically vanilla ISBD, re-organized using FRBR entities. One could conclude that this means that we could move into linked data without taking on the RDA step. I would very much like to see how these examples would be different if RDA were the applied cataloging rule set.

The view of Works, Expressions, plus of course Persons, Corporate bodies, and Subjects, as akin to today's authority records really serves to separate the linking entities from the description (which is in the Manifestation). The manifestation remains primarily text strings, but the examples in the BNE demo suggest that there are significant opportunities to create links between library resources and potentially to non-library resources without major modifications to descriptive cataloging. In fact, it looks like our attention should be on our authority data, which provides the linking opportunities.

More and more I come to the conclusion that in the linked data space the thing we seem to focus on today, descriptive cataloging, will be less useful than the entities that are represented by our authority data. For this reason I see VIAF as a very good start because links to and from those personal and corporate entities should give us a lot of connections to the rest of the world.

Monday, May 21, 2012

Google goes semantic

In a long-awaited move [1], Google has announced that its search will now be "semantic." They don't actually mean "semantic" in the sense of the semantic web, although there are similarities. While what Google is doing may not formally follow the W3C standards for the semantic web, there is no doubt that they are performing acts of "data linking" that make use of the concepts of linked data. The W3C standards for linked data are designed for openness, so that data from disparate communities can come together. Google has no obligation to play well with others and, as we saw with the development of schema.org, is in a position to make its own rules, many of which are known only within the giant Google-verse. They call their technology a "knowledge graph" and talk about "things not strings." I've used this same phrase myself in numerous presentations on linked data.

Google has always been about using links between things on the web to determine its brand of "relevance" of a web resource to a search query. By using existing linked data, via large stores of links like Dbpedia, Wikipedia, Freebase, and presumably others, Google can now expand its offerings from a single list of results to additional information about the topic that might be the intended topic of the searcher. I say "might be" without any irony; whether in a web search engine or a library catalog, the communication between the searcher's mind and the device that provides results is always only approximate. What the additional data provides is not only more context but a more ample explanation of the topics that have been retrieved. No longer do users have to guess from snippets the meaning of the results in the result set, but they can see a Wikipedia-like entry that not only gives them more information, but it contains links to other sources of information of the topic.
Snippet
"Knowledge Graph" result

"Knowledge graph" detail



At a meeting of the Northern California Technical Services Group in Berkeley last Friday, I said to the group:

Imagine that you have an 18-year-old user who finds a novel on your library's shelf by Oliver North. The user looks up the author in  your catalog and sees that this person has written a few other books, but oddly always with a "co-author." Is someone so inept worth reading? Now imagine that your catalog also presents the user with the context: Ollie North, Iran Contra, and related persons. Suddenly the user sees where North fits into US history, has a chance to find out what an interesting character he is, and the books take on a whole new meaning.

That was before I saw this Google result.

We treat library users as if they are all-knowing; as if they know each author in our catalog, as if the title of the book and the number of pages is sufficient for them to decide if it is a good read or has the information they need. This is so obviously false that I am at a loss to explain how we continue to work under this illusion.

[1] Google purchased the only linked-dated search system, Freebase, in July of 2010, thus tipping their hand that they were moving in that direction. Not only did they acquire Freebase and the skills of its employees, they eliminated a potential rival (although it may be silly to consider that anyone could really be a rival to Google).