Friday, October 11, 2013

Who uses Dublin Core - dcterms?

In my previous post I gave some data on Dublin Core field use. Today I look at who is using Dublin Core's dcterms vocabulary.

The LOV statistics show 212 datasets that use the vocabulary at http://purl.org/dc/terms/, and the number of instances of usage. I did some "back of the envelope" counts on what types of organizations or projects use the terms, and also the type of use. By these calculations, the highest use was from libraries (> 60%). The next highest use was in a single language study called Semantic Quran (~10%). Third was the use in government data at less than 1%.

If one looks at the type of data, bibliographic data makes up nearly 90% of the usage. In this category I included archives, eprint repositories, and a few databases of videos and teaching materials.

From this one might conclude that dcterms isn't used much outside of the bibliographic world, but in fact traditional libraries provide only 28 of the 212 datasets on this list. The range of users and uses is impressive. Here are a few to peak your interest:
  • Southampton University has a number of datasets of civic information, including a list of bus stops.
  • There is a biomedical data service called eagle-i used by 24 universities or departments that provides information on specimens, reagents and services. This contributes nearly 500,000 instance of dcterms usage.
  • The New York Times linked data service uses dcterms. This service consists of topics (persons, organizations, locations, topics) covered by the newspaper.
  • I've mentioned the Semantic Quran. This is a linguistic database consisting of 43 translations of the Quran. It contributes over 6 million instances of dcterms use.
  • There is government data covering a wide range of topic areas. By my estimate there are at least 70 sets of government data in this compilation (including international), with everything from the aforementioned bus stops to election data, patents, economic indicators and scientific information.
If one is to make conclusions from this evidence, it could be said that the dcterms vocabulary is a core vocabulary for the description of intellectual resources, such as the holdings of libraries and archives, but that it also provides functionality for a wide range of data types. 

There are also users of the original Dublin Core vocabulary, now referred to as "1.1". I will cover that usage next.

Wednesday, October 09, 2013

Dublin Core usage in LOD

Thanks to some projects that gather statistics on the growth of linked data, we can find out various interesting things about the vocabularies being used and the degree of linking between data sets from different communities. The data I report here comes from LODstats via the Linked Open Vocabularies (LOV) project.

The LOV project looks particularly at the interrelations between vocabularies. For example, it can show which vocabularies use terms from other vocabularies. This crossover of terms is one of the things that makes links between datasets possible. For example, this shows that the geoSpecies vocabulary is not itself referenced by other vocabularies, but can link through its use of vocabularies like FOAF and Dublin Core. You can watch the visualization grow here.
In contrast, this is what Dublin Core terms looks like at LOV:

With the animated visualization here.

Dublin Core does seem to have fulfilled its role as a core vocabulary that many different communities have found useful, at least in part. The set of terms often abbreviated as "dcterms" (or sometimes "dct") and whose namespace is http://purl.org/dc/terms/ has been used approximately 192 million times as reported in the LOD statistics. This is only the usage in the 2289 linked data datasets used by that project. The earlier set of Dublin Core terms, the original fifteen terms, whose namespace is http://purl.org/dc/elements/1.1/, has been used 24.2 million times. This gives us a total of 216 million uses of Dublin Core in this particular count.

The interesting question, then, is what parts of DC are heavily used? I have a sorted list, from most to least, of all terms in the http://purl.org/dc/ namespace. The top fifteen terms are all from the "dcterms" namespace:

count          term
24147876    subject
22575133    identifier
17120343    title
17065873    issued
14459601    publisher
11605978    language
9930733    medium
9795117    format
9792064    BibliographicResource
7700745    isPartOf
7371553    creator
7241777    contributor
6590791    description
6184994    type
5983236    extent

Of this list, only four were not part of the original "Dublin Core 15" vocabulary: issued, medium, BibliographicResource, and isPartOf. The terms of that original vocabulary cluster together beginning right after the last term in the above list. I believe this provides an interesting affirmation that the original fifteen terms were a fair definition of "core." 

However, these terms, in the "dcterms" namespace got less than ten uses, and some were even zero:

accrualPeriodicity
Frequency
AgentClass
dateSubmitted
isRequiredBy
Jurisdiction
LicenseDocument
LinguisticSystem
MediaType
MediaTypeOrExtent
PeriodOfTime
PhysicalResource
RightsStatement

The last term, which got zero in the LOD calculations, is particularly interesting because the element "rights" in the original "DC 15" got 398,361 uses, and is ranked 39th in the list of elements the overall http://purl.org/dc namespace.

Next, I'll take a quick look at which datasets are contributing to the use of Dublin Core terms, and who is creating those datasets.



Tuesday, October 08, 2013

Women in Science

Today's New York Times has an excellent article on women in science -- that is, of course, the lack of -- entitled Why are there still so few women in science? Coincidentally, this image hit the pages of Google+ in recent days even though it has been around since at least 2010:


The picture has set off a huge argument about religion and atheism on the post that carried it. There has been some less rancorous discussion of who really should or should not be in the picture. One woman posted that there are other women who should be there, but no one seems to have noticed the portrayal of the one women who is there, Marie Curie.

Of course it is absolutely right that Marie Curie be portrayed - as one of the few people who have earned more than one Nobel Prize, and the only person to have been awarded Nobels in two different scientific areas, she is obviously qualified. However, the problem is the portrait, which is not of Marie Curie, but Marie Curie as portrayed by Susan Marie Frontczak in "Manya," a one-woman drama on the life of Marie Curie. This is as if the portrait of Einstein had been actually that of the actor who played him in the episodes of Alien Nation. There are plenty of photographs of Marie Curie, from her early days to her later years.

I think it is only fitting that Curie have her own identity, and not be given the image of an actress portraying her.  So this is just a heads up for the Marie Curie fans among us: you can find plenty of photos with an image search, although you will have to color them yourself. That's hopefully not to much to ask as a way to honor such a significant scientist.

Tuesday, October 01, 2013

Cataloging as Observation*

"Last, there has been some spectacularly misguided and misinformed discussion of the need to create 'master records' for works that are manifested in different physical forms. It is hard for me to believe that this notion has been put about by people who are cataloguers. Let me spell it out. Descriptions are of physical objects (and, nowadays, of defined assemblages of electronic data). It is literally impossible to have a single description of two or more different physical objects…"

Michael Gorman, AACR3? Not! in: Schottlaender, Brian. The Future of the Descriptive Cataloging Rules: Papers from the Alcts Preconference, Aacr2000, American Library Association Annual Conference, Chicago, June 22, 1995. Chicago: American Library Association, 1998. p. 27
When I first read this aside in Michael Gorman's highly charged article in opposition to the cataloging rules that would succeed the AACR2 rules (that he edited), I was shocked that anyone would say that cataloging is primarily a "description of physical objects." I thought of library catalogs as being about content, about knowledge. But as Gorman surely has a finely honed grasp of the purposes behind library cataloging, it seemed best not to dismiss such a statement, and I marked it in my copy of the book and tucked it away in my memory.

It has come back to me as I've pondered not only RDA and the state of library catalogs, but in my attempts to explain library cataloging to non-librarians. At the meeting regarding the question of bibliographic metadata and copyright, I described library cataloging as analogous to a medical diagnosis: a great deal of testing, expert knowledge, and judgment result in a few scribbled lines in a medical file and a prescription. If you consider these latter two the "metadata" of the situation, you see that what is visible is merely the tip of the iceberg, with a great deal of intellectual activity hidden below the surface. What I didn't mention at that time, because it wasn't yet clear to me, was the role of observation in the two activities being compared. A good physician knows how to observe and analyze a patient, and a good cataloger knows how to observe and analyze a cultural artifact.

Cataloging rules actually instruct their users on how to observe. In fact, the very first rule in AACR2 (1.0.A) defines the sources of information for the catalog entry: the preferred source of information is always the thing being described. In essence the thing itself is the primary informant for the catalog record.

There are a couple of important things we can conclude from this. The first is that the act of cataloging is an act of describing what is being observed. This makes cataloging something like the act of a biologist who is describing a specimen before her. In theory, if both librarians and biologists follow the rules of their disciplines, the same specimen or artifact would be described similarly by two different professionals. (In fact, there are always edge cases that defy simple application of the rules, but these are also the cases that make the professional activity interesting.)

The next important aspect about library cataloging is that the content of the catalog record is in large part the expression of those who created the artifact itself, not that of the cataloger. As RDA (chapter 2) says:
"The elements reflect the information typically used by the producers of resources to identify their products—title, statement of responsibility, edition statement, etc. "
Significant parts of the cataloging description are either quotes from the thing or paraphrases of observable content. I am unaware that anyone has ever challenged the right of catalogers to copy this information from the artifact to the catalog record.

What I have addressed to this point follows a fairly strict definition of "descriptive cataloging" and presumably not terribly far from the division between description and access that is made by RDA, or from the division of AACR2 into description and headings. The access or heading portion of the catalog record adds information beyond the observations of the physical piece. Headings and access points are standardized forms of proper names (including persons, corporations, government bodies, and some titles). The standardized form of the name serves as an identifier for the named entity, and also normalizes display. Here are a few examples:
On the artifact: J. R. R. Tolkien
Heading: Tolkien, J. R. R. (John Ronald Reuel), 1892-1973
On the artifact: Beethoven's Ninth Symphony
Heading: Symphonies, no. 9, op. 125, D minor
On the artifact: T. C. Boyle
Heading: Boyle, T. Coraghessan.
Note that some of these add more information than is on the actual piece, and that additional information requires research. That doesn't mean necessarily that the additional information requires the level of creativity that qualifies it for copyright protection. This is one of the areas of bibliographic metadata that needs to be analyzed further. However, I think we can conclude that some portion of "descriptive cataloging" consists primarily of observations about real world objects; and some portion normalizes those observations to create standard identifiers for bibliographic entities that exist entirely independently of the cataloging act.

You have undoubtedly noticed that I have not mentioned subject headings or classification in this post. Subject analysis, although recorded on the same catalog entries as the bibliographic data, is a separate activity in the cataloging workflow, and is not covered by the above-mentioned cataloging rules.  As a topic in the "metadata and copyright" discussion it should be covered separately.



* My thanks to Tom Baker who, as I struggled to find another way to say "bibliographic description" suggested that catalogers make observations about things.

Tuesday, September 24, 2013

Hopes and fears for Google Books case

We're back in the saddle of the now epic lawsuit against Google for its massive scanning of the books held by libraries. I have very mixed feelings about the case and its outcomes, and the news reports from yesterday's hearing (transcript) in Judge Denny Chin's court are not making me feel any better about it. In brief, the Author's Guild is claiming that Google violated fair use by scanning in-copyright books. Since that act alone is not sufficient to address a defense of fair use, they also state (correctly, in my view) that although Google is not providing advertising on the individual book pages, that it overall makes money off of the scanned books because that digital corpus enforces its position against other search engines. There are some things that the Authors Guild has right (such as, that Google makes money off of search results pages that can include links to Google Books), but they miss the mark in other arguments:
"For all intents and purposes, it paid libraries for the right to digitize and copy much of our nation’s literary heritage and then used the resulting digital library to gain a competitive advantage over search engine competitors that respected the rights of authors by limiting their digitization programs to books that were either licensed or were no longer protected by copyright. Aided by its infringing conduct, Google’s search engine has proven remarkably successful—to the point where “google” has become a widely used verb in the English language.
First, the addition of Google Books to the search took place long after we were all "googling." Google's main value still comes from providing access to open web resources that otherwise would just be a massive digital junk heap. I suspect that those who are interested in using Google to search within the text of "closed" books (ones that are not available as full text online) consciously go to the Google Book Search pages. I don't know this for a fact, but I'd be willing to bet that user intent behind most Google searches is to access the actual content of a web page or document, not to be given a reference to an off-line resource.

Next is the statement that Google "paid libraries for the right to digitize..." This makes it sound like Google gave the libraries money, and that there was no cost to the libraries. The agreement between Google and libraries was an exchange that had costs for both (less for the libraries, more for Google) and benefits for both (less for the libraries, more for Google). In the end, Google got the better part of the deal, but libraries got something, even though something they have not yet been able to greatly benefit from: libraries got copies of the scans at a lower price than had they done the digitization themselves. Unfortunately, due to both copyright issues and the nature of the agreement between Google and the libraries, there are significant barriers to making the kind of uses that would make this a truly transformative corpus for research.

All of the news reports emphasized some comments by Judge Chin to the effect that Google Books appears to be both transformative (in the copyright law sense) and a benefit to society. What worries me a bit is that Judge Chin is not looking beyond the use of the resulting digital texts for search. I consider search to be the tip of the iceberg, and the visible part of Google Books that Google would like everyone to focus on. My assumption is that Google has a research interest in having exclusive access to 20 million non-Web digital texts in a myriad of languages, and that this research is aimed not only at search but at Google's desire to be THE interface between man and machine, which means that machines have to get better at human languages.

If Judge Chin rules that Google's book digitization is fair use, it's a huge win, not only for Google but also for libraries. After all, if it is fair use for Google to digitize works for the purposes of searching, there is no question that it is also fair use for libraries to do the same. If Judge Chin rules that Google's book digitization is NOT fair use because of profit-making, then we still do not know for sure whether library digitization would be considered fair use (although much would depend on exactly how the decision is worded). This of course makes me want to cheer on Chin toward the "is fair use" decision, but at the same time I know that this means that any research that Google is doing on its private cache of digital texts will continue, giving them great advantages over competitors in the arms race of technology advancement.

Once again, I so wish that large-scale digitization for search and research had been undertaken by libraries, not Google. The questions of "not for profit" and social value would be a slam-dunk, and I'd not be harboring this fear that there is a hidden agenda behind the project. Maybe if libraries had done this we'd only have two or three million digitized books, not 20 million (as is claimed for Google), but they'd be untainted, in my mind, and I could still consider them a cultural heritage resource rather than a commercial product.

Sunday, September 22, 2013

Copyright, Metadata, and Attribution

The Berkeley Center for Law and Technology (BCLT) has done some interesting research on copyright, including a white paper that details the issues of performing "due diligence" in a determination of orphan works.

Recently I attended a small meeting co-sponsored by BCLT and the DPLA to begin a discussion of the issues around copyright in metadata, with a particular emphasis on bibliographic metadata. Much of the motivation for this is the uncertainty in the library and archival community about whether they can freely share their metadata. As long as this question remains un-answered, there are barriers to the free flow of data from and between cultural heritage institutions.

At the conclusion of the meeting it was clear that it will take some research to fully define the problem space. Fortunately for all of us, BCLT may be able to devote resources to undertake such a study, similar to what they have done around orphan works.

One of the first questions to undertake is whether bibliographic metadata is copyrightable in the first place. If not, then no further steps need to be taken -- not even putting a CC0 license on the data. In fact, some knowledgeable folks worry that using CC0 implies that there do exist intellectual property rights that must be addressed.

However, before you can attempt to determine if bibliographic metadata can be argued to be a set of facts which, under US copyright law, do not enjoy protection, you must be able to define "bibliographic metadata." During the meeting we did not attempt to create such a definition, but discussion ranged from "anything about a resource" to a specific set of descriptive elements. As there were representatives of archives in the room, we also talked about some of the implications of describing unpublished materials, which have a different legal standing but also provide less self-identification than resources that have been published. Drawing the line between fact and embellishment in bibliographic metadata is not going to be easy. Nor will the determination of level of creativity of the data, a necessary part of the analysis for US law. Note that other types of metadata were also discussed, such as rights metadata and preservation metadata, as well as a recognition that the exchange of metadata will of course cross national boundaries. Any study will have to determine where it will draw the "metadata" line, and also whether one can address the the question with an international scope.

Another complexity is that bibliographic data is already "crowd-sourced" in a sense. For any given bibliographic record,  different contributions have been made by different librarians from different institutions and at different times. This recognition makes it hard to ascribe intellectual ownership to any one party. And while library catalog data may be considered to be factual, it is much more than a simple rendering of facts, as the complexity of the cataloging rules attests. I likened library cataloging to a medical diagnosis: the end result (some scribbles in a file and perhaps a prescription given to the patient) does not reveal all of the knowledge and judgment that went into the decision. Metadata is the tip of an iceberg. That may not change its legal status, but I think that unless you have delved into the intricacies of cataloging it is hard to appreciate all that goes into the fairly simple display that users see on the screen.

The legal question is difficult, and to me it isn't entirely clear that solving the question on the legality of bibliographic data exchange will be sufficient to break the logjam. In a sense, projects like DPLA and Europeana, both of which have declared their metadata to be available with a CC0 license, might have more real impact than a determination based in law. Significant discussion at the meeting was about the need for attribution on the part of cultural heritage institutions. Like academics, the reputation and standing of such institutions depends on their getting recognition for their work. Releasing metadata (including thumbnails in the case of visual materials) needs to increase the visibility of those institutions, and to raise public awareness of the value of their collections. It is possible that solving the attribution problem could essentially dissolve the barriers to metadata sharing, since the gain to the institutions would be obvious.

Perhaps my one unique contribution to the group discussion was this:

We all know the © symbol and what it means. What we need now is an equally concise and recognizable symbol for attribution. Something like "(@)The Bancroft Library" or "(@)Dr. Seuss Collection". This would shorten attribution statements but also make them quickly recognizable, and a statement could also be a link to the appropriate web page. Standardizing attribution in this way should make adding attributions easier, and would demonstrate a culture of "giving credit where credit is due." The symbol needs to be simple, and should be easy to understand. It's time to comb through the Unicode charts for just the right one. Any suggestions?

See Also:


Unicode 1F6A9 - Triangular flag meaning "location"

Friday, August 09, 2013

Green paper on copyright, II


"To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries" (US Constitution)
 In 1993 the US government issued a green paper on copyright and the Internet. At the time the latter was being referred to as the "National Information Infrastructure" or NII. That green paper led to a white paper called: Intellectual Property and the National Information Infrastructure. The conclusions in this paper led to government efforts like the DMCA, as well as the as-yet unresolved questions about the library exceptions in section 108 of the copyright law.

Twenty years later we have another green paper on copyright and the Internet. It is, as are nearly all government documents, long and complex, and I hope to find time to do a comparison of the two papers to see if we have progressed in this area. But from a first reading I believe I can say that there is something that the two papers have in common: who they consider to be a creator and a rights holder.

It won't surprise you to learn that the emphasis in papers I and II is on the commercial content production communities: books, movies, music, newspapers.   As quoted in the press release,
“We see a digital future in which the relationship among digital technology, the Internet, and creative industries becomes increasingly symbiotic,” said Assistant Secretary of Commerce for Communications and Information and NTIA Administrator Lawrence E. Strickling. “In this digital future, the rights of creators and copyright owners are appropriately protected; creative industries continue to make their substantial contributions to the nation’s economic competitiveness; online service providers continue to expand the variety and quality of their offerings; technological innovation continues to thrive; and consumers have access to the broadest possible range of creative content.”
Oddly, he seems to be describing the Internet that I interact with today, but calls it the digital future. Nowhere, however, is there a mention of the rights of creators of Facebook pages, Google+ entries, Youtube videos, or tweets. A distinction is often made between creators and consumers, yet through social media and Internet publication methods (like blogging), that division is less valid than it was in the past.

Nor does the report admit that there is a the vast amount of content that is user-supplied and highly consumed, like Youtube, and that this content exceeds, both in publication and in consumption, the commercial offerings that the authors of the report are so concerned about. Hearings held to discuss the first green paper had industry representatives saying that the Internet would not be a "success" if commercial and entertainment content was not available, as if the Internet at the time (1994) were just an empty shell waiting for Time-Warner and Disney and Thomson-Reuters to come along and fill it up. Bruce Lehman, chairman of the committee producing the green and white papers, said to Congress in 1995:
 Creators, publishers and distributors of works will be wary of the electronic marketplace unless the law provides them the tools to protect their property against unauthorized use. Thus, the full potential of the NII will not be realized if the education, information and entertainment products protected by intellectual property laws are not protected effectively when disseminated via the NII.
In case you weren't there at the time, the Internet in 1994 was a thriving community with huge amounts of content. It was less flashy than today's Internet, since the technology did not yet allow for the efficient streaming of video and sound, and downloading a photograph could take a while. But it was not languishing for lack of content. This goes many-fold for the Internet today, but the representatives of the commercial content industries are somehow oblivious to any content that isn't making money for them. And that includes every Youtube upload, every Facebook page, bazillions of Flickr photographs, countless tweets.

Here are some stats from Youtube and Hulu, taken from their sites. Admittedly, Youtube is not 100% non-commercial and Hulu is not 100% pay-for-view, but this still shows that Youtube and its crowd-sourced content should be counted as real content in the sense of the green paper on copyright. But it isn't.

Hulu
 Number of Hulu video views in the past year     457 million
Total number of people who watched Hulu at least once in the past year     38 million
Number of Hulu video views in the past year     457 million
Average number of videos a Hulu watcher views     12
Average length of time a person spends on Hulu     1hr 13min
Number of devices in use that Hulu is available on     120 million
Percentage of people who use Hulu and only watch television shows     73 %
Percent of people who use Hulu to watch movies     9 %
Percent of videos viewed online that Hulu makes up     4 %
Percent growth for Hulu from 2010 to 2011     60 %
Total revenue made in 2011 by Hulu     $420 million
Youtube
More than 1 billion unique users visit YouTube each month
Over 6 billion hours of video are watched each month on YouTube—that's almost an hour for every person on Earth, and 50% more than last year
100 hours of video are uploaded to YouTube every minute
70% of YouTube traffic comes from outside the US
YouTube is localized in 56 countries and across 61 languages
According to Nielsen, YouTube reaches more US adults ages 18-34 than any cable network
So Youtube got 12 billion unique visitors in a year, to Hulu's 457 million. You'd think data like that would have an impact on the thinking of these folks. And the fact that Youtube reaches more viewers in the key demographic than cable TV should also shake up some media executives. In addition, those who upload content to Youtube are rights holders, as the Youtube terms of use makes clear:
For clarity, you retain all of your ownership rights in your Content. However, by submitting Content to YouTube, you hereby grant YouTube a worldwide, non-exclusive, royalty-free, sublicenseable and transferable license to use, reproduce, distribute, prepare derivative works of, display, and perform the Content in connection with the Service
You would also think that the report would address these Youtube content providers as creators, just as those who add content to Facebook pages are creators. But these creators are somehow not real creators in the minds of the writers of the report, and their needs are not addressed. No where does the report decry the exploitation, without compensation, of the work of these creators by industry giants like Facebook and Google. If someone else is making money off of your content it's bad, unless you aren't one of us, and therefore it's ok. The 99.99% of us who do not own the means of production and distribution must cede our rights in order to participate in the creation of content.

The report urges Congress to criminalize the unauthorized streaming of copyrighted works. Copyright law violations are subject to civil penalties, so this greatly ups the threat level for violations. This also revives one of the disputed and hated aspects of the famed "Stop Online Piracy Act," (SOPA) which was defeated through massive activism.

The upshot is that copyright will benefit those who create industrially, but not the millions who create individually. We should just admit that copyright law no longer has anything to do with creativity or social impact, and instead rename it to the "copyright industries law." As the 2013 green paper explains:
The industries that rely on copyright law are today an integral part of our economy, accounting for 5.1 million U.S. jobs in 2010—a figure that has grown dramatically over the past two decades. In that same year, these industries contributed 4.4 percent of U.S. GDP, or approximately $641 billion.
That's what this green paper, and the previous green paper, and all of the changes to copyright in the centuries since the time of the US Constitution, are really saying. Very little of what is protected today, and very little of what contributes to that $641 billion, is either Science or a useful Art. The founding fathers were writing at a time when science and technology defined progress, and progress defined prosperity. They did not, and could not have, anticipated the rise of leisure time and the media that would allow the creation of a huge economic center based on entertainment (including infotainment, which covers much news reporting today). Ironically, real Science is being encouraged to provide its content as Open Access. It's time to openly state that copyright law today is not what the founding fathers had in mind.

Update: Excellent blog post on the Green Paper and copyright by Kevin Smith.


Addendum: I just re-discovered the CPSR statement on the NII from 1994, and it contains this paragraph which is so sadly insightful:
An imaginative view of the risks of an NII designed without sufficient attention to public-interest needs can be found in the modern genre of dystopian fiction known as "cyberpunk." Cyberpunk novelists depict a world in which a handful of multinational corporations have seized control, not only of the physical world, but of the virtual world of cyberspace. The middle-class in these stories is sedated by a constant stream of mass-market entertainment that distracts them from the drudgery and powerlessness of their lives. It doesn't take a novelist's imagination to recognize the rapid concentration of power and the potential danger in the merging of major corporations in the computer, cable, television, publishing, radio, consumer electronics, film, and other industries. We would be distressed to see an NII shaped solely by the commercial needs of the entertainment, finance, home shopping, and advertising industries.