Monday, January 28, 2013

Wikipedia and bibliography

One thing that I have noticed on many Wikipedia pages is that you get references, and some external links, but not what I would call a useful bibliography. There are bibliography pages, but these are usually huge, comprehensive lists of works on broad topics.

What I appreciate about Wikipedia is that it is a great place to start when you are delving into a new topic. The links between Wikipedia pages within the text can be very helpful (albeit at times a bit too much of a distraction when your curiosity takes you a great distance from where you started). I would like for at least some Wikipedia pages to serve also as a beginning bibliography for the topic.

What constitutes a beginning bibliography is obviously not easy to define, but Wikipedia has never shied away from such difficulties. When I was in college we had a separate undergraduate library that had the basic books in each field. If you'd never thought about, say, anthropology, you could find the LC class number for that topic, go to the shelf, and you'd be looking at the books that most professors teaching an undergraduate course would consider "must reads" in the field. There was even a published list, called "Books for College Libraries" that listed the key books that college libraries of various sizes should have in their collections. This is now an online resource called "Resources for College Libraries" (behind a paywall) that has over 70,000 titles in 61 different subject areas. What this means is that doing something similar in Wikipedia is neither impossible nor radical.

I got to do a small experiment in this area yesterday at a local "Wikipedia edit-a-thon." I had brought with me some books that I thought would yield interesting explorations - one of which is a marvelous book called "Woman in Science" published in 1913.  Although the author is listed as "H. J. Mozans" that turns out to be a pseudonym for John Augustine Zahm. Zahm does have a Wikipedia page, and it did list, within a paragraph, a number of books that he had written. Oddly enough, Woman in Science wasn't one of them. I added it, then decided that since he had written a handful of books that I would add a bibliography on his page. The Wikipedia bibliography format is, well, you know, like so many Wikipedia structures, something less than friendly. But I discovered something that I probably should have known.

I went to the Open Library page for the book, and near the bottom found the list of export formats.


Clicking on "Wikipedia Citation" I got this:

which can be pasted directly into Wikipedia. If you are using it for an inline citation, you need to surround this code with <ref></ref>, which will then create a number reference and will add this to the references at the bottom of the page.

Unfortunately, the Open Library code doesn't include a link to the full text, most likely because that isn't part of the Wikipedia format. To do that I added a link to the Internet Archive digitized version of the book after the citation. You can look on the Zahm page and see how that looks. (I'm still looking for a better way to format it so that the link to the full text stands out without looking ugly.)

There is another way to add bibliographic data to Wikipedia which is to click on the menu at the top of the edit window and select a citation type, which then gives you a form to fill out. But if you can find the item in the Open Library, you can avoid all of that typing.


Now that I have learned that it is easy to add bibliographic data to Wikipedia I'm interested in exploring ways that Wikipedia pages can be starting places for essential reading on topics. It naturally makes sense to point to any existing digital materials, but a next logical step would be to find a way to point to libraries for more recent (and in copyright) materials. 


Wednesday, January 16, 2013

Hackers and heroes

Recent events have led me again to a contemplation of the equation of hackers and heroes. How is it that an essentially cerebral and sedentary activity gets equated with heroics? And why computing and not, say, bioscience?

If you've read your obligatory Joseph Campbell you know that the hero myth is ubiquitous in human cultures. Each culture adds its own flavorings and decorations, but the general story is the same: a usually young, alone male goes through transformational trials, performs some task that makes a difference to the world, and is then declared a hero.

In the story-telling world, it ends there. You don't get the post-hero narrative, although, like love stories, there is an implicit "and they lived happily ever after." This makes it easy to forget that in real life "heroism" is a moment, not a lifetime. The fireman who saves the baby from the burning building, the batter who hits the World Series-winning home run: this is a moment of glory before the person goes back to being an ordinary Joe.

When Steven Levy wrote "Hackers, heroes of the computer revolution" in 1984, the hero myth was perhaps new to the computer world. By the early 1990's consumer computing magazines were full of hero and superhero images. This presents an odd contrast to the stereotype of the out-of-shape, asocial, code-writing computer geek. Since then graphics capabilities have allowed the hero myth to move to the screen in the form of first-person games in which anyone with the time and inclination can play at being a hero.

Levy's "hacker heroes" were in fact ordinary computer geeks, and not even the first. Levy focuses on a group of young men at MIT starting in the late 1950's, but they had been preceded in the 1940's and early 1950's by those who were truly the first computer programmers, many of whom were women. There was no attribution of heroics to those pioneers, neither at the time nor retrospectively. What changed?

Many of the studies of the decline of womens enrollment in computer science ask a similar question, which is: how did computer science become a male bastion when it had once seemed welcoming to women? And why did it take on a hyper-masculinized culture, with home brew, skateboarding the hallways, pizza delivered to midnight coding frenzies, and heroes?

I don't have, and have not encountered in my reading, an answer to that question. I do want to caution, however that the hero aspiration has a down side that is played out as tragedy. It might be best to limit our heroes to the mythological realm and leave computing to mortals. It just might become a friendlier place for everybody.

Wednesday, January 02, 2013

OCLC Top 50

OCLC recently released a file of 1.2 million metadata records for the most widely held items in its catalog. These are all items with 250 library holdings or more. I created a list on WorldCat of the top 50, mostly out of curiosity. I was quite surprised at the results, however.

Here's how it breaks down:
  • 16 periodicals, with Time and Newsweek being numbers 1 and 2, respectively
  • 29 kid and YA books, four of which (and very high even in this small list) from the Diary of a Wimpy Kid series
  • 5 adult books
The five adult books are:
  1. McCullough, D. G. (1992). Truman. New York: Simon & Schuster. 
  2. Brown, D. (2003). The Da Vinci code: A novel. New York: Doubleday.
  3. Johnson, S. (1998). Who moved my cheese?: An a-mazing way to deal with change in your work and in your life. New York: Putnam. 
  4. Haley, A. (1976). Roots. Garden City, N.Y: Doubleday.  
  5. Peters, T. J., & Waterman, R. H. (1982). In search of excellence: Lessons from America's best-run companies. New York: Harper & Row
This small set gives me many ideas of things to investigate in the full set. First, the monographs in this set are all recent dates, with the oldest being 1976, and most after 2000.
I am hoping to graph the full set by date. What I expect is that the items will be overwhelmingly recent publications because libraries tend to hold what people read, and my guess is that readers are mainly reading new books. Also, libraries buy from the set of things that are in print, so even if they are buying a so-called classic (as they do every time yet another movie is made of Pride and Prejudice) they are buying a current edition which will have a recent date.

The next obvious bit of information would be correlation between holdings and date, which I expect to be high for the very reasons given above.

The overall distribution of holdings is unsurprising, starting high (at almost 7000 holdings), dropping off dramatically, and creating a long tail. (I had managed to coax a chart of out ooCalc but it crashed before I captured it. Am now studying how to deal with large files and visualization. Advice gladly received.) Of course, the tail would be very, very long if you could chart the entire WorldCat database. (Anyone know how many items in WC are held by only one library? I can't find that in the available WC stats.)

I think it would be interesting to be able to analyze library holdings in correlation with the FRBR-ization that OCLC has done. In fact, I would really like to see the top 1% (or .5%) of FRBR-ized items. Related to FRBR I am mainly wondering if we can estimate how frequently FRBR might fulfill its promise of saving the time of the cataloger. But that's for another day.

Friday, December 07, 2012

Invisible women 2: Cognitive Surplus

I've just read Clay Shirky's 2010 book Cognitive Surplus: creativity and generosity in a connected age. The short summary of the book is: since the 50's people have had more and more leisure time. Until recently that leisure time was taken up with the passive activity of watching television. The Internet has given us the possibility to use our leisure time for social and creative activities, like creating Wikipedia, engaging in online discussion, and even creating lolcats.

Yet I have to ask: how could a smart, well-read professor write an entire book about what people do with their leisure time and not address the well-known and well-documented gender inequality in time available? The OECD did an entire report on what is called "unpaid labor:"
"Most unpaid work is cooking and cleaning – on average 2 hours 8 minutes work per day across the OECD – followed by care for household members at 26 minutes per day. Shopping takes up 23 minutes per day across the OECD on average" Visualized it looks like:


(The left-hand column is minutes, and obviously not all countries are listed.  The full data is available as an Excel file.)

The Economist put it more bluntly (and I do think this image is unnecessarily demeaning, but I haven't found one with the same message that is more neutral):
If they had read Shirky's book, this fellow would have been sitting at his computer, updating Wikipedia pages or adding his cognitive surplus to a discussion group on health issues. But he still would have had more leisure time than a woman.

A social world based on "cognitive surplus" will be one that is not gender neutral. It will have more participation by males, and therefore will be socially skewed to the masculine -- at least until we have gender parity in taking care of the home, the children, the elderly, etc. That is something that I would expect an intelligent observer of society to notice. Not only notice, but to ponder: what does this tell us about the nature of the things being created with this cognitive surplus? Does this explain, in whole or in part, the masculine view of "hacking," the participation in Open Source, the gender nature of games and gaming? 

** I woke up this morning realizing something that is both not in this book but that I hadn't mentioned: the difference in leisure time and income level. I don't have any figures on that right now, but will do some investigation. My assumption is that leisure is not evenly distributed, and that the working poor have much less leisure time than the middle and upper classes. 

Friday, November 23, 2012

Fair Use(-ful)

The beauty and the aggravation of Fair Use in US copyright law is that one cannot pre-define particular uses as "fair." The countries that have, instead, the legal concept of "Fair Dealing" have an enumerated set of uses that are considered fair, although there is obviously still some need for interpretation. The advantage to Fair Use is that it can be re-interpreted with the times without the need for modification of the law. As new technologies come along, such as digitization of previously analog works, courts can make a decision based on the same four factors that have been used for earlier technologies. However, until such a decision is made in a court of law, it isn't possible to be sure whether a use is fair or not.

We have recently seen a court case that decided that HathiTrust's use of digitized books to provide an index to those books is fair. There is another court case that will decide a similar question regarding Google's digitization of books for its Google Book Search. Note, however, that even if both of these are determined to be fair use, each is a particular situation in a particular context. Both organizations have developed their services in an attempt to meet what they judged to be the letter of the law, and yet there is a considerable difference in the services they provide.

HathiTrust stores copies of digitized books from the collections of member libraries. In this case, HT is not itself doing the digitization but is storing files for books mostly digitized by Google. A search in the full text database of OCR'd page images returns, for in-copyright items, the page numbers on which the terms were found, and the number of hits found on each page. There are no snippets and no view of the text unless the text itself is deemed to be out of copyright.

Google has a different approach. To begin with, Google has performed mass digitization of books (estimated at about 20 million) without first obtaining permission from rights holders. So the Google case includes the act of digitization, whereas the HathiTrust case begins with digital files obtained from Google. Therefore the act of digitizing was not a factor in that case. In terms of use of the digitized works, Google also provides keyword searching of the OCR'd digital images, but takes a different approach to the results viewable by the searchers. Google provides short (about 3-5 lines) snippets that show the search terms in context on a page.
Google, however, places specific restrictions to avoid letting users "game" the search to gain access to enough of the text to substitute for actually acquiring access to the book. Here is how Google describes this in its recent legal response:
"The information that appears in Google Books does not substitute for reading the book. Google displays no more than three snippets from a book in response to a search query, even if the search term appears many times in the book. ... Google also prevents users from view a full page, or even several contiguous snippets, by displaying only one snippet per page in response to a given search and by 'blacking' (i.e. making unable for snippet view in response to any search) at least one snippet per page and one out of ten pages in a book." p.8
Google also exempts some types of books, like reference works, cookbooks, and poetry, from snippet display entirely.

The differences in the results returned by these two services reflect the differences in their contexts and their goals. HathiTrust has member institutions and their authorized users. The collection within HathiTrust reflects the holdings of the member institutions' libraries which means that the authorized users should have access, either in their library or through inter-library loan, to the physical book that was scanned. The HathiTrust full text is a search on the members' "stuff." The decision to give only page numbers makes some sense in this context, although providing snippets to scholars might have been acceptable to the judge. The return of page numbers and full word counts within pages reflects, IMO, the interest in quantitative analysis of term use. It also gives scholars some idea of the weight the term has within the text.

Google's situation is different. Google has no institutions, no members, no libraries; it provides its service to the general public (at least to the US public). There is no reason to assume that all of the members of that public will have access to the hard copy of any particular digitized book. Google seems to have decided that promoting its service as having primarily a marketing function, with the snippets as "teasers," would mollify the various intellectual property owners. In its brief of November 9, Google reiterates that it does not put advertising on the Google Book Search results pages, nor does Google make any money off of its referrals to book purchasing sites.

So here are two organizations that have bent over backwards to stay within what they deemed to be the boundaries of fair use, and they have done so in significantly different ways. This means that the fair use determination of each of these could have different outcomes, and each will provide different clues as to how fair use is viewed for digitized works.

It of course bears mentioning that both of these solutions provide hurdles for users. The HathiTrust user who is searching on a term that could have more than one meaning ("iron" "dive" "foot") does not have any context to help her understand if the results are relevant. The Google user, on the other hand, gets some context but cannot see all of the results and therefore does not know if there are key retrievals among those that have been blocked algorithmically. A use that is "fair" within copyright law may not seem "fair" to the user who is doing research. It makes you wonder if our idea of "fair use" couldn't be extended to be fair but also "useful."

Related posts
http://kcoyle.blogspot.com/2012/10/copyright-victories-part-ii.html

Thursday, November 01, 2012

Turing's Cathedral, or Women Disappear

"She features significantly in computing historian George Dyson's book, Turing's Cathedral: The Origins of the Digital Universe, ISBN 978-0375422775."
From the Wikipedia article for Klara Dan von Neumann

Unfortunately, she features significantly mainly as von Neumann's wife, even though she also was "a pioneer computer programmer," as per the Wikipedia article. In fact, of the 35 women whose names are in the book's index, 24 are in the book as wives, including Klara. Klara is the only one who gets a full bio and a fair amount of ink. Much of the ink comes from her unfinished memoirs about her life as von Neumann's wife. She was also one of the primary programmers working on the ENIAC, and Dyson's book names her as one of the first three programmers, along with her husband, programming ENIAC. (p. 104). Her work, however, is described as "help," one of the ways that women's activities are diminished in importance (men "do", women "help"):
"'With the help of Klari von Neumann,' says Metropolis, 'plans were revised and completed and we undertook to implement them on the ENIAC...'" p. 194
Yet she obviously provided more than "help." In fact, she invented:
"'Your code was described and was impressive,' von Neumann wrote to Klari from Los Alamos, discussing whether a routine she had developed should be coded as software or hardwired into the machine. 'They claim now, however, that making one more, 'fixed,' function table is so little work, that they want to do it. It was decided that they will build one, with the order soldered in." (p. 195)
Of the other women mentioned, one is a secretary, the other the manager of the cafeteria. The saddest story is that of Bernetta Miller, the fifth licensed woman pilot in the US who was a demonstration pilot for an airplane company, volunteered for duty in WWI and was wounded, then became secretary to the directory of the Institute for Advanced Study in Princeton. In the Dyson book, she is mainly remembered for her memoranda about dining room accounting, and for being fired by Oppenheimer. (p. 91-92)

There are eight women, other than Klara, who are in the book in their professional positions. Three of them are mentioned in a single sentence as "computers," that is people (mainly women) who did the hard math by hand before the machine computers were up to the job. (see: http://en.wikipedia.org/wiki/Human_computer, and I highly recommend the books by Grier in the bibliography if you wish to learn of the sophistication of methods that were developed by the "girls.")

One woman, Mina Rees, is named twice as someone who was written to:
"... Goldstine had written to Mina Rees of the Office of Naval Research." p. 147
"... 'The best change for a real undersatnding of protein chemistry lies in the x-ray diffraction field,' he wrote to Mina Rees at the Office of Naval Research." p.229
Later there is a quote from a report that states,
"... was informed by Dr. Mina Rees and Colonel Oscar Maier, representing the Office of Naval Research, and the Air Material Command, respectively..." p. 321
In themselves these quotes are not important, but this is one of the few professional women who gets mentioned in the book, and this is all that is said about her. Dr. Mina Rees was an amazing character: "She earned her doctorate in 1931 with a thesis on "Division algebras associated with an equation whose group has four generators," published in the American Journal of Mathematics, Vol 54 (Jan. 1932), 51-65. Her advisor was Leonard Dickson." (Wikipedia article) At the time of these references she was head of the Mathematics Department at the Office of Naval Research.

There are some other minor mentions, like one of Meg Ryan in a parenthetical sentence about a named location that was later used in a movie, and one woman mathematician who was named with two male mathematicians in a single sentence. These obviously are not major characters in the book, and the book is wide-ranging with everything from Aldus Huxley to George Washington, also not major characters.

The real mystery woman is Hedvig Selberg.
"'... says Atle Selberg, whose wife, Hedi was hired by von Neumann on September 29, 1950, and remained with the computer project until its termination in 1958.'" p. 152
Later we get a short bio of her: born in 1919 in Transylvania, graduated with a master's degree in mathematics at the head of her class, and was the only family member to survive Auschwitz. She came to the U.S. and was hired to work on the first computer project. She seems to have worked closely with a Martin Schwarzschild on a complex model of stellar evolution that related to the radiation effects of the bomb that was being designed. Schwarzschild went on to fame, as did Selberg's husband, a mathematician. Hedvig didn't even rate an obituary in the big newspapers (nor a Wikipedia article), although she is mentioned in her husband's obit where he first marries her, then she dies (in 1995) and he remarries.
"His first wife, Hedvig Liebermann, a researcher at the institute and Princeton's Plasma Physics Laboratory, died in 1995." (NYT Aug 17, 2007)
(Note: Mina Rees did get a NY Times obit. )

Among the other striking aspects of this treatment of women (and this book isn't by any means unusual in this respect) is that women tend not to exist until they marry a man of interest, and then suddenly they appear on the scene. Men, on the other hand, have parents and educations and often interesting stories that are told in the book, both as character building but also as bone fides. It is therefore a bit of a shock to learn in some aside that the wife has a PhD in "trans-sonic aerodynamics" as in the case with Kathleen Booth. (p. 133)

Admittedly the opportunities for women in science were very limited in the period being discussed in this book, the 1950's. However, the role of a historian is to go beyond the period's view of itself and tease out a deeper meaning from the privileged position of hind-sight. I have read other histories of computing that also failed to notice that there were women involved in the invention of this field, but this one has come out in 2012. Really, we didn't need another book on the topic written with male blinders. What a shame.

Suggested reading:
Noble, David F. The Religion of Technology : the Divinity of Man and the Spirit of Invention. 1st ed. New York: A.A. Knopf :, 1997.
Mozans, H. J. Woman in Science; with an Introductory Chapter on Woman’s Long Struggle for Things of the Mind,. New York,: D. Appleton and company, 1913.
Toole, Betty A., and Ada King Lovelace. Ada, the Enchantress of Numbers : a Selection from the Letters of Lord Byron’s Daughter and Her Description of the First Computer. 1st ed. Mill Valley, Calif.: Strawberry Press ;, 1992.
Grier, David Alan. When Computers Were Human. Princeton: Princeton University Press, 2005. Print.



Thursday, October 18, 2012

Is Linked Data the Answer?

I recently gave keynote talks at Dublin Core 2012 and Emtacl12 with the title "Think 'Different'." Since the slides of my talks don't generally have much text on them, I wrote up the talk as a document. The document has a kind of appendix covering the point in my presentation where I took advantage of my position on stage to ask and answer what I think is a common question: Is linked data the answer?

Many would expect me to answer "yes" to this question, but my answer is a bit more complex. Linked data is a technology that I believe we will make use of to connect library data to other information resources. That's what the "linked" in linked data is all about -- creating a web of information by connecting bits of data in different documents and datasets. However, we have to be very cautious about having "an answer." When you have an answer you tend to stop looking at the questions that arise, and you also tend to ignore questions that aren't going to be solved by the answer you have chosen. There is no technology that will do everything that we need, so while linked data can be useful for some things we may need to do, it cannot be the answer to all of our technical requirements.

Note that I describe linked data as "connecting bits of data." The origin of the semantic web is in the need and desire to make actionable data that today is essentially hidden within the text of documents. For example, if I say:

"My name is Karen. I will be holding a webinar on June 4 at 3:00 Pacific time for anyone who wants to learn about my paperweight collection."

That's text. There is interesting information in there, but it isn't available for any computational uses. The Semantic Web, as implemented through linked data, would make that information actionable. There are various ways to do this, and one is through the use of microformats which mark up data within a document. This could look something like:

<p>My name is <span class="name">Karen</span>. I will be holding a <span class="event">webinar</span> on <span class="datetime" title="2012-06-04T03:00-09:0000">June 4 at 3:00 Pacific time</span> for anyone who wants to learn about my <span topic="paperweights">paperweight</span> collection.</p>

This text now also has bits of data that can be used for various purposes, including linking. The linking capabilities in this particular example are low, but some additional information, like standard identifiers for the person and for the topic, would then increase the linkability of this data.

<p>My name is <span class="name" id="http://viaf.org/viaf/48369992/">Karen</span>. I will be holding a <span class="event">webinar</span> on <span class="datetime" title="2012-06-04T03:00-09:0000">June 4 at 3:00 Pacific time</span> for anyone who wants to learn about my <span topic="paperweights" id="http://id.loc.gov/authorities/subjects/sh85097666.html">paperweight </span> collection.</p>

This isn't a perfect example, but I wouldn't claim that we're heading toward perfect data. What we need is to get more out of the information we have. 

I perceive an assumption in the library linked data movement that what the Web needs (because linked data is data on the Web) is our bibliographic data. I disagree. The Web is awash in bibliographic data - from Amazon to Google Books, from fan sites like IMDB or MusicBrainz, and from sharing sites like LibraryThing and GoodReads. Libraries may have some unique bibliographic data, but most of what we have would duplicate what is already there, many times over.

There's also the fact that much of bibliographic data isn't DATA in the linked data sense. It isn't actionable data elements for the most part. In fact, bibliographic data is more like a structured document: it mainly has text, and that text is to be displayed to humans. It is possible to extract actual data (dates of publication, numbers of pages, various identifiers), but the text itself is a large part of the point about bibliographic data.

What this means for us in libraries is that we shouldn't be thinking that linked data will replace bibliographic data. It will encode the aspects of bibliographic data that will give us the most and the best links.

Then we need to ask: why are we linking? What will we get? Well, we can get connections between books and maps, between books and documents, and between search retrievals and libraries. This latter interests me especially. Google is experimenting with using microformat data, in particular the schema.org data that it is fostering along with Yahoo!, Bing, and Yandex (the Russian search engine).  Schema.org microformat data allows a search engine like Google to enrich the snippets with more than just a block of text from the page. This is an example from the Google Webmaster pages on Rich Snippets:
Below is my conceptualization of what we could do with library data. The bibliographic data, as I've said before, often already exists on the Web and we may not be helping things by adding many more duplicate copies of that data. But what we have in libraries that no one else has is library holdings data. We know where Web users can find "stuff" in their local community. If that could be linked to the Web, a future rich snippet might look like:

Obviously there are steps to be taken to make this possible, but if you want to think about how library data might fit into the Web of data that information seekers make use of millions or billions of times a day, this is one option. It's a start, and it uses data we already have.

You can take a look at the schema.org data that is created for WorldCat records simply by doing a WorldCat search and scrolling down to the section called "Linked data." The number of holdings is included (and this in itself is something that might interest Google as a measure of popularity). Making the link to the holdings of an actual library, and making that possible for all libraries, not just OCLC member libraries, is something I consider a worthy experiment for linking library data.