I attended Code4lib 2008 in Portland Oregon at the end of February, and I must say it was the most intense and rewarding and even thrilling conference I've ever been to. Imagine 200 coders in a room, laptop on each lap, programs running, chat flying, boosterism abounding. Lightening talks (5 minutes or less) introduced quick hacks with stunning results (none of which I have been able to replicate with my own very modest skills, but I am sure that others have had more success).
It was an honor to be speaking at that event, and I used the opportunity to present the reasoning behind the RDA Vocabularies project. I have to say that it was an easy audience because a room full of coders understands the difficulty of working with unstructured textual data, which is a lot of what we work with in libraryland. The RDA Vocabularies project is developing a set of elements (or "properties" in RDF-speak) that will facilitate a more machine-friendly approach to bibliographic data, without compromising on user-friendliness. But the main point of the Vocabularies project for me is that we can create better user services with data that plays well on the Web.
The talk was filmed, but the videos haven't been set up for streaming yet. I added text to my slides (PDF), making some of them much too wordy (it seemed to take fewer words when I was speaking it). If I learn to do videocasts (a current goal of minie), I will try to get some of this content available for viewing+listening.
(Note: Jon Phipps, code4lib attendee and programmer on the NSDL registries project, has an interesting post on using RDF triples with DC and RDA. I think I'll get it if I read it 3 or 4 more times.)
Monday, March 17, 2008
Thursday, February 21, 2008
Girls and the Internet
Today's New York Times has an article about the dominance of young females on the Internet. According to the article, 35% of girls (age 12-17) blog, compared to 20% of boys. And 32% of girls create or work on their own Web pages, compared to 22% of boys. They also greatly outnumber boys in having social networking sites. The article gives examples of girls who create and advise on CSS code creation and who design icons and animated icons. Some of these young women are making money (at least some money, no figures are given) from their sites.
So where was this article placed in the paper? In the business section with other "Internet entrepreneurs"? In the technology section? No, it was in Fashion & Style, right under an article on wedding dresses.
*sigh*
So where was this article placed in the paper? In the business section with other "Internet entrepreneurs"? In the technology section? No, it was in Fashion & Style, right under an article on wedding dresses.
*sigh*
Tuesday, February 19, 2008
Libraries -- Open
I think of libraries as the quintessential open institutions. At least in the U.S. Most libraries are physically open to the public (even those in private institutions), and many serve as community spaces focused on learning, exploring, and simply being. They also promote the open use of what we might call "cultural heritage resources." Libraries fight for open access and they fight censorship all of the way to the Supreme Court.
Yet - there seems to be a real barrier when it comes to libraries being open with their own catalog data. This seems rather odd because most libraries' catalogs are available for open access on the web. But try asking a library for some or all of that data and you suddenly hit a wall. Libraries don't like to say "no" so there's a lot of hemming, hawing, "we'll think about that," that goes on. But an out-and-out "yes" is rare.
I speak about this based on my experience with the Open Library, a project of the Internet Archive. The OL wants to create a humongous bibliographic database (right now only of book records) on the Web, using a wiki-like front-end that would allow anyone to edit the bibliographic data. To me the most interesting aspect of the project is that it would bring bibliographic entries to the web's surface; they could be the subject of links from other documents, and potentially could begin the creation of a bibliographic web linking books to each other. But in spite of putting out a call for bibliographic records and making personal and direct pleas to a number of libraries, the OL has received only a lukewarm response.
To be sure, any data submitted to the Internet Archive becomes publicly available. And at some time in the future it may be possible for people to download individual bibliographic records for their own use. I know that there is some speculation that OCLC "owns" the data and that the OCLC license may not allow this level of re-distribution. I also know that some records in library databases are covered by vendor licenses (other than OCLC). Presumably those could be excluded from the data set. But it still surprises me how un-open libraries are with their own data, given how much they encourage others to be open with theirs.
During the comment period for the Future of Bibliographic Control report, the Open Knowledge Foundation posted a call for library data openness on the OKF wiki. Many dozens signed their names to the OKF's call for open licensing of bibliographic data, including important people like Larry Lessig and Tim O'Reilly. The arguments in OKF's document seemed pretty clear to me:
Libraries complain that they don't get the kind of attention that Web resources like Google and OL get. They complain about the lack of transparency of the commercial data vendors; that Google won't say how many books it has online nor will it reveal its work on attempts to rank book retrievals. Libraries could be doing this experimentation themselves, and in the open, if their data were on the Web. They could be visible, out there, allowing incredible innovation to happen based on the hundreds of years of collecting materials and creating relatively consistent metadata for those materials. Their reluctance to let their data out of the databases just baffles me, and isn't in concert with their stated goals of open access.
Yet - there seems to be a real barrier when it comes to libraries being open with their own catalog data. This seems rather odd because most libraries' catalogs are available for open access on the web. But try asking a library for some or all of that data and you suddenly hit a wall. Libraries don't like to say "no" so there's a lot of hemming, hawing, "we'll think about that," that goes on. But an out-and-out "yes" is rare.
I speak about this based on my experience with the Open Library, a project of the Internet Archive. The OL wants to create a humongous bibliographic database (right now only of book records) on the Web, using a wiki-like front-end that would allow anyone to edit the bibliographic data. To me the most interesting aspect of the project is that it would bring bibliographic entries to the web's surface; they could be the subject of links from other documents, and potentially could begin the creation of a bibliographic web linking books to each other. But in spite of putting out a call for bibliographic records and making personal and direct pleas to a number of libraries, the OL has received only a lukewarm response.
To be sure, any data submitted to the Internet Archive becomes publicly available. And at some time in the future it may be possible for people to download individual bibliographic records for their own use. I know that there is some speculation that OCLC "owns" the data and that the OCLC license may not allow this level of re-distribution. I also know that some records in library databases are covered by vendor licenses (other than OCLC). Presumably those could be excluded from the data set. But it still surprises me how un-open libraries are with their own data, given how much they encourage others to be open with theirs.
During the comment period for the Future of Bibliographic Control report, the Open Knowledge Foundation posted a call for library data openness on the OKF wiki. Many dozens signed their names to the OKF's call for open licensing of bibliographic data, including important people like Larry Lessig and Tim O'Reilly. The arguments in OKF's document seemed pretty clear to me:
Bibliographic records are a key part of our shared cultural heritage. They too should therefore be made available to the public for access and re-use without restriction. Not only will this allow libraries to share records more efficiently and improve quality more rapidly through better, easier feedback, but will also make possible more advanced online sites for book-lovers, easier analysis by social scientists, interesting visualizations and summary statistics by journalists and others, as well as many other possibilities we cannot predict in advance.Nothing of this was included in the final report.
Libraries complain that they don't get the kind of attention that Web resources like Google and OL get. They complain about the lack of transparency of the commercial data vendors; that Google won't say how many books it has online nor will it reveal its work on attempts to rank book retrievals. Libraries could be doing this experimentation themselves, and in the open, if their data were on the Web. They could be visible, out there, allowing incredible innovation to happen based on the hundreds of years of collecting materials and creating relatively consistent metadata for those materials. Their reluctance to let their data out of the databases just baffles me, and isn't in concert with their stated goals of open access.
Sunday, February 03, 2008
The ILS minus the catalog
The greatest amount of action happening today regarding library user services is the separation of the user interface from the integrated library system (ILS). This seems odd, perhaps, since only two decades ago the integration of all of the functions of library systems was seen as a real step forward. Until then, one system had handled acquisitions, another circulation, and another cataloging. Many functions, such as serials check-in and bindery management, were not managed through automation. This situation had a number of problems: different data about the same book were stored in multiple databases or in card files, leading to inconsistencies throughout the system; the data had to be keyed or copied multiple times; system-wide updates were nearly impossible. The "integration" of the integrated library system was the creation of a single database for bibliographic and management data, where all of the information about an item would be stored once and only once. This also was the first time that the full bibliographic record was linked to the library management functions. Independent systems like acquisitions and circulation systems primarily used brief records only. At best, these brief records contained an identifier (such as the item's barcode or call number) that could connect it to other records in other systems. Sometimes even that wasn't possible.
In theory, this database integration is a dandy way to organize your data and the activities that use your data. In reality, the user interface suffered in this design. Not that anyone purposely shorted the user interface, but in a world of scarcity, there are things that just have to get done; and then there are other things. In the have to category, libraries have to make purchases, manage accounts, receive and check-in serials, perform interlibrary loans, and check items out to borrowers. These are clear, quantitative, auditable library functions, and ones that library administrators focus on. These are the functions that can have dollar amounts attached to them in terms of staff time. These functions are the inside view of the library, the library being a library.
User success, on the other hand, is qualitative, hard to define, and does not have a direct effect on the library's bottom line. We count the countables, like numbers of bibliographic records, items circulated, and online database accesses, but there appears to be no penalty for a lousy user interface and no premium for the creation of a good one. If at any point there is a conflict between quantitative library management and qualitative user service, my gut feeling is that the latter loses out.
Users have everything to gain from the separation of the user interface from the library management system. Libraries, however, are in a bit of a bind. The new "user interface on top of the ILS" adds features for users but it doesn't result in any less work in the ILS. Libraries are still hanging all of their management functions off of full bibliographic records in the catalog (which the users no longer see). Librarians still see the data creation functions in the early management steps of acquisitions and receipt to be a direct line to the standard bibliographic record that in the end will appear on the users' screens. They are still storing the full bibliographic records in a local database, although these records are siphoned off nightly to the "real" user interface.
Much of the objection to using more EDI (electronic data interchange) functions with our vendors is that their data doesn't conform to library cataloging. Yet our library management systems are getting further from the user interface. We may need to rethink the library management workflow as well as the basis of our cataloging activity. What could we achieve if we move cataloging and catalogs out of our individual library databases to the network level? Could this provide the basis for increased sharing of the cataloging effort? Are there other efficiencies that could be gained in the "back room" functions of purchasing and managing the library inventory?
In theory, this database integration is a dandy way to organize your data and the activities that use your data. In reality, the user interface suffered in this design. Not that anyone purposely shorted the user interface, but in a world of scarcity, there are things that just have to get done; and then there are other things. In the have to category, libraries have to make purchases, manage accounts, receive and check-in serials, perform interlibrary loans, and check items out to borrowers. These are clear, quantitative, auditable library functions, and ones that library administrators focus on. These are the functions that can have dollar amounts attached to them in terms of staff time. These functions are the inside view of the library, the library being a library.
User success, on the other hand, is qualitative, hard to define, and does not have a direct effect on the library's bottom line. We count the countables, like numbers of bibliographic records, items circulated, and online database accesses, but there appears to be no penalty for a lousy user interface and no premium for the creation of a good one. If at any point there is a conflict between quantitative library management and qualitative user service, my gut feeling is that the latter loses out.
Users have everything to gain from the separation of the user interface from the library management system. Libraries, however, are in a bit of a bind. The new "user interface on top of the ILS" adds features for users but it doesn't result in any less work in the ILS. Libraries are still hanging all of their management functions off of full bibliographic records in the catalog (which the users no longer see). Librarians still see the data creation functions in the early management steps of acquisitions and receipt to be a direct line to the standard bibliographic record that in the end will appear on the users' screens. They are still storing the full bibliographic records in a local database, although these records are siphoned off nightly to the "real" user interface.
Much of the objection to using more EDI (electronic data interchange) functions with our vendors is that their data doesn't conform to library cataloging. Yet our library management systems are getting further from the user interface. We may need to rethink the library management workflow as well as the basis of our cataloging activity. What could we achieve if we move cataloging and catalogs out of our individual library databases to the network level? Could this provide the basis for increased sharing of the cataloging effort? Are there other efficiencies that could be gained in the "back room" functions of purchasing and managing the library inventory?
Friday, January 25, 2008
Books as Social Vectors
Ursula Le Guin has a fabulous article in Harper's (Feb. 2008, v. 316, n. 1893) responding to the NEA report on Reading at Risk. That report states that there has been a sharp decline in the reading of books of "literature" (which I couldn't find a definition for in the report).
Le Guin's article is called "Staying Awake," which comes from one person's statement "I just get sleepy when I read." As Le Guin points out, there are "people who read wide awake," but the corporate culture of today's publishing isn't interested in cultivating anything except the "best seller" product. (Some books are art "And the relationship of art to capitalism is, to put it mildly, vexed.")
She talks about what reading has meant to culture ("Books are social vectors..."), from the early use of books to spread a uniform view of religion, to the late 19th century serial books that had everyone discussing what would happen next. It is this aspect of books as social vectors that I think we in the library world need to come to grips with.
The public library of the 19th century was about bringing book culture to the masses. (See Dee Garrison's book Apostles of Culture for a good account.) Somewhere in the 20th century we swung the pendulum in the opposite direction and began aiming for maximum neutrality. But people don't respond well to neutrality. In fact, they are... well, neutral on it. It takes a certain interest, perhaps even passion, to stay awake.
There are obvious issues for libraries (many of them government agencies) should they become instigators of passion for books. However, I see a somewhat less problematic possibility, which is allowing the library to itself be a "social vector" by connecting the library, and in particular its catalog, to the world of social networking. This is starting to happen in a small way, such as links from web sites or social bookmarking tools to WorldCat, but I think it's time to really ratchet up our efforts in this area.
Le Guin's article is called "Staying Awake," which comes from one person's statement "I just get sleepy when I read." As Le Guin points out, there are "people who read wide awake," but the corporate culture of today's publishing isn't interested in cultivating anything except the "best seller" product. (Some books are art "And the relationship of art to capitalism is, to put it mildly, vexed.")
She talks about what reading has meant to culture ("Books are social vectors..."), from the early use of books to spread a uniform view of religion, to the late 19th century serial books that had everyone discussing what would happen next. It is this aspect of books as social vectors that I think we in the library world need to come to grips with.
The public library of the 19th century was about bringing book culture to the masses. (See Dee Garrison's book Apostles of Culture for a good account.) Somewhere in the 20th century we swung the pendulum in the opposite direction and began aiming for maximum neutrality. But people don't respond well to neutrality. In fact, they are... well, neutral on it. It takes a certain interest, perhaps even passion, to stay awake.
There are obvious issues for libraries (many of them government agencies) should they become instigators of passion for books. However, I see a somewhat less problematic possibility, which is allowing the library to itself be a "social vector" by connecting the library, and in particular its catalog, to the world of social networking. This is starting to happen in a small way, such as links from web sites or social bookmarking tools to WorldCat, but I think it's time to really ratchet up our efforts in this area.
Friday, January 18, 2008
Being Careful
The Library of Congress has put some great collections of photos up on Flickr. Take a look at the group on the 1930's and 40's. There are some great photos of "Rosie the Riveter" women building various implements of war, especially aircraft. (Note: the photos do look staged, or at least they gave the women a chance to freshen their lipstick before the shoot.)
Type the word "careful" into the search box at the top and you'll get an idea of that era's nervousness about having women work on technology, and the qualities that women were seen as bringing to the job. (Hint: it's not innate mechanical ability.) These women really did show that "we can do it." Training, schmaining -- just give me a power tool and turn me loose!
Type the word "careful" into the search box at the top and you'll get an idea of that era's nervousness about having women work on technology, and the qualities that women were seen as bringing to the job. (Hint: it's not innate mechanical ability.) These women really did show that "we can do it." Training, schmaining -- just give me a power tool and turn me loose!
Friday, January 11, 2008
ALCTS CCS Discussion of RDA Draft
Committee on Cataloging: Description and Access (also known as CC:DA)
This is an informal discussion on the draft of RDA. Focuses on sections 2 & 3; the remainder will be covered on Monday.
Overall Comments
Discussing:
Chapter 2. Identifying manifestations and items
Chapter 3. Describing carriers
This was a discussion of RDA by the ALCTS cataloging group Presumably the comments here become the ALA comments to the JSC.
It's hard to characterize this discussion. It varied between comments about the need to improve the definitions and the problems with the structure of the document, and with statements like: What we have here is a crisis of confidence; pedantic adherence to structural hierarchy; Why is it that the rules from AACR2 do not inspire confidence in this new setting?
Here are some of the more interesting issues that came up (filtered, obviously, through my viewpoint).
- Some of the text is just AACR2, pulled into the RDA document. This is considered by some to be a step backward -- that this "new" code doesn't take advantage of the opportunity to make changes in these areas.
- There seems to be a great deal of confusion on what the final RDA product will actually be. Some see it as the final cataloging code that they will use daily in their cataloging. Others (possibly the members of JSC) see RDA as being a basis for cataloging, but a neutral background for the creation of actual cataloging rules. This is particularly odd when we consider that the RDA text is being transferred to an online system which will be the primary product allowing people to access RDA.
- This also means that there is tension between creating a general code and getting all of the special rules in for music, law, cartographic materials, etc. This tension does not seem to be resolved, and there are people with different expectations.
- The who RDA direction seems to be in incredible flux. You may know that they recently announced a restructuring of the document to make it in more line with FRBR. They also seem to be attempting to do some redesign of concepts. For example, there is no longer any reference to authority records -- it is assumed that in the future, those will not exist as they do today, although the same information will be carried somewhere.This is a major change - at least in thinking, to happen just months before the full draft is due to be available.
- There was a fair amount of dissent over the format of the text and the fact that it will be thousands of pages in length when created. There's a deep contradiction in the process, because the RDA folks say that they are writing the online version, but they are creating this as a print document. As some members of the audience pointed out, they are really creating neither -- what they have doesn't work as a print document, and the web document will undoubtedly look quite different. So... what is it that we are looking at now?
- Although there is quite a bit of dissent in the US over RDA, there is great enthusiasm among the non-US members of the Joint Steering Committee. We don't have any explanation as to why we have these polar opposites, and it would be very interesting to hear WHY they think it's so good, sonce here it seem to be almost universally disliked.
- There was a fascinating, but not quite coherent, discussion of persons and personal names: are we identifying persons, or are we identifying names? If the same person uses more than one name, how many identities is that? It was said that we are now treating persons like corporate bodies: a difference in naming is a different entity. This has some practical elements, of course, but it also seems to be deeply philosophical and something that we have to be very clear on if we are going to exchange data with communities who emphasize persons over named identities.
This is an informal discussion on the draft of RDA. Focuses on sections 2 & 3; the remainder will be covered on Monday.
Overall Comments
Discussing:
Chapter 2. Identifying manifestations and items
Chapter 3. Describing carriers
This was a discussion of RDA by the ALCTS cataloging group Presumably the comments here become the ALA comments to the JSC.
It's hard to characterize this discussion. It varied between comments about the need to improve the definitions and the problems with the structure of the document, and with statements like: What we have here is a crisis of confidence; pedantic adherence to structural hierarchy; Why is it that the rules from AACR2 do not inspire confidence in this new setting?
Here are some of the more interesting issues that came up (filtered, obviously, through my viewpoint).
- Some of the text is just AACR2, pulled into the RDA document. This is considered by some to be a step backward -- that this "new" code doesn't take advantage of the opportunity to make changes in these areas.
- There seems to be a great deal of confusion on what the final RDA product will actually be. Some see it as the final cataloging code that they will use daily in their cataloging. Others (possibly the members of JSC) see RDA as being a basis for cataloging, but a neutral background for the creation of actual cataloging rules. This is particularly odd when we consider that the RDA text is being transferred to an online system which will be the primary product allowing people to access RDA.
- This also means that there is tension between creating a general code and getting all of the special rules in for music, law, cartographic materials, etc. This tension does not seem to be resolved, and there are people with different expectations.
- The who RDA direction seems to be in incredible flux. You may know that they recently announced a restructuring of the document to make it in more line with FRBR. They also seem to be attempting to do some redesign of concepts. For example, there is no longer any reference to authority records -- it is assumed that in the future, those will not exist as they do today, although the same information will be carried somewhere.This is a major change - at least in thinking, to happen just months before the full draft is due to be available.
- There was a fair amount of dissent over the format of the text and the fact that it will be thousands of pages in length when created. There's a deep contradiction in the process, because the RDA folks say that they are writing the online version, but they are creating this as a print document. As some members of the audience pointed out, they are really creating neither -- what they have doesn't work as a print document, and the web document will undoubtedly look quite different. So... what is it that we are looking at now?
- Although there is quite a bit of dissent in the US over RDA, there is great enthusiasm among the non-US members of the Joint Steering Committee. We don't have any explanation as to why we have these polar opposites, and it would be very interesting to hear WHY they think it's so good, sonce here it seem to be almost universally disliked.
- There was a fascinating, but not quite coherent, discussion of persons and personal names: are we identifying persons, or are we identifying names? If the same person uses more than one name, how many identities is that? It was said that we are now treating persons like corporate bodies: a difference in naming is a different entity. This has some practical elements, of course, but it also seems to be deeply philosophical and something that we have to be very clear on if we are going to exchange data with communities who emphasize persons over named identities.
Subscribe to:
Posts (Atom)