Friday, January 25, 2008
Books as Social Vectors
Le Guin's article is called "Staying Awake," which comes from one person's statement "I just get sleepy when I read." As Le Guin points out, there are "people who read wide awake," but the corporate culture of today's publishing isn't interested in cultivating anything except the "best seller" product. (Some books are art "And the relationship of art to capitalism is, to put it mildly, vexed.")
She talks about what reading has meant to culture ("Books are social vectors..."), from the early use of books to spread a uniform view of religion, to the late 19th century serial books that had everyone discussing what would happen next. It is this aspect of books as social vectors that I think we in the library world need to come to grips with.
The public library of the 19th century was about bringing book culture to the masses. (See Dee Garrison's book Apostles of Culture for a good account.) Somewhere in the 20th century we swung the pendulum in the opposite direction and began aiming for maximum neutrality. But people don't respond well to neutrality. In fact, they are... well, neutral on it. It takes a certain interest, perhaps even passion, to stay awake.
There are obvious issues for libraries (many of them government agencies) should they become instigators of passion for books. However, I see a somewhat less problematic possibility, which is allowing the library to itself be a "social vector" by connecting the library, and in particular its catalog, to the world of social networking. This is starting to happen in a small way, such as links from web sites or social bookmarking tools to WorldCat, but I think it's time to really ratchet up our efforts in this area.
Friday, January 18, 2008
Being Careful
Type the word "careful" into the search box at the top and you'll get an idea of that era's nervousness about having women work on technology, and the qualities that women were seen as bringing to the job. (Hint: it's not innate mechanical ability.) These women really did show that "we can do it." Training, schmaining -- just give me a power tool and turn me loose!
Friday, January 11, 2008
ALCTS CCS Discussion of RDA Draft
This is an informal discussion on the draft of RDA. Focuses on sections 2 & 3; the remainder will be covered on Monday.
Overall Comments
Discussing:
Chapter 2. Identifying manifestations and items
Chapter 3. Describing carriers
This was a discussion of RDA by the ALCTS cataloging group Presumably the comments here become the ALA comments to the JSC.
It's hard to characterize this discussion. It varied between comments about the need to improve the definitions and the problems with the structure of the document, and with statements like: What we have here is a crisis of confidence; pedantic adherence to structural hierarchy; Why is it that the rules from AACR2 do not inspire confidence in this new setting?
Here are some of the more interesting issues that came up (filtered, obviously, through my viewpoint).
- Some of the text is just AACR2, pulled into the RDA document. This is considered by some to be a step backward -- that this "new" code doesn't take advantage of the opportunity to make changes in these areas.
- There seems to be a great deal of confusion on what the final RDA product will actually be. Some see it as the final cataloging code that they will use daily in their cataloging. Others (possibly the members of JSC) see RDA as being a basis for cataloging, but a neutral background for the creation of actual cataloging rules. This is particularly odd when we consider that the RDA text is being transferred to an online system which will be the primary product allowing people to access RDA.
- This also means that there is tension between creating a general code and getting all of the special rules in for music, law, cartographic materials, etc. This tension does not seem to be resolved, and there are people with different expectations.
- The who RDA direction seems to be in incredible flux. You may know that they recently announced a restructuring of the document to make it in more line with FRBR. They also seem to be attempting to do some redesign of concepts. For example, there is no longer any reference to authority records -- it is assumed that in the future, those will not exist as they do today, although the same information will be carried somewhere.This is a major change - at least in thinking, to happen just months before the full draft is due to be available.
- There was a fair amount of dissent over the format of the text and the fact that it will be thousands of pages in length when created. There's a deep contradiction in the process, because the RDA folks say that they are writing the online version, but they are creating this as a print document. As some members of the audience pointed out, they are really creating neither -- what they have doesn't work as a print document, and the web document will undoubtedly look quite different. So... what is it that we are looking at now?
- Although there is quite a bit of dissent in the US over RDA, there is great enthusiasm among the non-US members of the Joint Steering Committee. We don't have any explanation as to why we have these polar opposites, and it would be very interesting to hear WHY they think it's so good, sonce here it seem to be almost universally disliked.
- There was a fascinating, but not quite coherent, discussion of persons and personal names: are we identifying persons, or are we identifying names? If the same person uses more than one name, how many identities is that? It was said that we are now treating persons like corporate bodies: a difference in naming is a different entity. This has some practical elements, of course, but it also seems to be deeply philosophical and something that we have to be very clear on if we are going to exchange data with communities who emphasize persons over named identities.
ALA Friday "Big Heads"
Big heads meeting
RDA update - John Attig
Some drafts are still out for review and will have to be reconciled in this new structure. The committee is still working on specific comments on chapters that have been reviewed. They still need to do appendices and examples. They will meet for two weeks in April.
Some background: over last 6 weeks, after putting out the final draft for comment, the group got 150 pages of comments. Some contradict each other. Many resulted in revision of text to clarify meaning in the report.
Areas of comment: clarification of meaning of particular sections, which they did. There was also a desire to be more specific about the 'how'. That wasn't for group to do, and in any case there was not time for that kind of analysis. As an example, the call for a new carrier – this is not easy to do and this wasn't the right group. but hopefully more energy will be put to it.encouraging discussion, in groups like this.
People also wanted to know exactly WHO would do certain tasks. Many Cannot be delegated to a specific group; everyone needs to work on it. Also, some changes will not controlled by libraries, but involve more participants.
Chris Cole, NAL (working group member)
Q: What does carrier mean?
A: It means that we cannot modify MARC to be our future bibliographic record; we have to create something substantially different.
Q: What were comments?
A: The recommendation on RDA got the most comments. There were many comments about the statement that LC is not the national library. Also a considerable number of comments on the economics. Another set of comments was on "we're already doing that."
(added by Karen Coyle: there was a web site that gathered signatures asking for library data to be open, started by non-librarians. This shows that there are people outside of the library world who are interested in using library data.)
Q: There is also a similar economics question for RDA:
A: We need someone who knows economics to take a good look at this. Also, economics of standards development and maintenance hasn't been worked out. This is complex, and in the end we need to find ways to reduce costs.
Comment: The proposed work with Dublin Core pulls out the structure of a possible carrier, so this won't be part of RDA economic model.
Comment (UCLA): We shouldn't spend time perfecting FRBR.
Comment (Yale): RE: special collections and manuscripts – This was a good section, and we would have liked to have seen it go farther. There are large hidden collections, some printed (pamphlets). This is a cultural issue. We want to see more on priorities for LC and for all of us. E.g. LC provides expertise on non-western materials.
LC (Beacher): The is an LC internal group on the nature of bibliographic control at LC – it will now take on this report, analyze it and comment.
- The Bibliographic access group (beacher's) will look at the report
- public services area will react.
The last two will come together to report to Deanna and the five directors who report to her. By ala annual they should have a plan of action to share. The Library is pledging that each recommendation will be addressed and will get a statement and reaction to explain why it is accepted or not accepted by LC and how it will be carried out. Some are immediate, some are already underway, some will be longer term. LC has not done a good job of sharing with the community what it has been doing, and there are many projects underway. They don't have a timetable today. Mid-spring is the target for the first LC response, and they will have something to discuss at Annual.
Tuesday, January 08, 2008
More on RDA and "literals"
literal: an alphanumeric string representing a value.
example: "Moby Dick, or The whale"
example: "Herman Melville"
example: "ISBN:123456789x"
non-literal: a surrogate for the value itself.
example: uri:lccn: n 79006936 [identifies the LC name authorities record for Person Herman Melville]
example: http://authoritylists.info/uri/RDACarr/1052 [identifies RDA carrier type "volume"]
In programming, non-literals are those data elements that you give a name to. This means that in programming you are mainly working with non-literals:
mainTitle = 245ab
In data records, the values are often literals:
dc:title = "Moby Dick"
When we think about RDA and the literal/non-literal difference, we have some choices. RDA could treat every value as a literal, and let another standard, the data standard, define some as non-literals. This would provide for the maximum flexibility for implementing RDA.
Another possibility is that RDA could define some elements, such as the vocabulary lists included in RDA, as non-literals. These are, indeed, defined as non-literals in the RDA Element Analysis. This would mean that any use of RDA would need to define vocabularies for those data elements, and would have to assign those vocabularies and their entries with identifiers.
There are data elements that might be a literal in one implementation, and a non-literal in another. Author names are an example of this. Today we actually embed the author name as a literal in our MARC records; in the future we could use an identifier to link to an author record, as shown in the example above. Whether or not something is a literal is often a matter of implementation.
There are some data elements, or pieces of information that we think of as single data elements, that could be a combination of a literal and a non-literal. We already have this in the MARC record in the date element in the 008 field. The date itself is a literal ("1984"); it's just a string of numbers. Included with the date is a code that tells you what kind of date it is (single, range, copyright date). The same could be true of the extent statement: "345 p." could consist of a literal ("345") and a code for the unit (pages or leaves or volumes, etc.).
There are data elements that are hard to think of as a non-literal, and in fact they may never be one. Titles, explanatory notes, values like numbers of pages or dates -- all of these are likely to be simple text values in a data record.
Conclusion
RDA, as a set of cataloging rules, should not pre-determine whether elements are transcribed as literals or whether they are represented with surrogates for the values.
A step related to RDA in which RDA is defined as data elements that can be encoded for processing should allow literals for all data elements, but should be defined in such a way that non-literals could be used for any data element.
Another step, that encodes RDA as the library world's bibliographic record, should define non-literals for all vocabulary lists, and, where possible, for all units of measure or data element attributes (such as the type of publication date). It should also define optional non-literals for all authority-controlled elements. This would allow us to move increasingly in the direction of using non-literals.
Of course, our data elements themselves are (or should be) defined in such a way that they are identified with URIs, and therefore are non-literal values. This should be an obvious step in moving our data in the direction of the semantic web.
OK, I've stuck my neck out here -- all comments welcome!
Tuesday, December 18, 2007
Definitions in RDA Scope
The RDA scope document defines some basic concepts that presumably will be used throughout RDA. Some of these concepts it takes from the Dublin Core Abstract Model. In particular, it uses "literal value surrogate" and "non-literal value surrogate." These are defined in footnotes of the scope document as:
The term literal value surrogate is used as defined in the DCMI Abstract Model: “a value surrogate for a literal value, made up of exactly one value string (a literal that encodes the value)”.
The term non-literal value surrogate is used as defined in the DCMI Abstract Model: “a value surrogate for a non-literal value, made up of a property URI (a URI that identifies a property), zero or one value URI (a URI that identifies the non-literal value associated with the property), zero or one vocabulary encoding scheme URI (a URI that identifies the vocabulary encoding scheme of which the value is a member), zero or more value strings (literals that represent the value)”.
I found a more concise definition of this in a PPT by Lutz Maicher, University of Leipzig:
- a resource which is a non-literal value is represented by a proxy
- a resource which is a literal value is represented as literal
In the above, "literal" means a text string. So "Melville, Herman" is a literal, while "http://www.loc.gov/names/#n_79006936" is a non-literal proxy (because it points to the authority record, which is where the actual value is held).
The scope document then states:
- A label is represented by a literal value surrogate.
- A quantity is represented by a non-literal value surrogate
- A quality is represented by a non-literal value surrogate.
- A type is represented by a non-literal value surrogate
- A role is represented by a non-literal value surrogate.
However, in the element analysis in the scope document, it shows that quantities can be represented identically to labels (and I suspect that all other data types can as well). So that document has (and here there is a diagram that I cannot reproduce in email):
label
[resourceURIref] -> rda:title_proper -> [plain value string]
quantity
[resourceURIref] -> rda:extent -> [typed value string]^^[syntax encoding scheme]
- or -
[resourceURIref] -> rda:non_linear_scale -> [plain value string]
Given that the label example and the second example under quantity are structurally the same, I don't see how one can be a literal and one a non-literal.
I see two possibilities here. One is that all of the above has no real effect on the development of RDA, and therefore any errors in interpretation of the DCMI model can be ignored. The other is that the misunderstanding (which I think it is, but wait to be proven wrong) is significant, and therefore needs to be corrected as part of the development of RDA.
My gut feeling is that it is the former -- I don't see references to these definitions in the RDA text itself, and all values are treated as simple value strings. For example, dates are just text:
Record the date of the expression by giving the year or years alone.
1940 (p. 6-47 5rda_sec2349.pdf)
And quantities also seem to be just text strings as well:
46 slides
12 cm (from 5rda-parta-ch3rev.pdf)
Thus, at least as far as the RDA text is concerned, there are only literal values.
If this is not the case, would some please present the argument for a different understanding. Thank you.
Friday, December 07, 2007
Interpretations of FRBR Classes
This is an admirably short list of basic building blocks for bibliographic data. The question is: is it enough? Can we really express our bibliographic data with just these basic concepts? The answer is: probably not. Although we should take a lesson from FRBR and try to keep our set of basic entities small, while allowing for extension of them to express more complex concepts.
As an exercise, I took two well-known attempts to model FRBR using formal definitions. One is the FRBR in RDF, the other is FRBRoo. I also took the RDF entries that Martha Yee created for her cataloging rules and added those to the comparison although it is important to note that Yee's set of RDF statements is intended to go beyond FRBR since it is an expression of cataloging rules, not just the FRBR model.
In each of these three efforts, the FRBR entities are recorded as classes, and the FRBR relationships are recorded as properties. This is in keeping with the definitions in the RDF schema. What is interesting is the number of classes that are defined:
- FRBR in RDF: 13 classes
- FRBRoo: 23 classes, 18 sub-classes, 41 total
- Yee's schema: 23 classes
- FRBRoo does not include Manifestation, but instead has Manifestation product type and Manifestation singleton
- Yee's substitutes Event as subject for the FRBR class Event and substitutes Place as geographic area and Place as Jurisdictional Corporate Body for the FRBR Place
FRBR in RDF adds only three classes. Two of these (Endeavor and ResponsibleEntity) are supersets of FRBR classes. Endeavor is a generalization that can be related to a work, expression, or manifestation. Similarly, ResponsibleEntity is a more general term that can relate to either a corporate body or a person. Both of these seem fairly sensible, allowing you to refer to the intellectual content or some actor without having to specify more information. It's like being able to say "it" without having to saying exactly to what you are referring.
The third class that is added is Subject. As a matter of fact, all three of these include some instance of subjects as classes in their schemas. FRBR clearly treats subject as a relationship. (And I would like to understand why these three interpreted subject as a class -- so post if you have ideas/knowledge on that, please.)
FRBRoo
FRBRoo is a very interesting interpretation of FRBR. As they state in the document, attempting to re-define FRBR using object-oriented rules rather than entity-relationship rules is a way to test the underlying concepts in FRBR. They also tackle the elements that in FRBR that are called "attributes." (Aside: The FRBR attributes are a bit odd, IMO. They seem to be all over the place and there is no explanation of how they were determined or any way to give them some organization. I don't think they actually fit the definition of attributes in E-R, which seem instead to be on the order of identifiers). The folks working on FRBRoo decided to treat the attributes as properties, that is, relationships between the classes.
FRBRoo defines 23 primary classes with 18 subclasses. They address the issue of complex items, such as articles within serials or collections of essays, by creating classes for aggregate and serial works. Some of the classes seem to be what I would normally understand as genres. As an example, there is a class Performance Plan that is described as:
This class comprises sets of directions to which individual performances of theatrical, choreographic, or musical works and their combinations should conform.Another example of a new class is Publication Event. This is an action that is part of the work flow of publication, such as
Establishing in 1972 the layout, features, and prototype for the publication of “The complete poems of Stephen Crane, edited with an introduction by Joseph Katz” (ISBN “0-8014-9130-4”), which served for a second print run in 1978.Being an action, I would tend to express this as a property (a verb). So the layout, features, etc. could be subclasses of a manifestation, there would be an actor (a noun, or a class, probably the publishing house, or more specifically a book designer), and a time. The verb (or property) could be "designed" "typeset" "printed" etc. This makes me wonder about the FRBR class Event as a noun, but I think I could buy into a concept of named events ("WWII" "Election day 2008" "Beatles first appearance on Ed Sullivan"). Interestingly, it does appear that all of these are events as subjects, as Event is defined in FRBR; the FRBRoo event does not appear to have this noun-ish characteristic.
Yee Schema
Martha Yee's set of classes (23 of them, but not the same 23 as FRBRoo) includes Genre/Form as a class. Genre/form seems to be more of an attribute about a work rather than something that has "thingness" in itself. It's hard to imagine how you can have genre/form without it relating to a work. (As opposed to: you can have a person or a corporate body that are things in and of themselves -- that have specific, unique identities.)
It has some classes that might be considered sub-classes. For examples, Place as geographical area and Place as jurisdictional corporate body would seem to be sub-classes of Place, although Yee does not include Place itself in her schema. I'm less clear about classes such as Corporate Subdivision, which has a part/whole relationship with Corporate Body, not a sub-class relationship. (Sub-class would be an "is a type of" relationship, and corporate subdivision is not a type of corporate body, it's a part of a corporate body.) Ditto the subject-related terms: Subject, Subject subdivision, Subject chronological subdivision, Subject form subdivision, Subject geographical subdivision, Subject topical subdivision. In FRBR, the subject is a relationship with the work. These look to me to be relationships with the subject heading, although there is no class for subject headings (unless that is what is meant by the class Subject, but I don't think it would be a good idea to equate subject with subject heading because it makes it impossible to include classifications as subjects or keywords as subjects).
What's the upshot? Well, it would take a good sit-down with all involved to hash out the differences, to understand what each group or person was thinking, and to see if we can formulate a theory of how one extends FRBR to meet ones needs. If a number of people turn out to have the same needs, then it may be that the FRBR model itself needs to take in those ideas. The only way to work this out is to keep modeling and sharing. So I thank the three featured here for the extensive work that they have done in this area.