Notes from the meeting on Nov. 13 of the Working Group on the Future of Bibliographic Control.
These are my notes and should NOT be taken to represent accurately the thoughts of the working group, only my quick recording of what I understood at the meeting. Also, I must add the disclaimer that I have been engaged as a consultant to the group for the writing of the report. I attempt in that work to be as faithful to the outcomes desired of the group as I can. However, I admit that pure objectivity is a chimera, so my own opinions may come through in the text below.
There was an introduction explaining the creation of the working group (which you can read about on the working group's web site: http://www.loc.gov/bibliographic-future/). The group presented an interim report to the Library of Congress. The full report will be available by December 1 for public comment. The comment period will end on December 15, and the final report will be presented on January 8, 2008.
The report was commissioned by the Library of Congress, but it many of its recommendations involve the the library community and other players in its environment. There are over 100 individual recommendations in five general areas.
The working group concluded that there are three major "sea changes" that are needed in the library community:
1. We must redefine bibliographic control broadly to include all materials, a widely diverse community of users, and a multiplicity of venues where information is sought.
2. We must redefine the bibliographic universe to include all stakeholders, including the for-profit organizations that are involved in information delivery and digitization
3. The role of the Library of Congress must be redefined as a partner with other libraries and with non-library institutions, working to achieve the goals of the library community.
The five areas of recommendations are:
1. Increase the efficiency of bibliographic production for all libraries through cooperation and sharing of bibliographic records and through the use of data produced in the overall supply chain.
2. Transfer effort into high value activity. In particular, provide greater value for knowledge creation through leveraging access for unique materials held by libraries, materials that are currently hidden and under-used.
3. Position our technology by recognizing that the Web is our technology platform as well as the appropriate platform for our standards. Recognize that our users are not only people but also applications that interact with library data.
4. Position our community for the future by adding evaluative, qualitative and quantitative analyses of resources. Work to realize the potential provided by the FRBR framework.
5. Strengthen the library and information science profession through education and through development of metrics that will inform decision-making now and in the future.
Under each of these areas there are sets of recommendations. The full set of recommendations is fairly detailed, and the group presented high level groupings of recommendations in the first four areas. (Area five was not presented in detail at the meeting.)
In area 1, the recommendations are grouped:
1.1 Eliminate redundancies in the production of bibliographic metadata. This means making use of data that is created elsewhere in the supply chain, and increasing the sharing of bibliographic records and modifications to records. In particular, the group asks for an examination of barriers to sharing.
1.2 Increase the distribution of responsiblity for bibliographic record production. Increase the number of institutions that participate in shared cataloging activities.
1.3 Collaborate on authority record creation. Similar to 1.2, this recommends that the number of participants in authority record creation be increased, but it also asks that we look at the possibility of sharing across sectors and internationally, to reduce the number of times that an authoritative heading must be created.
Area 2 is called "Enhance Access to Rare and Unique Materials." In this area the group states that any efficiencies gained in other areas should allow the redirection of energy to providing access to unique materials that are held by libraries and other cultural heritage institutions. In particular, the group recommends:
2.2 Integrate access to rare & unique materials with other library materials
2.3. Share bibiographic data relating to these materials. The sharing of bibliographic data must not be limited to those areas where copy cataloging is desired.
2.4 Encourage digitization to allow broad access
Area 3 is about technology and the Web:
3.1.1 Integrate library standards into the Web environment
3.1.2 Extend the use of standard identifiers for bibliographic entities, and include those identifiers in bibliographic records.
3.1.3 Develop a more flexible, extensible metadata carrier that can be readily exchanged with non-library applications.
Area 3 also addresses standards:
3.2.1 Develop standards with a focus on return on investment. Do analysis before beginning the standards process.
3.2.2 Incorporate usage data and lessons from use tests in the standards development process
Area 4 is about positioning the library community toward a more progressive future. In this area there are three main recommendation areas:
4.1 Design for today's and tomorrow's user. This means that we must design into our catalogs and other tools the ability to present evaluative information, and to allow and encourage users to interact with bibliographic data. We must also make use of statistical and other computationally-derived information in our user services.
4.2 Realize FRBR. The framework known as FRBR has great potential but so far is untested. It is being used as the basis for RDA, even though FRBR itself is not clearly understood. The working group recommends that no further work be done on RDA until there has been more investigation of FRBR and the basis it provides for bibliographic metadata. [Note: this recommendation is likely to change such that there will be specific recommendations relating to RDA; FRBR will be treated separately.]
4.3 Optimize LCSH for Use and Re-use. Encourage an analysis of LCSH that would move the system toward a more facetted subject system. Work to create more links between LCSH and other subject heading systems in use. Recognize that with the digitization of works the act of subject assignment may benefit from computational analysis.
In the time that I was at the meeting (I had to leave before the question period ended) there were two questions/comments. The first had to do with the fact that while there are costs to today's methods of bibliographic control, that changes in bibliographic control will have costs as well. (Here it would be good to listen again to the talk given by Rick Lugg at the meeting held at LoC. He spoke of the costs of NOT changing, something that is hard to measure but is very real.) The other comment (from Barbara Tillett) mentioned many of the recommendations and stated that LoC is already engaged in, or has rejected, analogous activities. It was acknowledged, however, that LoC had not made these activities public, so the community is generally unaware of the progress made. To me this points out one of the areas that we all need to work on, which is sharing information about our projects and their progress so that the community as a whole can benefit from work done by a single institution.
Thursday, November 15, 2007
Wednesday, November 07, 2007
Hierarchy v. Relationships
The use of hierarchy as an organizing principle keeps coming up. I think we are attracted to hierarchy because of its neatness, even though in fact the real world is organized more like fuzzy sets. Fuzzy sets are hard to comprehend, nearly impossible to draw, and can't be slotted neatly into an application.
When people talk about FRBR, they are often focussed on the Group 1 entities, and those are seen as hierarchical. They tend to be shown as:
- Work
-- Expression
--- Manifestation
---- Item
as if we'll fit all of our intellectual works into such a neat hierarchy. T'ain't so. Of all of the relationships that are talked about in FRBR (I almost said "expressed" but that term has now been given a new meaning in this discussion) I think these are the least interesting. And they become even less interesting when we move beyond the traditional inventory control function of the library catalog and begin to see ourselves as navigating in a knowledge universe. But first let me tackle the Group 1 entities.
There are complaints (or remarks, depending on the context) that we don't have an agreed on definition of Work, and that the division between Work and Expression is unclear. They are unclear because in real life there isn't a neat hierarchy that just needs to be modeled. What is a Work is entirely contextual -- when I'm looking for an article, the article is a Work. When I'm subscribing to a journal, the journal is a work. When I'm on iTunes a song is a work, when I'm in the music store the album is a work. A Work is the content I am seeking at that time. In the imaginary universe where I get to create my bibliographic system, a Work will be defined as: anything you wish to talk about, point to, address. So a book-length text is a work, an article in a journal is a work, a journal is a work, a book chapter is a work -- all at the same time and in the same system. For one person the book Wizard of Oz and the movie Wizard of Oz will be a single Work. To a film buff, the director's cut of Blade Runner and the original release are distinct works. To be a Work, it just has to be definable and have a way to name it, that is it has to have an identifier. But anything can be a Work. As a matter of fact, I probably won't use the term Work at all in my universe.
As for Expressions, there will be very obvious Expressions of Works, and there will be fuzzier Expressions. There will be Expressions that express more than one Work. Expression is a relationship, not a subset. If you don't have to organize your bibliographic universe in a hierarchical way, then the need to slot each Expression under a Work goes away, although the relationship can remain.
I'm less sure about Manifestation and Item, even though these are the most concrete of the Group 1 entities. Are they a legitimate focus of a Knowledge Management system, or are they about managing physical objects? When I think about some of the uses of bibliographic data, for instance as citations in a text publication, Manifestation seems to be mainly about locating -- so if I've quoted a passage from a book, I need to cite the manifestation and the page because that's the only way that someone else can find that exact quote. When I include a URL in a document that links to a particular digital manifestation, I am giving the user a direct link to the location. Manifestations and Items will be of interest in some instances, say to rare book collectors, but I'm not at all sure that those instances justify the emphasis they have been given. And if the purpose is primarily inventory control, then I think those relationships will be managed to the extent that they matter to the library. For example, a public library may not terribly care which manifestation of the book Moby Dick is on its shelves, although its inventory system will need to know the barcode, and its acquisitions system will need to store how much the library paid for it and the provider.
The truly interesting relationships in FRBR are those between and among these entities, and those are ones that I have not seen explored. These are the relationships between things: thing1 is a translation of thing2; thing3 is an abridgment of thing4; thing5 extends thing6 in this certain way; thing7 cites thing1; thing8 continues thing3. This is where we get real value, where we provide various interesting paths through which seekers can navigate. This is what we don't provide explicitly in our catalogs today, although a human user may be able to intuit some of these relationships among the works we present.
We have so narrowly defined bibliographic control in libraries that it doesn't really include the relationships between intellectual products, except to the degree that we might make a note that one thing is a translation of another thing. But we see those relationships as "extra" or "secondary," and yet they are the very essence of knowledge creation. It astonishes me that we have focused so completely on the physical items that we have essentially missed what would make our catalogs intelligent.
When people talk about FRBR, they are often focussed on the Group 1 entities, and those are seen as hierarchical. They tend to be shown as:
- Work
-- Expression
--- Manifestation
---- Item
as if we'll fit all of our intellectual works into such a neat hierarchy. T'ain't so. Of all of the relationships that are talked about in FRBR (I almost said "expressed" but that term has now been given a new meaning in this discussion) I think these are the least interesting. And they become even less interesting when we move beyond the traditional inventory control function of the library catalog and begin to see ourselves as navigating in a knowledge universe. But first let me tackle the Group 1 entities.
There are complaints (or remarks, depending on the context) that we don't have an agreed on definition of Work, and that the division between Work and Expression is unclear. They are unclear because in real life there isn't a neat hierarchy that just needs to be modeled. What is a Work is entirely contextual -- when I'm looking for an article, the article is a Work. When I'm subscribing to a journal, the journal is a work. When I'm on iTunes a song is a work, when I'm in the music store the album is a work. A Work is the content I am seeking at that time. In the imaginary universe where I get to create my bibliographic system, a Work will be defined as: anything you wish to talk about, point to, address. So a book-length text is a work, an article in a journal is a work, a journal is a work, a book chapter is a work -- all at the same time and in the same system. For one person the book Wizard of Oz and the movie Wizard of Oz will be a single Work. To a film buff, the director's cut of Blade Runner and the original release are distinct works. To be a Work, it just has to be definable and have a way to name it, that is it has to have an identifier. But anything can be a Work. As a matter of fact, I probably won't use the term Work at all in my universe.
As for Expressions, there will be very obvious Expressions of Works, and there will be fuzzier Expressions. There will be Expressions that express more than one Work. Expression is a relationship, not a subset. If you don't have to organize your bibliographic universe in a hierarchical way, then the need to slot each Expression under a Work goes away, although the relationship can remain.
I'm less sure about Manifestation and Item, even though these are the most concrete of the Group 1 entities. Are they a legitimate focus of a Knowledge Management system, or are they about managing physical objects? When I think about some of the uses of bibliographic data, for instance as citations in a text publication, Manifestation seems to be mainly about locating -- so if I've quoted a passage from a book, I need to cite the manifestation and the page because that's the only way that someone else can find that exact quote. When I include a URL in a document that links to a particular digital manifestation, I am giving the user a direct link to the location. Manifestations and Items will be of interest in some instances, say to rare book collectors, but I'm not at all sure that those instances justify the emphasis they have been given. And if the purpose is primarily inventory control, then I think those relationships will be managed to the extent that they matter to the library. For example, a public library may not terribly care which manifestation of the book Moby Dick is on its shelves, although its inventory system will need to know the barcode, and its acquisitions system will need to store how much the library paid for it and the provider.
The truly interesting relationships in FRBR are those between and among these entities, and those are ones that I have not seen explored. These are the relationships between things: thing1 is a translation of thing2; thing3 is an abridgment of thing4; thing5 extends thing6 in this certain way; thing7 cites thing1; thing8 continues thing3. This is where we get real value, where we provide various interesting paths through which seekers can navigate. This is what we don't provide explicitly in our catalogs today, although a human user may be able to intuit some of these relationships among the works we present.
We have so narrowly defined bibliographic control in libraries that it doesn't really include the relationships between intellectual products, except to the degree that we might make a note that one thing is a translation of another thing. But we see those relationships as "extra" or "secondary," and yet they are the very essence of knowledge creation. It astonishes me that we have focused so completely on the physical items that we have essentially missed what would make our catalogs intelligent.
Sunday, November 04, 2007
Our subject mess
Lately I've had occasion to work with a few different groups of people who are delving into library bibliographic data for the first time. Believe me, it is quite revealing to view it from the viewpoint of these novices. Novices only in this one area, because they generally are quite savvy about computing and data. Each new revelation gives me a chance to regale them with an amusing story about "how it got that way." I can explain (note: explain, not justify) why we have no identifiers for key elements like authors and works. I can pretty much explain why we seem more concerned about the package than the content. I can reminisce about moments in the history of library systems development that happened before some members of these groups were born. But I get totally stuck when they point out the mess that is our subject access.
We have two classification systems, Dewey (DDC) and Library of Congress. (LCC) That in itself is not a problem, and it's fairly easy to explain how they developed in different contexts, always making sure to explain that these systems classify the items in a library, not the world of thought.
What is hard is to try to explain what either of them has to do with the Library of Congress Subject Headings.(LCSH) Many folks assume that LCSH is the entry vocabulary into LCC. Thus if there is a classification code in a record that stands for "vocal music, choruses" that there will be a heading in the record that is "vocal music, choruses," and vice versa. They also assume that the two subject systems (classification and subject headings) have the same structure, which would mean that you can "drill down" from music to vocal music then to choruses in either or both. Nothing could be further from the truth. So it is quite confusing to them when they see a record with a call number that would ostensibly be about "vocal music, choruses" based on the classification, but instead the subject heading is "Cantatas, Secular -- Scores." And they are equally confused when the record has another subject heading ("Funeral music") but only the one classification number.
I can't explain this disconnect between the subject headings and the classification scheme, except to say: that's how it is.
Recently, I was browsing through my beloved copy of the DDC from 1899 that still has both its numeric and alphabetical tabs relating respectively to the classification and the "Relativ Subject Index." The RSI is indeed an index to the classification scheme, and it appears that Dewey originally intended it also as the access to the collection:
Meanwhile, no wonder users are confused.
We have two classification systems, Dewey (DDC) and Library of Congress. (LCC) That in itself is not a problem, and it's fairly easy to explain how they developed in different contexts, always making sure to explain that these systems classify the items in a library, not the world of thought.
What is hard is to try to explain what either of them has to do with the Library of Congress Subject Headings.(LCSH) Many folks assume that LCSH is the entry vocabulary into LCC. Thus if there is a classification code in a record that stands for "vocal music, choruses" that there will be a heading in the record that is "vocal music, choruses," and vice versa. They also assume that the two subject systems (classification and subject headings) have the same structure, which would mean that you can "drill down" from music to vocal music then to choruses in either or both. Nothing could be further from the truth. So it is quite confusing to them when they see a record with a call number that would ostensibly be about "vocal music, choruses" based on the classification, but instead the subject heading is "Cantatas, Secular -- Scores." And they are equally confused when the record has another subject heading ("Funeral music") but only the one classification number.
I can't explain this disconnect between the subject headings and the classification scheme, except to say: that's how it is.
Recently, I was browsing through my beloved copy of the DDC from 1899 that still has both its numeric and alphabetical tabs relating respectively to the classification and the "Relativ Subject Index." The RSI is indeed an index to the classification scheme, and it appears that Dewey originally intended it also as the access to the collection:
"HOW TO USE THIS INDEXFrom this I can only presume that the shelves and the subject catalog were in classification order, and the alphabetical index was the index to that classification. I can only guess at this point, from what he says here, that the subject catalog was in classification order, as is the shelf, but also contained the verbal translation of what the decimal classification numbers meant.
Find the subject desired in its alphabetical place in the index. The number after it is its class number and refers to the place where the topic will be found, in numerical order of class numbers, on the shelves or in the subject catalog."
"Under this class number will be found the resources of the library on the subject desired. Other subjects near the one sought may often be consulted with profit; e.g., Communism is the topic wanted and the index refers to 335.4, but 335, Socialism, and even the inclusive division 330, Political economy, also contain much on this subject. The reverse is equally true; the full material on socialism can only be had by looking at its divisions 335.3, Fourierism, 335.4, Communism, etc. The topics which are thus subdivided are plainly marked in the index by heavy faced type."My copy is #3933, originally owned by the Roger Williams Park Museum in Providence, Rhode Island. The current incarnation of the institution appears to be the Museum of Natural History and Planetarium. My copy has many penciled notes in the area of Zoology (DDC 590), which would fit the natural history nature of the institution. (I don't see any evidence of a current library.) By 1900 the "dictionary catalog" would have taken root, so I don't know if the library would have followed Dewey's instructions for the creation of a classified catalog. But I do wonder how we got from a single system that had an alphabetical index to a classification system to a system with an alphabetical index and two classification systems, but in which the index and the classification have essentially each gone their own ways. This is obviously a gap in my education, which I will gladly rectify if you have suggestions for readings.
Meanwhile, no wonder users are confused.
Sunday, October 28, 2007
Bibliographic ER
No, I'm not sending libraries to the emergency room, although there are days when I feel like we're at that point. The ER in the title refers to Entity-Relationship, a way to look at data that emphasizes the general viewpoint that there are things, and those things exist in relation to each other.
In one sense, this is what we have done for over a century with our library data. The bibliographic records that we create have in them many relationships: Person authored Book; Publishing House published Book; Book is in Series; Book has Topics. Those relationships are implicit in our records, but the data isn't formatted in an entity-relationship model. Our records, instead, talk about the relationships but don't make it easy to give the various entities their own existence. So we create a record that contains:
Author
Book title.
Place, publisher, date
Series
Subject A
Subject B
The record represents all of the information about the book, but there is no record that represents all of the information about the author, or all of the information about the publisher, etc. Instead, those "entities" are buried in bibliographic records scattered throughout the file.
An E-R model would give each of these entities an identity on which you could hang information about the entity.

OK, I can't draw worth beans. But basically the idea is that authors, subjects, publishers, topics, all become entries in their own right. This means that you can add information to the author record or the series record, because they have their own place in the design. It also makes it easy to look at your data from many different points of view, while still retaining all of the richness of the relationships. So from the point of view of the person who is the illustrator in the book above, the bibliographic world may look like this:

This type of model is expressed in FRBR, but the E-R aspect of FRBR does not seem to be incorporated into RDA as it stands today. Instead, RDA appears to be aimed at creating the same flat structure that we have in library data today.
If you take a look at the OpenLibrary you will see that books get a page that is about the book, and authors get a separate page that is about the author. This is very simple, but it is also very important. It means that the catalog is no longer just a list of books with authors but can become a rich source of information about authors. You can add bios for authors, link to web sites about the author, launch a discussion group about a favorite author. Because the author is an entity, not just a data element in a record about the book, it becomes a potentially active part of your information system.
In the future, I hope that we can give life to many more entities in the OpenLibrary, and also that we can give them meaningful relationships between each other. This would mean taking a semantic web approach to library data. I don't have a clear picture of where we'll end up, but I'm glad that folks there are interested in experimenting. If you've already thought this through or have ideas in this direction, please step forward. I'd love to hear from you.
In one sense, this is what we have done for over a century with our library data. The bibliographic records that we create have in them many relationships: Person authored Book; Publishing House published Book; Book is in Series; Book has Topics. Those relationships are implicit in our records, but the data isn't formatted in an entity-relationship model. Our records, instead, talk about the relationships but don't make it easy to give the various entities their own existence. So we create a record that contains:
Author
Book title.
Place, publisher, date
Series
Subject A
Subject B
The record represents all of the information about the book, but there is no record that represents all of the information about the author, or all of the information about the publisher, etc. Instead, those "entities" are buried in bibliographic records scattered throughout the file.
An E-R model would give each of these entities an identity on which you could hang information about the entity.
OK, I can't draw worth beans. But basically the idea is that authors, subjects, publishers, topics, all become entries in their own right. This means that you can add information to the author record or the series record, because they have their own place in the design. It also makes it easy to look at your data from many different points of view, while still retaining all of the richness of the relationships. So from the point of view of the person who is the illustrator in the book above, the bibliographic world may look like this:
This type of model is expressed in FRBR, but the E-R aspect of FRBR does not seem to be incorporated into RDA as it stands today. Instead, RDA appears to be aimed at creating the same flat structure that we have in library data today.
If you take a look at the OpenLibrary you will see that books get a page that is about the book, and authors get a separate page that is about the author. This is very simple, but it is also very important. It means that the catalog is no longer just a list of books with authors but can become a rich source of information about authors. You can add bios for authors, link to web sites about the author, launch a discussion group about a favorite author. Because the author is an entity, not just a data element in a record about the book, it becomes a potentially active part of your information system.
In the future, I hope that we can give life to many more entities in the OpenLibrary, and also that we can give them meaningful relationships between each other. This would mean taking a semantic web approach to library data. I don't have a clear picture of where we'll end up, but I'm glad that folks there are interested in experimenting. If you've already thought this through or have ideas in this direction, please step forward. I'd love to hear from you.
Labels:
FRBR,
library catalogs,
OpenLibrary,
RDA,
semantic web
Saturday, October 20, 2007
Great Minds...
As if in response to my post on name authorities, OCLC has come up with a version of the Virtual International Authority File (acronym VIAF). Type in "Fitzgerald, Michael" and you'll see that each name has associated with it what they are calling a "sample title." The titles are unattractive, being normalized forms, but still give you some idea of what each author has written, and you might be able to sort the Michael Fitzgerald who writes on XSL from the one who has written the guide to better business letters. At this point, that authority control has already determined that these are different people is incredibly valuable, where the value was much harder to see when all you had were names and dates.
Friday, October 12, 2007
Cataloging as Industry
Something pointed me to this paper by Alan Danskin of the British Library:
Tomorrow never knows: the end of cataloging?
It has some well-spoken statements about the great increase in materials, the need to collaborate better with others in the publishing supply chain, etc. But what really stood out for me was this:
He qualifies this by saying
I can't disagree with what he says here, but I must say that I have a different take on the idea of industrialization of cataloging, and that is that we should consider taking cataloging out of the library and giving it to others who will actually industrialize it. Just as we don't hand craft our own library shelves, and we don't hand craft our own library systems, perhaps we shouldn't be hand-crafting our own catalog records.
What I refer to here would probably come under the rubric of "outsourcing," some of which already takes place, especially for works in less common or more difficult languages. But what if, just what if, someone could develop a cataloging service that was cheaper than what libraries can do themselves, and had comparable quality? Is there any reason why we shouldn't go for it?
Tomorrow never knows: the end of cataloging?
It has some well-spoken statements about the great increase in materials, the need to collaborate better with others in the publishing supply chain, etc. But what really stood out for me was this:
The future of cataloguing depends on transforming the process from a craft into an industry.
He qualifies this by saying
This requires unambiguous identification at different levels of granularity to facilitate repurposing of metadata created at the different stages of the process of creating and publishing resources. It also means we may have to be less precious about some of our cherished practices.
I can't disagree with what he says here, but I must say that I have a different take on the idea of industrialization of cataloging, and that is that we should consider taking cataloging out of the library and giving it to others who will actually industrialize it. Just as we don't hand craft our own library shelves, and we don't hand craft our own library systems, perhaps we shouldn't be hand-crafting our own catalog records.
What I refer to here would probably come under the rubric of "outsourcing," some of which already takes place, especially for works in less common or more difficult languages. But what if, just what if, someone could develop a cataloging service that was cheaper than what libraries can do themselves, and had comparable quality? Is there any reason why we shouldn't go for it?
Sunday, September 30, 2007
Glut? Gunk!
You've probably had the experience of participating in some activity that was later covered by print or TV news. In many cases, the report of the event is so wrong, so different to what you experienced, that you could hardly recognize it as being the same event. Similarly, when reporters write about something you know intimately, the reports are almost always aggravatingly wrong.
The same is true about books, of course. I thoroughly enjoyed Bill Bryson's A Short History of Nearly Everything, which drove real scientists nuts for everything it got wrong. Now I'm going out of my mind reading Alex Wright's Glut, which I can only describe as poorly researched, and in some cases just outright wrong.
I became suspicious when I read on page 21
OK, we all can slip up when we get going at the keyboard, and I figured that his editors just hadn't paid attention. Then I got to page 79 where he says:
I have kept reading, I guess because I wanted to get to his treatment of more modern times. I've gotten as far as Panizzi, but had to get all of this out of my system before going on. On page 167, Wright quotes a biographer, one Louis Fagan, on Panizzi's appearance. I looked at the citation for the quote and found:
It's not an important point nor a particularly important passage, but it is sloppy scholarship. It means he took his information from someone else and did not verify the original source. In fact, of the about 260 citations in the book (and I'm counting all of the "ibid's" in this) a full 52 are "quoted in" or "cited by," and mainly the former. The entire first half of the book, which is on ancient and medieval history, uses modern sources almost exclusively. One chapter, on memory, cites only six discrete works, and takes quotes of Thomas Aquinas, Giulio Camillo, John Willis, John Wilkins, and Francis Bacon second-hand from books published mainly in the 1990's. In that chapter, only one "ancient" quote is from an original source. One of the citations referring to Wilkins is to a BBC web site page. It's no longer available. I might be just being mean, but I can find the BBC page cited on the Wikipedia entry for John Wilkins in the Wikipedia version prior to the date of Wright's citation, although it has since been removed. I don't at all mind people using Wikipedia for its basic purpose: to give one a clue and lead one on to sources. And of course we all jump on to the nearest bit of information on the web. But when researching a well-known historical figure, it really is important to cite a good, permanent resource, and in terms of Wilkin, other resources should be available.
As for Panizzi, Wright talks about his creation of a schedule of tiered subject headings. On page 168 he has a quote from Elaine Svenonius that implies some criticism of Panizzi's work.
Must I go on? I was able to check this one reference carefully because I happened to have the Svenonius book on my own bookshelf. I have no reason to believe that the rest of his text is any more accurate or faithful to the sources he cites. I suppose the one consolation is that in spite of his MLS from Simmons, Alex Wright calls himself an Information Architect, eschewing the "L" word. I wouldn't want people to think that librarians don't know how to do research.
The same is true about books, of course. I thoroughly enjoyed Bill Bryson's A Short History of Nearly Everything, which drove real scientists nuts for everything it got wrong. Now I'm going out of my mind reading Alex Wright's Glut, which I can only describe as poorly researched, and in some cases just outright wrong.
I became suspicious when I read on page 21
"It is no coincidence that snakes have been a leading cause of human mortality throughout our species' history, so it should come as no surprise that the occurrence of serpent imagery tracks closely to the prevalence of poisonous snakes in particular regions."I don't doubt that snakes are scary creatures and they sure do seem to show up in all kinds of ancient imagery and tales, but "a leading cause of human mortality"? I don't think so. Famine, pestilence, war -- those are leading causes of human mortality. Snakes? A drop in the bucket.
OK, we all can slip up when we get going at the keyboard, and I figured that his editors just hadn't paid attention. Then I got to page 79 where he says:
"... a new form of document: the codex book, so named because it originated from attempts to 'codify' the Roman law in a format that supported easier information retrieval."Codex comes from "codify"? Were the Romans speaking English? And besides, I'd recently read a few books on book history myself and those all referred to that origin as being from the Latin term "caudex" referring to wood used as the first book covers. The use of "code" for groups of laws came from the term "codex," not vice versa. I began to wonder where he would have gotten such a definition, and on a hunch decided to look at the Wikipedia entry on Codex. There had been some confusion between codex and code in an early Wikipedia version of the codex page, and it was removed:
"Mistaking Codex for CodeI have no idea if that is where Wright got his information, but this statement makes the same mistake that Wright does.I moved this mis-stated misunderstanding here: "A legal text or code of conduct is sometimes called a codex (for example, the Justinian Codex), since laws were recorded in large codices." This is simply an error, one that doesn't come into educated or official discourse. --Wetman 20:14, 9 May 2006 (UTC)
I have kept reading, I guess because I wanted to get to his treatment of more modern times. I've gotten as far as Panizzi, but had to get all of this out of my system before going on. On page 167, Wright quotes a biographer, one Louis Fagan, on Panizzi's appearance. I looked at the citation for the quote and found:
"3. Louis Fagan, quoted in Teresa Negrucci, 'Historiography of Antonio Panizzi,' 2001, http://www.gseis.ucla.edu/faculty/maack/Panizzi.doc"I looked up the paper online, and Ms. Negrucci was a student in the UCLA library school at the time of writing this paper, done for IS 281 "Historical Methodology for Library and Information Science." (The citation above is no longer valid. You can find it linked from this page of student writings.) A perfectly fine school paper, but probably not an authoritative source. Plus, I was taught that you only took quotes from someone else if the original is terribly hard to get to. Fagan's book is available in at least 80 US libraries, according to WorldCat, although today I was able to get to it online. Now, I admit that the book may not have been available via Google Book Search when Wright was composing his work, but by no means is the original inaccessible. In fact, if he had looked at the original, rather than the student paper, he would have understood that Fagan was quoting someone else in his description of Panizzi, not making the statement himself, as Wright states.
It's not an important point nor a particularly important passage, but it is sloppy scholarship. It means he took his information from someone else and did not verify the original source. In fact, of the about 260 citations in the book (and I'm counting all of the "ibid's" in this) a full 52 are "quoted in" or "cited by," and mainly the former. The entire first half of the book, which is on ancient and medieval history, uses modern sources almost exclusively. One chapter, on memory, cites only six discrete works, and takes quotes of Thomas Aquinas, Giulio Camillo, John Willis, John Wilkins, and Francis Bacon second-hand from books published mainly in the 1990's. In that chapter, only one "ancient" quote is from an original source. One of the citations referring to Wilkins is to a BBC web site page. It's no longer available. I might be just being mean, but I can find the BBC page cited on the Wikipedia entry for John Wilkins in the Wikipedia version prior to the date of Wright's citation, although it has since been removed. I don't at all mind people using Wikipedia for its basic purpose: to give one a clue and lead one on to sources. And of course we all jump on to the nearest bit of information on the web. But when researching a well-known historical figure, it really is important to cite a good, permanent resource, and in terms of Wilkin, other resources should be available.
As for Panizzi, Wright talks about his creation of a schedule of tiered subject headings. On page 168 he has a quote from Elaine Svenonius that implies some criticism of Panizzi's work.
"Some would argue [the subject headings] were too ambitious -- that there was no need to construct elaborate Victorian edifices since jerrybuilt systems could meet the needs of most users most of the time."The bracketed words "the subject headings" was added by Wright. In fact, Svenonius was not referring to Panizzi's headings. The quoted passage is about "systems produced during the second half of the nineteenth century," ("Victorian" should be a hint) which would be after Panizzi, whose primary work was done earlier in that century. And the full quote, with no reference to subject headings, is:
"The systems produced during the second half of the nineteenth century, a period regarded as a golden age of organizational activity, [cites Cutter 1904] were ambitious, full-featured systems designed to meet the needs of the most demanding users. Some would argue that they were too ambitious -- that there was no need to construct elaborate Victorian edifices since jerrybuilt systems could meet the needs of most users most of the time. [cites Coffman]" Svenonius, p. 3The Cutter reference is to his 4th edition of Rules for a Dictionary Catalog. The sentence quoted by Wright is a reference to American Libraries article by Steve Coffman called "What If You Ran Your Library Like a Bookstore?".
Must I go on? I was able to check this one reference carefully because I happened to have the Svenonius book on my own bookshelf. I have no reason to believe that the rest of his text is any more accurate or faithful to the sources he cites. I suppose the one consolation is that in spite of his MLS from Simmons, Alex Wright calls himself an Information Architect, eschewing the "L" word. I wouldn't want people to think that librarians don't know how to do research.
Subscribe to:
Posts (Atom)