Wednesday, January 31, 2007

Legislating technology - RFID and California SB 30

The California state Senate will be considering a bill that would regulate the use of RFID in identity documents. The bill is over 5700 words long, so it's a bit hard to address its details (and boy, does it have details!) in a short blog post. so I'll just make a few provocative statements and encourage you to read the text for yourself.

Too Much, Too Soon

It is almost always a very bad idea to legislate specific technology, especially when the technology is in its infancy. As far as I know, there have been no privacy breeches related to the use of RFID in any form, much less on identity documents. This legislation is 5700 words and many complex do's and don'ts for something that may not even be an issue five years from now. To its credit, it has a sunset date of December 31, 2013, allowing for change over that time period. Still, I regard this legislation as being too much, too soon.

Background

The legislation is at least partly fueled by work of the Electronic Frontier Foundation on RFID. EFF's view is that RFID is the camel's nose under the tent for a ubiquitous surveillance society. The EFF and the local ACLU argued against the implementation of RFID for materials in the San Francisco Public Library. At this moment, funding to purchase an RFID system by SFPL has been removed by the city council.

What the Bill Says

Having read the bill, I can say that its content would be a very interesting "best practices" document for the use of RFID for identification documents. Of course, such a document would not be written as law, and therefore might be easier to digest. Basically, the bill sets some very strict rules for the implementation of government-issued identity documents that can be read "remotely." These documents would be given much greater scrutiny than, for example, those with visible data or those with data in a magnetic strip. All cards using RF must meet these two standards:

(1) In order to prevent duplication, forgery, or cloning of the identification document, the identification document shall incorporate tamper-resistant features.
(2) In order to determine to a reasonable certainty that the identification document was legitimately issued by the issuing entity, is not cloned, and is authorized to be read, the identification document and authorized reader, in conjunction with related, functionally integrated software, shall implement an authentication process.
I wish I could tell you what this means for someone wishing to implement an identity card, but I'm afraid I don't know. Is this a level of complexity that adds a great deal to the cost? To what extent do these features create vendor lock-in, making it hard for an agency to move to a different vendor's platform? I have lots of questions of this nature throughout the document, and hope that some explanatory documentation will be provided at some point.

The technical details in the document are fairly specific, such as:

(3) If personally identifiable information is transmitted remotely from the identification document, the identification document and authorized reader, in conjunction with related, functionally integrated software, shall not only meet the requirements of paragraph (2) but also shall implement mutual authentication in order to prevent the transmission of personally identifiable information between identification documents and unauthorized readers.

(4) If personally identifiable information is transmitted remotely from the identification document, the identification document shall make the data unreadable and unusable by an unauthorized person through means such as encryption of the data during transmission, access controls, data association, encoding, obfuscation, or any other measures, or combination of measures, that are effective to ensure the confidentiality of the data transmitted between the identification document and authorized reader.

Like most complex legislation, you have to keep track of the "subdivision (a) shall not apply to" or the "except as provided in subdivision (b)", which is why I wish this had been written as a best practices document before getting turned into legislation-eze. However, if I read the document correctly, at one point it suggests that the holder of the document must be able to exercise control over whether the data is transmitted or not, and this requires some kind of physical contact, such as keying on a keypad, or having another person visually checking one's identity. This control is even suggested in the case of identity cards for school students, and I am immediately struck by the need to successfully get elementary school students to use PINs. That aside, this brings us to a dilemma -- if the person is there and available to key in a PIN, why would you go to the expense of using RFID rather than the less expensive magnetic strip or even a barcode?

Exceptions are included for the incarcerated, who presumably have little right to privacy but must be carefully identified, and those in government-run hospitals. In the latter case, however, each new hospitalization requires the creation of a new number. What puzzles me in this section is the last statement below (and note the all in the first sentence):

(5) An identification document issued to a patient who is in the care of a government-operated or government-owned hospital, ambulatory surgery center, or oncology or dialysis clinic if all of the following requirements are met:
(A) The identification document is valid for only a single episode of care.
(B) The identification document may be removed and reattached when used on a nonemergency outpatient.
(C) The identification document does not transmit or enable the remote reading using radio waves of personally identifiable information. [My emphasis]
Would this bill prevent hospitals from coming up with a wrist ID that could provide information about the patient's condition? To have their chart number, or their date of birth so the patient's identity could be easily verified on the way into surgery? Without a discussion with the bill's drafters, it's hard to extract from all of this detail what capabilities the bill will and will not allow.

There are also sections that would cover law enforcement use of IDs (I'm thinking of mass arrests during riots, but I'm sure some people will imagine even more dastardly motivations), and use by emergency response personnel. There are some exceptions for locating people in immediate physical danger, but if I've read the technical protections sections correctly, there will be few opportunities to make use of the radio frequency device to perform these kinds of operations. I think that this section arises from our "McGiver Miracle" wishful thinking -- that if it did happen that I were buried alive, I could somehow turn on my cell phone and the hero(ine) of my fantasy would be able to use the technology in some highly creative way to discover and rescue me. Comforting, but unlikely.

The bill prescribes user education and notification, and states that the agency must provide a notice on each reader, or a list of the location all of the readers that can be used to read the card, or a web site address where such locations can be found.

Personally Identifying Information

Library cards are mentioned in a list of possible uses for identity cards, however a card that only has an assigned identifier such as a patron ID appears not to be covered under the restrictions of the bill relating to personally identifiable information as it is defined:

(o) "Personally identifiable information" includes any of the following data elements to the extent that they are used alone or in conjunction with any other information to identify an individual:

(1) First or last name.

(2) Address.

(3) Telephone number.

(4) E-mail address.

(5) Date of birth.

(6) Driver's license number or State identification card number.

(7) Any unique personal identifier number contained or encoded on a driver's license or identification card issued pursuant to Section 13000 of the Vehicle Code.

(8) Bank, credit card, or other financial institution account number.

(9) Credit or debit card number.

(10) Any unique personal identifier number contained or encoded on a health insurance, health benefit, or benefit card issued in conjunction with any government-supported aid program.

(11) Religion.

(12) Ethnicity or nationality.

(13) Photograph.

(14) Fingerprint or other biometric identification.

(15) Social security number.
I can understand most of these but I'm puzzled by numbers 11 and 12, Religion and Ethnicity or nationality. There is no question that these are sensitive bits of information, but they can hardly be considered "personally identifiable" under most circumstances. If anything, they are elements of group identification.

As with any statement about what is personally identifiable, however, it comes down to the fact that the right context can link almost any information to you. Your library card number becomes you when combined with the library's patron database. Your credit card number identifies you if one has access to the bank's records. Quibbling over what is and what isn't personally identifiable just doesn't jive with the reality of our data mined world, and it is unclear to me why a bank card number is personally identifiable but a library card number is not (if it isn't, by this definition).

Bottom Line

There must be some way to promulgate best practices for new technologies without creating laws. Since this legislation relates to identity documents issued by state or local government agency, couldn't the agencies refuse to do business with anyone who can't provide the required level of security? I must say, however, that if the goal of this legislation is to make it just too complex and too expensive to use RFID in government-issued identity documents, it is probably a good vehicle for achieving that goal.

Tuesday, January 16, 2007

Comments on D-Lib Article: "RDA... for the 20th c."

Diane Hillmann and I wrote an article called "Resource Description and Access: Cataloging for the 20th Century." It is a critique of what is being developed as the successor to the current library cataloging rules. I'm posting this here primarily to provide a place for comments... so feel free to comment, criticize, add to, or simply vent. Note that I have to approve comments so there will be some delay, and something of an interruption at times due to my ALA schedule.

Also check out the article in the same issue of D-Lib by Karen Markey. One of her points is similar to what Diane and I say, which is that it may be time to de-emphasize descriptive cataloging (at least for regularly published materials) and put that energy into better subject access. Markey suggests adding tables of contents and index terms to records, and developing ranking algorithms to help get the most appropriate material in front of users. Some of what she suggests I would put under the heading of "context" -- categories like reader level vis-a-vis the topic (beginner, expert), general topic area (science, history).

Tuesday, January 02, 2007

RDA at MARBI

There are two interesting documents on RDA that are being presented to the MARC standards group, MARBI, at ALA in Seattle.

The first is a "crosswalk" from the MARC format to RDA. Those of us who think about the MARC standard have been rather anxiously awaiting a look at MARC from the RDA perspective. We're still waiting, because this document is a look at RDA from the MARC perspective. It concludes that with a few "tweaks" you can fill in a MARC21 record using RDA data. By this same token you could show that you can fill in a Dublin Core record using RDA data. That's backwards, of course -- MARC is supposed to allow markup of the cataloging record, the cataloging record is not supposed to fit into MARC. But given how strong the MARC culture is, I wouldn't be surprised if some people consider "fitting the cataloging rules to MARC" to be a logical step.

This report is comforting because it appears to show that MARC does not need to change. The forthcoming report that maps from RDA to MARC will be considerably less reassuring. Even some of the suggestions in this report, such as
RDA has elements that are recorded using terms for an English language context, e.g., publisher unknown. It may be useful to identify such elements through MARC 21 encoding.

could have significant implications for the MARC record.

The document suggests that the various code and authority lists in MARC21 might be better managed as part of the RDA standard. I'd go this one further and say that values in authority lists should not be part of either standard. One of the big problems with MARC21 today is that it takes a change to the standard to add values to a list. Because the standards process is slow, by the time you've added a new physical format to the appropriate list, you've got two years of cataloged materials that you have to go back and add the code to. Code lists should be managed by the communities for whom they are relevant. There should be a process for updating them and a standard location for them on the net. Just like there is for the larger lists managed by Library of Congress for geographical names, languages, and others.

Note: this document refers at points to sections 9-13 of the RDA draft. This appears to be Part B of RDA, which I cannot find on the JSC site. If anyone knows where it is, please let me know.

The second document is a work in progress to categorize media types for resources. This was developed in conjunction with the publishing industry standards group that has produced the ONIX standards. The problem tackled is what is often referred to as "content versus carrier." (See recent article by Gorden Dunsire in D-Lib on this project.) These two have become rather hopelessly muddled in the MARC format, so this is an opportunity to get it straightened out. The level of abstraction here is high, so an item's content can be described as being of Character=language, SensoryMode=sight, ImageDimensionality=two-dimensional, Interactivity=non-interactive, and the carrier could be StorageMediumFormat=sheet, HousingFormat=binding, BaseMaterial=paper, IntermediationTool=not required.

In the end, however, users will come to the library catalog looking for a book, or a DVD, or a music CD. I hope we can present the catalog data in the user's language. Note that today's MARC21 record does not unambiguously identify books. The closest it gets is "language material" plus "Monograph/item". Unfortunately, things other than books fit that bill, including pamphlets and digital documents. Many library catalogs extrapolate the designation "book" from that coding because that's right most of the time. But we really need to keep the users in mind when we start categorizing materials.

FRBR OO - Not?

Posted on the FRBR blog was a link to an article by Allen Renear and Yunseon Choi
Allen H. Renear and Yunseon Choi: Modeling Our Understanding, Understanding Our Models: The Case of Inheritance in FRBR (95 KB PDF). In Grove, Andrew, Eds. Proceedings 69th Annual Meeting of the American Society for Information Science and Technology (ASIST) 43. Here’s the abstract:
They argue against seeing FRBR as having inheritance between the Group 1 entities because only the Item entity is concrete, the others are abstract.

The argument is simple: FRBR describes works as abstract and items as concrete. If all properties of “higher” entities are inherited by “lower” entities then items inherit the property of being abstract, and therefore items will be both abstract and concrete. But nothing is both abstract and concrete - therefore there is no unlimited general property inheritance in FRBR.
They make their point using a symbol set that isn't part of my vocabulary, so I'm taking on faith that they've proven this adequately. I have to say that I tend to consider all aspects of metadata to be abstract in nature, since it is a representation of something else, so their argument doesn't quite work for me.

This brings up for me, however, some larger issues, such as: Do we need a bibliographic concept that we can describe as a formal model? The FRBR model doesn't appear to survive formal analysis (see citations in the article), but does that really matter? I'm not a great fan of formality (at least not compared to some other folks), but it worries me that we are embracing a concept that we may not all understand in the same way. I have twice seen references to the "Work" entity as being "the idea." This strikes me as being horribly wrong, but without something a bit more (pardon the expression) concrete to go on, I don't see how we are going to come out with a definition that we can all agree on. And if we need to jigger the FRBR model a bit to make it work better, what's the mechanism for doing so?

Sunday, December 17, 2006

Digitization and the Catalog

I have just posted the preprint of my current column for the Journal of Academic Librarianship, titled "Mass Digitization of Books." It takes about 4-6 months for the columns to be published, and as I read over this one I can see that things have already changed. For example, when I wrote the column, Google was not yet allowing the download of its public domain books.

However, I should have included one more very important issue in the article, but it hadn't occurred to me at the time: the effect of this mass digitization on our catalogs. The cataloging rules require that the digital copy be represented in the catalog with its own record. This means that a library that undergoes a mass digitization project on its book collection faces doubling the number of book records in its catalog. Leaving aside the issues of user display for now, and assuming that the creation of the records requires very little human intervention, we can probably still calculate a significant cost in storage space (albeit cheap these days), the size of backups, the time to load and index all of those records, and a general overhead in the underlying database.

This brings up the issue of creating catalog entries that represent "multiple versions," that is, having a single record that contains the information for all of the different formats in which the book is available -- regular print, e-book version, digitized copy, large print. There are good arguments both for and against, and it's a complex discussion, but I'll just say that I am convinced that we could structure our catalog records in a way that would make this work.

Wednesday, December 13, 2006

Section 108, oh my!

The Library of Congress Study Group on Section 108 (of Title 17, the US copyright law) has issued a "notice of a public roundtable with a request for comments" in the ever-popular Federal Register. Which we all read daily, right? (I checked - no RSS feed that I could find, thank you very much.)

I have only read through the section on Topic A (there's also a Topic B), but I don't think I can go any further. This is about the worst mish-mash I have ever seen. If this is intended to clarify things, we are in deep doo-doo. (Believe me, I'm trying hard not to sound any less professional than that.)

OK, first, Section 108 is the section of the US copyright law with exceptions for libraries. In essence, section 108 allows libraries to make copies of items that are still under copyright in certain prescribed cases. Library of Congress formed a group to study Section 108 and make recommendations on how to update it for the digital environment. The group has been meeting, behind closed doors, for over a year. The group consists of lawyers, librarians, publishers, and lawyers. Oh, I said that, didn't I? They have held public meetings and have issued a document outlining what they see as the issues. This most recent call is proof that the study group is getting absolutely nowhere.

The subsections of Section 108 under question in this "notice" are the two that allow copying for lending, both within the library and over interlibrary loan. Because the study group's meetings are not open to the public, and because this is a highly political issue, the notice asks many questions that are suspiciously leading but there is no clue as to WHOSE issue it is. There also isn't much to explain the assumptions about technology that are behind some of the questions, so I often find myself unable to understand WHY a certain question is being asked.

That said, here are some examples of what I think are very strange statements and questions:

  • There is a great deal of concern about users receiving a copy of an item from a library through Interlibrary Loan without going through their own library. In other words, direct user borrowing. This violates what someone sees as the "natural friction" of ILL:

    it was presumed that users had to go to their local library to make an interlibrary loan request. ... for any user electronically to request free copies from any library from their desks, that natural friction would break down, as would the balance originally struck by the provision.

    Now this is just weird. Essentially they are implying that ILL was ok, even digital delivery, as long as it was inefficient and costly. If it becomes efficient, then it's just too much, and competes with sales. (I don't really see a difference between a user sending a request to their own library for an ILL rather than directly to the lending library -- except for the cost to the local library to pass the request through. And if that becomes efficient enough, the user won't even know how many middle-men there are in her request.)

  • Question 1:

    How can copyright law better facilitate the ability of libraries and archives to make copies for users in the digital environment without unduly interfering with the interests of rightsholders?

    What? Isn't this exactly what the study group has been discussing for 18 months? Now they put out a public notice asking the rest of us to answer the question? Haven't they at least worked it out to a set of choices or options? What have they been doing?

  • Question 3 (and Question 4 is very similar)

    How prevalent is library and archives use of subsection (d) for direct copies for their own users? For interlibrary loan copies? How would usage be affected if digital reproduction and/or delivery were explicitly permitted?

    Uh, isn't this something that someone should study? I mean, this is not something you ask people's (even educated) opinions on -- you've got to get facts and figures. It would be very interesting to know how much digital copying and delivery does go on in libraries. Without that information, we're just jabbering into the wind here, aren't we?

  • Question 5
    ... should there any any conditions on digital distribution that would prevent users from further copying or distributing the materials for downstream use?

    Well, there are conditions, and they are called copyright law. And of course they deter more than they prevent, but this really seems to be a silly question.
    Should persistent identifiers on digital copies be required?

    I wonder what they think that identifiers will accomplish? Do they see them as acting like watermarks, that would identify whose digital copy it is?

  • Question 7
    Should subsections (d) and (e) be amended to clarify that interlibrary loan transactions of digital copies require the mediation of a library or archives on both ends, and to not permit direct electronic requests from, and/or delivery to, the user from another library or archives?


OK, I'll stop here. As I have said, these statements and questions are so odd that I have no idea what happened in that closed room but it was weird.
Let me remind you that anonymous comments are allowed on this blog. So if you have some inside information on what the real problems are that are behind these questions, I would love to hear from you.

Sunday, December 10, 2006

The keyboard

I spend a lot of time each day "working the keyboard." It's easy to take it for granted; I learned to touch type in junior high school when the ability to type with speed and accuracy was part of a common job description. Little did we know at the time that we were heading into a future when everyone typed, and that typing would no longer be considered a special skill. (Nor would it be considered something "girly".)

There has been some questioning of the keyboard in the form of criticism of the QWERTY design. I tried switching to a Dvorak keyboard for a while, but didn't have the patience to work up to an approximation of the unconscious ease with which I type today. Recent ads I've seen are touting voice recognition as the replacement for typing, but I don't want to say all of my thoughts out loud, and in most offices with open or cubicled designs voice recognition would lead to cacophony. No, I'm happy to type, I just want it to be more efficient.

What I haven't seen questioned, yet it must have occurred to someone, is why we are still typing every letter when software could fill in or complete most words for us. Remember the ads that used to be on the back of magazines: "if u cn rd ths u cn gt a gd jb"? That's how I'd like to type. Yes, I can add those into my MS Word autocorrect, and I have placed a select number of long words I hate to type into the list. But we know that our language is very predictable and we should be able to take advantage of that. There are interesting IM keyboard options like T9 Word -- although obviously, the IM vocabulary doesn't need a large dictionary behind it. Open Office tries to help out by auto-completing words as you type, but this is useless for a touch typist because you have to 1) watch the screen (I often type while staring into space) and 2) take your fingers off their normal home row positions to hit the enter key. The Open Office method might work with a re-organized keyboard with a special key that means "go for it" when the screen shows the correct word, but I still think that would be slower than touch typing.

A neighbor of mine is a court reporter. She has the chorded court reporter "typewriter" which today hooks into a computer that auto-translates from the shorthand coming out of the device to words. The output isn't perfect, but it's good enough to be used in a courtroom in real time to feed the text to lawyers. That shows me that it can be done. Yes, of course, we'd all have to learn something new. But upcoming generations would benefit from a better solution to getting words onto a screen.