Saturday, December 27, 2008

FRBR and Group 2 & 3 Oddities

You've probably realized by now that I cycle back to FRBR frequently, each time discovering something new. New to me, at least. Perhaps because of not being a cataloger it seems that I have missed some key concepts in earlier readings. This might help explain some misunderstandings between me and more catalog-savvy folks.

This time I was thinking about the way that the entities are used with the subject relationship. But before I get to that, there's always the publisher to torment me.

Creators and Publishers in FRBR and RDA

The Group 2 entities have what is called "responsibility relationships" with the Group 1 entities. The diagram (Figure 3.2, p. 14) shows the two Group 2 (G2) entities, person and corporate body, to related to the Group 1 entities in the following way:
Work is created by... G2
Expression is realized by ... G2
Manifestation is produced by ... G2
Item is owned by ... G2
(Note that I find it odd that FRBR limits the Group 1 to Group 2 relationships to only four, and only one per Group 1 entity, but that is how it is written. It makes me wonder what one does with, say, an illustrator of a particular expression of a book. Surely the addition of illustrations doesn't make it a new work?)

In section 4 of FRBR, the Group 2 entities are not included in the lists of attributes of the Group 1 entities. In other words, when you read the list of attributes of a work, there is no mention of creator, and the list of attributes of an item does not include owner.

I was therefore surprised to find among the attributes of a manifestation:
4.4.5 Publisher/Distributor
The publisher/distributor of the manifestation is the individual, group, or organization named in the manifestation as being responsible for the publication, distribution, issuing, or release of the manifestation. A manifestation may be associated with one or more publishers or distributors.
Since Group 2 entities are not listed as attributes in the Group 1 attribute lists, this pretty clearly states that publisher is not a person or corporate body entity.
Yet, the section on relationships between Group 1 and Group 2 entities says:
5.2.2 Relationships to Persons and Corporate Bodies
The entities in the second group (person and corporate body) are linked to the first group by four relationship types: the “created by” relationship that links both person and corporate body to work; the “realized by” relationship that links the same two entities to expression; the “produced by” relationship that links them to manifestation; and the “owned by” relationship that links them to item.
Essentially, this apparent inconsistency between the definitions of the entities and the attribute list for the manifestation has to do with the practice of transcribing data from the manifestation:
At first glance certain of the attributes defined in the model may appear to duplicate objects of interest that have been separately defined in the model as entities and linked to the entity in question through relationships. For example, the manifestation attribute “statement of responsibility” may appear to parallel the entities person and corporate body and the “responsibility” relationships that link those entities with the work and/or expression embodied in the manifestation. However, the attribute defined as “statement of responsibility” pertains directly to the labeling information appearing in the manifestation itself, as distinct from the relationship between the work contained in the manifestation and the person and/or corporate body responsible for the creation or realization of the work. (Section 4.1)
What this points out is that while FRBR supposedly puts forth an entity-relation model, in fact it is no more ER than our current bibliographic model with its mixture of transcribed data, cataloger supplied data, and controlled headings.

Then Comes Group 3

This is easier to explain, because it is very simple: The Group 3 entities (concept, object, event, place) can ONLY be used as subjects, e.g.:
For the purposes of this study places are treated as entities only to the extent that they are the subject of a work (e.g., the subject of a map or atlas, or of a travel guide, etc.). (section 3.2.10)
This eliminates any thought of using place as in "place of publication." Not to mention that each of these has a very limited attribute list; in fact, they each have exactly one attribute:
term for the concept/object/event/place
The Upshot

The upshot is that FRBR does not give us a true entity-relation model for our bibliographic data. This is frustrating for those of us trying to move library data in an ER direction, and it means that to achieve the ER model we will have to go beyond what exists today in FRBR, and beyond the version of FRBR that has been realized in RDA. I've kind of known this, but it's discouraging to have it confirmed in the FRBR document itself. Even more frustrating that it's been there the whole time and I missed it.

I've looked again at FRBR in RDF and the Scholarly Works Application Profile, and both make some interesting extensions to the FRBR concepts, taking them further along the ER road. It seems to me that the DC/RDA work will need also to deviate from FRBR in order to achieve its goals. The big question is: how far can we go and still be compatible with library data?

Tuesday, December 23, 2008

Monday, December 22, 2008

Google Replies on OCA Blog

The Open Content Alliance blog has a post on the Google/AAP agreement with a lengthy reply from Dan Clancy of Google Books, and my reply to Dan's reply.

LC forces take-down of lcsh.info

I am beside myself with fury. I hardly know where to begin. Not long ago, Ed Summers took the LCSH authority file and created an online site with the LC Subject Heading authority file re-formatted as a SKOS vocabulary. For the first time, Web services could link directly to LC subjects as represented in the authority file. And some did.

But the Library of Congress, our Federal, if not National, library, has required Ed to take down the site. A site that contained nothing more than LCSH in a usable form. Data that SHOULD be in the public domain, for anyone to use as they wish. This is an assault against libraries everywhere, an act of censorship.

You can read Ed's statement on lcsh.info.

I would very much like to hear LoC's statement about this. They should not be allowed to control the use of this data, data that belongs to all of us.

Ed couldn't refuse the Library's demand, but anyone who isn't an employee of LoC should have greater freedom. Let's gather around a find a new home for LCSH, one that can't be removed from the public.

Thursday, December 04, 2008

Google and Fair Use

There's some background to the Google/AAP settlement that I believe is key to understanding the subtext around it. This won't be news to most folks, but I thought it would be good to re-articulate it in the context of the settlement, lest we forget.

Google's first business is that of indexing resources that are on the web. I'll talk about them as if they were all texts because it's easier, but the same thing could be said for images and other resources.

To do the indexing, Google must make a copy of the web page or document. Using this copy, it adds the page to its search engine. As a good citizen, Google pays attention to the robots.txt file, and does not index pages where the site owner has opted out of being included in search engines.

This is all fine and unremarkable until you look at it from the point of view of copyright law. Copyright is specifically about... making copies, and it gives the right to make copies, or to authorize the making of copies, to the copyright holder. That can be the author, or someone to whom the author has passed along the right. Copyright holders must opt in to the making copies: they have to give permission. The default in copyright law is that copies cannot be made unless the copyright holder gives approval.

So the big question is: Is Google violating copyright law by making copies of web pages without the permission of the copyright holders? There are two main ways of looking at this:
  1. The web is different from the print environment. Anyone who has put their works out on the web has agreed to copying because no one can even view the work without making a copy. If they don't want people copying, they need to hide their works behind a security screen. However, there is no such exception or wording in copyright law that would support this.
  2. The web is not different from the print environment. But Google is just producing an index and there is nothing in copyright law that would prevent someone from producing an index of words in texts. The incidental copies that Google makes in order to produce the index are allowed under the Fair Use aspects of the copyright law.
So then we move on to the Google Books project. Initially, Google claimed that it was doing the same thing with books as it does with the web: making incidental copies in order to create keyword indexes to the texts. In terms of copyright law, argument #1 is pretty much out because these works can be read without making a copy, so the copyright holders haven't agreed to let their works be copied. This leaves us with argument #2: it must be fair use.

In fact, Google did and does make the fair use argument. The libraries that partnered with Google also came to the fair use conclusion in at least some cases. The CIC project FAQ says:

University of Michigan said this in 2007:

Does this project comply with copyright law?

Yes. This project was undertaken with careful attention to the law and to the rights and responsibilities of the various parties involved. The purpose of copyright law is to promote progress in society. We are confident that the Books Library project is fully consistent with the fair use doctrine under U.S. copyright law and the principles underlying copyright law itself. Copyright law strikes a balance between rewarding creators of intellectual property for their creations and facilitating public access to these works in ways that do not create a business harm. For books, this means ensuring authors write books, publishers sell them and libraries lend them. By making books more discoverable, Google is enhancing the ability of authors and publishers to sell books to an audience beyond the traditional book market.

What was at stake with the AAP lawsuit was exactly this decision about Fair Use. If copying the books for the purpose of indexing were determined to not be fair use, then this decision could bleed over into the web. And of course it would mean the end of Google Book Search (which has now become Google Book Store). Although Google has always provided a confident posture to the public, declaring unwaveringly that what it does as a search engine is perfectly within copyright law, the idea of going to court over the issue would have put their entire operation at risk.

Now back to libraries. Fair use is not a list of things you can do but a judgment call relating to some complex factors. Some key factors have to do with whether your use is commercial in nature or could compete with the exploitation of works by the copyright holders. There are, in addition, exceptions in the copyright law relating to research and study, and special exceptions for libraries. In fact, in relation to copyright law, libraries and educational institutions get considerably more latitude in using works than do commercial enterprises. As an example, a teacher can make copies of an article for her students as part of a lesson, and that is generally considered fair use. A company manager who wants his staff to read an article cannot rely on fair use for copying, but must apply to the copyright holder (usually through an intermediary such as CCC) and pay a fee. (See the Texaco case.)

What happened with Google Book Search and the AAP is that the digitization of the libraries' books and subsequent use of those was judged not by the criteria that would be used normally for libraries, of course, but by the criteria that would be used for a commercial entity. That's totally logical, since although Google was partnered with the libraries, the primary use of the materials was to fuel Google Book Search, an obviously for-profit activity.

Libraries have gotten the short end of the stick because their use of their own materials became commercialized through their partnership with Google. If instead libraries had managed to digitize the books on their own, the outcome would have likely have been entirely different (if any lawsuit had been brought, which might not have happened). I believe that libraries could be found to have a fair use case for digitizing their works for the purposes of searching, and could be allowed to use those digitized copies for the exceptions spelled out in section 108 of the copyright law (such as providing access to the sight impaired, or for replacement of deteriorated originals). Unfortunately, the concept of digitization of the contents of libraries has now been tainted with the air of commercialization and has earned the wrath of the publishers and authors. The Google/AAP settlement has created a mechanism that ignores the inherent rights of the libraries, but also makes it more difficult for them to justify undertaking their own digitization project.

This is why I disagree heartily when I hear statements like:

We're delighted that this agreement creates new opportunities for libraries and universities to offer their patrons and students access to millions of books beyond their own collections. (from Google)
The settlement might look good from the point of view of a commercial entity facing copyright law, but it binds the non-profit educational and cultural heritage community to legal decisions designed for the for-profit sector. This is not only not a win for libraries, but it will hinder libraries in their efforts to make use of current technologies to further the arts and sciences.

Friday, November 28, 2008

OCLC Use Policy Details: Use and Transparency

An interesting aspect of this policy is that it is entirely about the use of WorldCat records. That may seem obvious from its title, but what I am interpreting from the policy language is that the policy covers all WorldCat records currently in existence, regardless of when they were created or the policy in force at the time that were first used. Creation or update of records take place at a particular time, while use is an ongoing activity. I'd like to cover some possible consequences of that.

Agreement to the policy

OCLC has stated that the Policy will go into effect in mid-February. It appears that current Members will be "grandfathered" in under the policy, their continued use of OCLC being their agreement to the terms. The Policy also covers Non-OCLC Members, who will not have made any agreement with OCLC, and I am hard pressed to understand why those organizations would abide by the terms of the Policy. 

Versioning and records already "in play"

Section E.7 says that OCLC can make changes to the policy, and that those changes will apply to use from that point on, essentially what is happening now with this Policy. Although they have agreed to place a version indication in the policy statement field in the WorldCat MARC records, I'm unclear as to what role that version would play. Instead, it seems to me that the policy implies that all WorldCat records will be covered by the current policy, whatever version that is. If this is not the case, then it isn't clear how the new policy can apply to records obtained from WorldCat before the Policy was in force. Yet this is exactly what is implied in the section on adding 996 fields on page 8 of the FAQ:
B. Retrospectively. For records that already exist in your local system, we encourage you to add the 996 field to WorldCat records transferred to others. Should you choose to use it, the field should have an explicit note like the examples below:

MARC:
996 $aOCLCWCRUP $iUse and transfer of this record is governed by the OCLC® Policy for Use and Transfer of WorldCat® Records. $uhttp://purl.org/oclc/wcrup/1.0
"Retrospectively" in this case means for records that were created before OCLC began adding 996 fields, and thus before the Policy goes into effect.

With this control over the use of all WorldCat records in existence, OCLC could become a highly disruptive force for anyone with ongoing relationships around bibliographic records. Because the policy could change again regarding records that have already been transmitted, anyone developing applications around use of WorldCat records is left with great uncertainty. Absent a good survey of the OCLC record use landscape, it is hard to know how many organizations and uses could be affected by this because we don't know all of the many ways that organizations are transmitting, receiving and using WorldCat records. However, with a policy based on use, possession of WorldCat records is like having a ticking time bomb since you have no assurance that your use will be permitted in the future.

Transparency

The "out" for all of these areas where it isn't clear what use is or is not allowed is to file a WorldCat Record Use Form with OCLC.  OCLC will then determine if the use is allowed. Section E.6 says:
OCLC has the sole discretion to determine whether any Use and/or Transfer of WorldCat Records complies with this Policy.
If I were an OCLC Member organization, I would want this process to be as clearly defined and as transparent as possible, if for no other reason than to avoid any semblance of discrimination against parties making requests. For publicly funded libraries, participation in a process that even appears to some to exhibit prejudices could be a public relations disaster. The only way to demonstrate fairness is to have a process that is open and auditable. The same section says:
In the event OCLC identifies a Use and/or Transfer which does not comply with this Policy, OCLC shall notify the relevant OCLC Member(s) and/or Non-OCLC Member(s) and such parties agree to work with OCLC to resolve the noncompliance.
I would go further and ask for the development of a publicly available set of guidelines for use of the records, and a formal appeals process that has member input. 

OCLC Use Policy Details: Your Records

There has been a lot of excellent commentary about the proposed OCLC record use policy. What I want to do here is highlight a few details about the policy that I haven't seen discussed elsewhere. The first is...

Your original cataloging


There are two areas where it becomes important to identify "your records." The first is in section B.3 where "WorldCat record" is defined. In the final paragraph (top of p.2) it states:

An OCLC Member or Non-OCLC Member may Use or Transfer the following without complying with this policy: (i) a WorldCat Record designated in WorldCat as the Original Cataloging of the OCLC Member or Non-OCLC member...
In other words, your own original cataloging is not covered by this policy. That's good news, but the practical application of this may not be simple. The way to determine this is by reading the MARC 040 $a subfield, presuming that the system you used at the time set this correctly. There is also the fact that OCLC merges duplicate records, so two instances of original cataloging could become one in OCLC...

Then there's the issue of how this affects down-stream users. For example, if Library A gives a copy of all of its original cataloging to Library B, and says: "no restraints on use," is Library B still held to the policy in terms of its use of WorldCat Records? According to the policy (E.5):

Regardless of the source from which WorldCat Records are received, Use and Transfer of WorldCat Records is authorized solely by OCLC pursuant to this Policy.
This seems to contradict the "your original cataloging is not covered" clause, although perhaps contract law deals with these kinds of apparent conflicts in some neat way. I would say that your original cataloging is not considered a WorldCat Record (as defined in the policy) except that the language of the exception refers to the original cataloging records as WorldCat records.

Also not clear is how this relates to the request to include the OCLC policy field in exported records. Although it isn't stated here, it would seem that original cataloging records should not contain the statement. (Those records could, however, be given a CC license by the originating library.)

Your holdings

Another key area relating to a library's own records is section D on the transfer of WorldCat Records. Section D.1.a states that libraries can transfer WorldCat records of their own holdings to other Members and Non-Members. Holdings is defined in the glossary as the OCLC institutional symbol on the record.

Section D.3 gives the logical converse of that: that to transfer WorldCat records that aren't of your own holdings, you must obtain permission from OCLC. This places restrictions on any institution that has received records from others, and could have implications for union and consortial catalogs. There isn't any mention of consortial agreements in the policy, yet many libraries already share their records in one or more such databases.

-------

Even if we work out the conceptual issues, both of these pose some real challenges in implementation since our bibliographic data today often does not clearly define the origin nor the source of the record, especially data that is not transmitted in MARC format. I'm really not at all sure that we could actually do what the policy requires.