Showing posts with label open access. Show all posts
Showing posts with label open access. Show all posts

Friday, April 05, 2013

The "Mellen Mess" and the changing role of publishers

Reading about the "Mellen Mess" -- the case of the publisher that is suing a librarian who criticized the quality of the houses's output -- I found the most interesting discussion to have taken place in the comments area of the original post (available via the Wayback Machine). One poster says:
On the other hand, I would say that few if any publishers do not publish a number of books that I would not buy.
To which Dale Askey replies:
The fact is, however, that libraries have to be able to trust presses to turn out good titles, or our work becomes impossible given the sheer global output of scholarship... libraries lack enough qualified subject expertise to make such judgments at the necessarily granular level, and the trend here is not encouraging. Subject librarianship is dismissed as a relic of a past age, and we now talk about “patron-driven” acquisition as if it were the Holy Grail. Having spent a brief but wonderful portion of my career as a focused subject librarian for an area where I have expertise, I know the benefit of reading substantive reviews and making intelligent choices about individual titles, but even that library no longer has the funds (or perhaps just lacks the will to commit the funds) for such esoteric enterprises.
What I think we see here is evidence of a substantial change in what it means to be a publisher in this age of "everyone can be a publisher." First, a little history.

Turin book fair, 2007
The first followers of Gutenberg were equal parts scholar, technician and businessman. There was never any question that producing print was a for-profit activity, and the same printers who turned out carefully edited classics also printed the first advertisements as well as a large number of indulgences to be sold to wealthy (but not well-behaved) Catholics. Well into the late 19th century, publishers were also printers, and often saw themselves as having a key role in scholarship and culture. The reputation of the publisher was what made the introduction of new, unknown authors possible.

Turin book fair, 2007
Although I am at my very core a "book person," I was unaware of the culture of publishers before visiting Europe and attending both bookstores and a few book fairs there. What struck me immediately was that the book covers represented the publisher more than the book itself. Near a university I found a bookstore that was entirely organized by publisher -- not by topic -- so that the only access other than "known item" was browsing by publisher.

 By my own observation, by the 1950's the role of the publisher in the US was subordinated to the book, preferably a best-seller. We could all name key books (Catcher in the Rye, To Kill a Mockingbird, The Spy Who Came in from the Cold), but I doubt if many of us could name the publishing house that issued them.

As Epstein and Schiffrin explain (see Further Reading), the purchase of publishing houses in the late 20th century by companies with a primary interest in profits, unhindered with cultural concerns, has made the publishing house no more than another business. From scholar-printer-businessman, only the latter role remains. If "best-selling" is your idea of quality, then these publishers can be considered consistent and trustworthy. If you are looking for greater cultural pursuits, you will probably be disappointed.

While that describes popular publishing, scholarly publishing has retained the publisher reputation... at least until very recently. While there still are known scholarly publishers whose output can be trusted sight un-seen (as Askey explains), there are many new entrants to this business area whose primary goal is income, not scholarship itself. This seems to be following a similar path to that of popular publishing, but with a twist: scholars must publish. The real culprit in this story is the "publish or perish" culture of academia. It matters not that there is no audience for a scholar's work; in fact, being actually read is rather icing on the cake. The main thing is that a scholar must get his or her work produced by someone acting as a publisher. It is therefore unremarkable that publishers have come on the scene to address this market.

The big "however" here is that while author fees may cover the cost (plus profit) of publishing an open access article, printed books still need to have some sales. Throughout the history of publishing, vanity books have been known as money-losers,* and some publishers have contracted with the authors to buy back any un-sold copies. This is more than an un-tenured faculty member can afford, however, so the business of publishing books by academics is one that wise investors would avoid.

The upshot of the story here is that we've gotten ourselves into an untenable position between the pressure to publish and the actual market for published works. Something has to give, and it has to give at both ends of the equation.

The next step, then, is improving the social media that the academic community uses so that the "post publication peer review" becomes the filter for quality and importance. 

---------------

* I ran into a great rant by a 19th c. Italian publisher about vanity publishing while doing research on Natale Battezzati. I unfortunately didn't mark it, but if I find it again I will link it here.


Further Reading

Epstein, Jason. Book Business: Publishing Past, Present, and Future. New York: W.W. Norton, 2001.

Schiffrin, André. The Business of Books: How International Conglomerates Took Over Publishing and Changed the Way We Read. London: Verso, 2000.

Wednesday, February 06, 2013

Book people v. article people

I am definitely a book person. When I want to learn about something, I want to read hundreds of pages about it. I have a half dozen books on copyright, more than a dozen about the social "questions" around the Internet, a handful on the Semantic Web, two shelves of books on libraries (history, cataloging, theory of knowledge organization), and now four books on cognitive science and the theories around concepts.

I've done work with people who are definitely "article people." Mostly academics, these folks rarely delve into a book since their scholarly conversation takes place in articles published in journals. My guess is that once you reach the level of knowledge that these folks possess, the breadth of a book contains nothing new and all of the interesting stuff comes out in article-sized chunks.

I also like to follow-up on my reading. When a bit of reading focuses around a place, like Bletchley Park or Los Alamos (as recent books on computer history do) I have to find them on a map. Concepts mentioned but not covered in detail require a visit to Wikipedia. I hunt down works cited in particularly intriguing passages. And it is in the midst of this last activity that I run into perhaps a hint about my attraction to books.

Because I am "unaffiliated" with an institute of higher education, it is easier for me to obtain books than to obtain articles. Books are available used or new in an open marketplace, and I find it to be rare that there is reference to a book that I cannot get at what seems to me to be a reasonable rate. But when I look up an article that I might be interested in, I often get something like:
 or

Yes, Wiley asks for $29.95 for an article, and JSTOR asks $38.00. I have seen these prices on articles as short as six pages. These prices are for the download of a PDF file, not an offprint to be delivered by express mail. I can only assume that they have no desire to sell access to individual articles, because the pricing is so out of whack with retail publishing. Remember, these are academic articles that quite frankly haven't a large audience. But they already exist in PDF and are available to members of subscribing institutions. In a world of $.99 pop songs and $9.99 best-selling e-books, these prices are just absurd.

One of the books I am reading at the moment is a compilation of essays called "Concepts: core readings." At least five of the essays were previously published in journals and when I looked them up the download price was $39.95 each. That's  about $200 for 100 pages of a 650-page book that retails for $55.

If we want "equal access to information," as we often claim we librarians do, then we need to do something about journal article pricing. I'd be quite willing to pay $2-$4 for an article, but the $30-$40 price range is ridiculous. I'm sure that these journal companies sell very few, if any, full-price articles. As we've seen with other media, when the price is right, it becomes as convenient to pay the price as it is to bother to pirate the materials (which in my case means borrowing someone's academic identity). Surely selling zero articles at $39.95 isn't better than selling a handful of articles at $2 each.

It's great that JSTOR is now offering some articles for free (although I have yet to be able to create an account since their site just hangs when I try), and I wouldn't suggest that JSTOR should be providing an entirely free service, since they have expenses. But $38 for an article is not just too much, it is prohibitive, and it unnecessarily creates an inequality of access. Someone needs to do to the journal publishers what Apple did to the music industry: show them the money.

Wednesday, September 28, 2011

Europe's national libraries support Open Data licensing

 "Meeting at the Royal Library of Denmark, the Conference of European National Librarians (CENL), has voted overwhelmingly to support the open licensing of their data. CENL represents Europe’s 46 national libraries, and are responsible for the massive collection of publications that represent the accumulated knowledge of Europe.

What does that mean in practice?
It means that the datasets describing all the millions of books and texts ever published in Europe – the title, author, date, imprint, place of publication and so on, which exists in the vast library catalogues of Europe – will become increasingly accessible for anybody to re-use for whatever purpose they want."

From an announcement by the Conference of European National Libraries.


Thursday, March 24, 2011

Open Data II

In this post I want to talk about some of the Open Government Data (OGD) projects taking place around the world.

Open government data is assumed to be a given by many in the US because our copyright law states that federal government data is not covered by copyright. (The situation in US states can vary, but the federal government's declaration sets the tone.) In other countries the situation is less clear and governments do not have a mandate to make data open. However, the open government data movement has purred on a number of fast-moving activities, many sponsored by governments themselves that encourage citizens to download and use government data.

The UK government has a site, Opening up government, where it not only shares data but encourages people to develop apps that use the data. Apps here can alert you to new building and planning projects in your area, and give you real-time public transportation information.

The EU has its own Open Government Data Initiative. It provides the data under these terms of use:
All Data on dev.govdata.eu is available under a worldwide, royalty-free, non-exclusive license to use, modify, and distribute the datasets in all current and future media and formats for any lawful purpose and that this license does not give you a copyright or other proprietary interest in the datasets.
There is a European site for public sector information, the European Public Section Information Platform: Europe's One-Stop Shop on Public Sector Information Re-use. You can search by country and see news and developments relating to public data, much of which is available for re-use. Because many countries for not have an explicit statement in their copyright laws covering government data, one of the important early steps for these jurisdictions is to develop blanket licenses that they can apply to the data. So when you visit the site you see recent news that Norway has developed a license for its government data and is asking for feedback (if you read Norwegian).

To understand the force of this movement, it is said that Albania and Bulgaria are on the verge of opening some government data.

The Obama administration announced its Open Government effort on the first day of his administration.
To the extent practicable and subject to valid restrictions, agencies should publish information online in an open format that can be retrieved, downloaded, indexed, and searched by commonly used web search applications. An open format is one that is platform independent, machine readable, and made available to the public without restrictions that would impede the re-use of that information.
Wired has a US-oriented "how-to" wiki on OGD. (Of course, they include in their "how-to" examples MarijuanaLobby.org, being Wired, but it's a good example of the range of utility of OGD. )

Not all data is at the country level, of course, and the movement is reaching into lower levels of government. Paris has an open data portal, while Enschede Netherlands has an open data declaration for its information. In Italy, the government of the Piemonte Region has a website for its open data.

The government open data movement is heavily influenced by grassroots efforts to convince governments that open data is a good thing -- not just for government watchdogs and opposition movements, but for heathy government and strong business. In the UK there is a Working Group on Open Government Data of the Open Knowledge Foundation, an independent not-for-profit that is promoting, as its name says, open knowledge. In Italy there is the wonderfully named "Spaghetti Open Data." Spain has a broad coalition of non-profits that form the "Coalición Pro Acceso." The CKAN web site, which is a general archive of available datasets of all kinds, has OGD under a number of tags, such as "gov". [Just out: Open Government Data video.]

We hear a lot about problems with copyright, with DRM, with information providers who want to lock down their products. Government data covers a huge variety of information types and is often the key information needed for a lot of civic and scientific decision-making. OGD can generate a mountain of new knowledge, and then tell you how high the mountain is.

Friday, March 04, 2011

Open Data I

The idea of open data has gone from an extremist rallying cry to a mainstream movement. In the next few posts I'll highlight just an iceberg tip's worth here, but expect to see more about this every day that passes.

The UK's educational research arm, JISC (something like NSF but more for education rather than pure science) and the research libraries' organization, RLUK, undertook a study about the advantages and possibilities afforded by opening data from libraries, archives and museums. They have produced the Open Bibliographic Data Guide, which investigates the business case for providing bibliographic data that can be re-used.


This is a practical, not a utopian, vision of open data.
"In earlier times, observers may have considered the ‘open data movement’ as the preserve of a certain type of fanaticism also associated with Open Source Software (OSS) and Open Content, emotionally and ideologically linked to the spirit of 1969.

However, OSS and Open Content have now morphed in to propositions with clear business cases of interest to corporations, institutions and governments. National strategies and Chief Information Officers espouse Open Source Software for financial and business benefit, whilst academic leaders are supporting Open Access Journals and Open Educational Resources (OER)."(link)
The report gives 17 different use cases -- situations in which an institution might want to provide its data with some degree of openness.

1 – Publish data for unspecified use
2 – Publish open linked data for unspecified use
3 – Supply data for Physical Union
4 – Allow Physical Union Catalogue to publish data
5 – Expose data for federation into Virtual Union Catalogue
6 – Publish grey literature data
7 – Contribute data to Google Scholar
8 – Publish activity data
9 – Supply holdings data for Collection
10 – Expose holdings / availability data for Closest Copy location
11 – Share data for Collaborative Cataloguing
12 – Supply data for Crowd Sourced Cataloguing
13 – Supply data to be enhanced for own
14 – Publish data for LIS research
15 – Allow personal use of data for Reference Management
16 – Publish data for lightweight application development
17 – Allow commercial use of data in mobile application

For each of the cases the report discusses pros and cons for the institution, its users, and the world, as well as the business case for making ones' data open. They acknowledge the complexity of our current environment of bibliographic data ownership:
"Our problems with bibliographic metadata are quite specific:
  • Non-profit and commercial players have built businesses around datasets of MARC records, indexing / TOC services and journal Knowledge Bases – but what is original about those accumulations?
  • Bibliographic records in the circulation amongst libraries are of uncertain and complex provenance, with the exceptions of those explicitly tagged by a ‘vendor’ or exclusive to a special collection" (link)
JISC doesn't stop at this report but is sponsoring projects and ongoing activities in this area. Already the British Library has produced its British National Bibliography data openly for reuse. You can keep up with these activities through the newsletter (subscribe here) whose logo reads: One to many; Many to one: Towards a virtuous flow of library, archival and museum data.

Wednesday, January 14, 2009

OCLC pushes back policy to fall, 2009

OCLC has just announced that it is pushing back the date on which the new record use and transfer policy will take effect. The actual new date isn't known, but the announcement says:
In order to allow sufficient time for feedback and discussion, implementation of the Policy will be delayed until the third quarter of the 2009 calendar year.
OCLC will form a "review board" to solicit info from members and others, and to advise the OCLC board of trustees about the policy. Jennifer Younger will chair this committee.

This delay is welcome, but I am dubious that a review board would be able to convince the trustees that OCLC must welcome open access to bibliographic data. Minor tweaks to the policy are not going to make much of a difference, and I doubt that any "advice" is going to force the board to do an about-face.

Those of us who promote open access must use this time wisely. First, we need to get some solid legal advice. It's clear that OCLC can propose any kind of conditions in a contract and hope to get signers; it's less clear that OCLC can impose a contract on members 1) without their explicit agreement 2) that covers data created before the contract becomes valid 3) that binds third parties to the contract. Next, anyone who has bibliographic data should release it "into the wild" as quickly as possible. Once the data is circulating, it will not be possible to withdraw it. One solution is to create database dumps and to upload these to the Internet Archive. They will be there for downloading by others, and some of the data may end up in the Open Library. Assuming that bibliographic records cannot be covered by copyright, all of this data ends up in the public domain to fuel innovation and creativity.

Note: if you are preparing a data dump, my advice is:
  • use a standard format (MARC21, MARCXML, UNIMARC, etc.).
Be sure to include in each record fields that give:
  • your local record ID (MARC 001)
  • something that identifies the source of the record (your system or institution) (MARC 003)
  • the version date (either the last date the record was updated, or the date of the data dump) (MARC 005)

Monday, December 22, 2008

LC forces take-down of lcsh.info

I am beside myself with fury. I hardly know where to begin. Not long ago, Ed Summers took the LCSH authority file and created an online site with the LC Subject Heading authority file re-formatted as a SKOS vocabulary. For the first time, Web services could link directly to LC subjects as represented in the authority file. And some did.

But the Library of Congress, our Federal, if not National, library, has required Ed to take down the site. A site that contained nothing more than LCSH in a usable form. Data that SHOULD be in the public domain, for anyone to use as they wish. This is an assault against libraries everywhere, an act of censorship.

You can read Ed's statement on lcsh.info.

I would very much like to hear LoC's statement about this. They should not be allowed to control the use of this data, data that belongs to all of us.

Ed couldn't refuse the Library's demand, but anyone who isn't an employee of LoC should have greater freedom. Let's gather around a find a new home for LCSH, one that can't be removed from the public.

Monday, November 03, 2008

Google/AAP settlement

This Google/AAP settlement has hit my brain like a steel ball in a pinball machine, careening around and setting off bells and lights in all directions. In other words, where do I start?

Reading the FAQ (not the full 140+ page document), it seems to go like this:

Google makes a copy of a book.
Google lets people search on words in the book.
Google lets people pay to see the book, perhaps buy the book, with some money going to the rights holder.
Google manages all of this with a registry of rights.

Now, replace the word "Google" above with "Kinko's."

Next, replace the word "Google" above with "A library."

TILT! If Google is allowed to do this, shouldn't anyone be allowed to do it? Is Jeff Bezos kicking himself right now for playing by the rules? Did Google win by going ahead and doing what no one else dared to do? Can they, like Microsoft, flaunt the law because they can buy their way out of any legal pickle?


Ping! Next thought: we already have vendors of e-books who provide this service for libraries. They serve up digital, encoded versions of the books, not scans of pages. These digital books often have some very useful features, such as allowing the user to make notes, copy quotes of a certain length, create bookmarks, etc. The current Google Books offering is very feature poor. Also, because it is based on scans, there is no flowing of pages to fit the screen. The OCR is too poor to be useful to the sight-impaired. And if they sell books, what will the format be?


TILT! Will it even be legal for a publicly-funded library to provide Google books if they aren't ADA compliant?


Ping! This one I have to quote:

"Public libraries are eligible to receive one free Public Access Service license for a computer located on-site at each of their library buildings in the United States. Public libraries will also be able to purchase a subscription which would allow them to offer access on additional terminals within the library building and would eliminate the requirement of a per page printing fee. Higher education institutions will also be eligible to receive free Public Access Service licenses for on-site computers, the exact number of which will depend on the number of students enrolled."


TILT! Were any public libraries asked about this? Does anyone have an idea of what it will cost them to 1) manage this limited access and pay-per-page printing 2) obtain more licenses when demand rises? Remember when public libraries only had one machine hooked up to the Internet? Is this the free taste that leads to the Google Books habit?


Ping! The e-book vendors only provide books where they have an agreement with the publishers, thus no orphan works are included. So, will Google's niche mainly consist of providing access to orphan works? Or will the current e-book vendors be forced out of the market because Google's total base is larger, even though the product may be inferior?


Ping! We already have a licensor of rights, the Copyright Clearance Center, and it was founded with the support of the very folks (the AAP) who have now agreed to create another organization, funded initially by Google and responding only to the licensing of Google-held content.


TILT! Google books gets its own licensing service, its own storefront... can anyone compete with that? And what happens to anything that Google doesn't have?


Ping! It looks like Google will collect fees on all books that are not in the public domain. This means that users will pay to view orphan works, even though a vast number of them are actually in the public domain. Unclaimed fees will go to pay for the licensing service. Thus, users will be paying for the service itself, and will be paying to view books they should be able to access freely and for free.


Ping! We have a copyright office run by the US government. I'm beginning to wonder what that Copyright Office does, however, since we now have two non-profit organizations in the business of managing rights, plus others getting into the game, such as OCLC with its rights assessment registry, and folks like Creative Commons. Shouldn't the Copyright Office be the go-to place to find out who owns the rights to a work? Shouldn't we be scanning the documents held by the Copyright Office that tell us who has rights? (Note: the famed renewal database is actually a scan of the INDEX to the copyright renewal documents, not the full information about renewal.) Even if we had access to every copyright registration document in the Copyright Office, would we know who owns various rights? I think not. And how much of this will change with the Google opt-in system? I get the feeling that we'll maybe resolve some small percentage of rights questions, somewhere in the order of 2-5%. And it will, in the end, all be paid for by readers, or by libraries on behalf of readers.


TILT! Rights holders can opt-out of the Google Books database. If (when) Google has the monopoly on books online, opt-out will be a nifty form of censorship. Actually, censorship aimed directly at Google will be a nifty form of censorship.


GAME OVER. All your book belong to us.

Tuesday, February 19, 2008

Libraries -- Open

I think of libraries as the quintessential open institutions. At least in the U.S. Most libraries are physically open to the public (even those in private institutions), and many serve as community spaces focused on learning, exploring, and simply being. They also promote the open use of what we might call "cultural heritage resources." Libraries fight for open access and they fight censorship all of the way to the Supreme Court.

Yet - there seems to be a real barrier when it comes to libraries being open with their own catalog data. This seems rather odd because most libraries' catalogs are available for open access on the web. But try asking a library for some or all of that data and you suddenly hit a wall. Libraries don't like to say "no" so there's a lot of hemming, hawing, "we'll think about that," that goes on. But an out-and-out "yes" is rare.

I speak about this based on my experience with the Open Library, a project of the Internet Archive. The OL wants to create a humongous bibliographic database (right now only of book records) on the Web, using a wiki-like front-end that would allow anyone to edit the bibliographic data. To me the most interesting aspect of the project is that it would bring bibliographic entries to the web's surface; they could be the subject of links from other documents, and potentially could begin the creation of a bibliographic web linking books to each other. But in spite of putting out a call for bibliographic records and making personal and direct pleas to a number of libraries, the OL has received only a lukewarm response.

To be sure, any data submitted to the Internet Archive becomes publicly available. And at some time in the future it may be possible for people to download individual bibliographic records for their own use. I know that there is some speculation that OCLC "owns" the data and that the OCLC license may not allow this level of re-distribution. I also know that some records in library databases are covered by vendor licenses (other than OCLC). Presumably those could be excluded from the data set. But it still surprises me how un-open libraries are with their own data, given how much they encourage others to be open with theirs.

During the comment period for the Future of Bibliographic Control report, the Open Knowledge Foundation posted a call for library data openness on the OKF wiki. Many dozens signed their names to the OKF's call for open licensing of bibliographic data, including important people like Larry Lessig and Tim O'Reilly. The arguments in OKF's document seemed pretty clear to me:

Bibliographic records are a key part of our shared cultural heritage. They too should therefore be made available to the public for access and re-use without restriction. Not only will this allow libraries to share records more efficiently and improve quality more rapidly through better, easier feedback, but will also make possible more advanced online sites for book-lovers, easier analysis by social scientists, interesting visualizations and summary statistics by journalists and others, as well as many other possibilities we cannot predict in advance.
Nothing of this was included in the final report.

Libraries complain that they don't get the kind of attention that Web resources like Google and OL get. They complain about the lack of transparency of the commercial data vendors; that Google won't say how many books it has online nor will it reveal its work on attempts to rank book retrievals. Libraries could be doing this experimentation themselves, and in the open, if their data were on the Web. They could be visible, out there, allowing incredible innovation to happen based on the hundreds of years of collecting materials and creating relatively consistent metadata for those materials. Their reluctance to let their data out of the databases just baffles me, and isn't in concert with their stated goals of open access.