Showing posts with label scholarly publishing. Show all posts
Showing posts with label scholarly publishing. Show all posts

Saturday, September 24, 2011

Are Neutrinos Superluminal? Judge for Yourself.

Yesterday, I couldn't restrain the physicist inside me. I just had to watch the seminar from CERN where Dario Autiero presented the result from CERN and CNGS of a measurement of the neutrino's velocity. The measurement is a tour de force of modern experimental physics, which harnesses amazing technologies such as the global positioning system, large scale data processing, picosecond lasers, ultrafast digital electronics, billion-volt particle beams, and highway tunnels a mile beneath a mountain in Italy. The scientific communication system is also state-of-the-art. The preprint (with 174 authors) appeared in ArXiv.org on thursday night: "Measurement of the neutrino velocity with the OPERA detector in the CNGS beam",  arXiv:1109.4897v1 [hep-ex]. And CERN webcasted the presentation live around the world, with not even a hiccup in the video feed.

What the news reports have failed to convey is that despite the impressive effort, the accuracy of the velocity measurement is teased out of the data with statistical procedures that are sure to come under intense scrutiny.

The way the experiment works is this. Pulses of neutrinos lasting 10.5 microseconds are generated by the accelerator at CERN, pointed at neutrino detectors 730km away under the mountain in Italy. Neutrinos are incredibly hard to detect (they have no difficulty traveling through 450 miles of rock), so only a tiny fraction of them are detected. Over 3 years, 16,111 of the CERN neutrinos were detected in Italy. The shape and timing of each generated pulse is measured and stored to be compared later with the timing of the detected neutrinos. The nub of the matter is shown in this graph, which shows only the leading and trailing edges of the accumulated data:
As you can see, the leading edge of the neutrino pulse is about 500 nsec wide. The red line is the cumulative shape of the generated pulse, the data show the counts and relative timing of the detected neutrinos.

The evidence for superluminal neutrinos is that the red curves at on the bottom, shifted by 60.7 nsec faster than the speed of light, are a better fit to the data than the red curves at the top. The claim is made that the bottom fit is 6 sigma away from the fit at the top. What do you think? Isn't physics fun?
Enhanced by Zemanta

Article any source

Thursday, April 28, 2011

Open Access eBooks, Part 1

No Shelf Required: E-books in LibrariesI've been working on on a book chapter for a book edited by No Shelf Required's Sue Polanka. My chapter covers "Open Access E-Books". Over the next week or two, I'll be posting drafts for the chapter on the blog. Many readers know things that I don't about this area, and I would be grateful for their feedback and corrections. Today, I'll post the introduction, subsequent posts will include sections on Types of Open Access E-Books, Business Models for Open Access E-Books, and Open Access E-Books in Libraries. Note that while the blog always uses "ebook" as one word, the book will use the hyphenated form, "e-book".

Open Access E-Books

As e-books emerge into the public consciousness, “Open Access”, a concept already familiar to scholarly publishers and academic libraries, will play an increasing role for all sorts of publishers and libraries. This chapter discusses what Open Access means in the context of e-books, how Open Access e-books can be supported, and the roles that Open Access e-books will play in libraries and in our society.

The Open Access “Movement”

Authors write and publish because they want to be read. Many authors also want to earn a living from their writing, but for some, income from publishing is not an important consideration. Some authors, particularly academics, publish because of the status, prestige, and professional advancement that accrue to authors of influential or groundbreaking works of scholarship. Academic publishers have historically taken advantage of these motivations to create journals and monographs consisting largely of works for which they pay minimal royalties, or more commonly, no royalties at all. In return, authors’ works receive professional review, editing, and formatting. Works that are accepted get placement in widely circulated journals and monograph catalogs.

In the late 1970’s and 1980’s academic libraries became acutely aware that an expansion of research activity had resulted in the growth of both the numbers of journals and the numbers of articles published in the journals. The combination of increased subscription prices and the number of journals needed to support research resulted in a so-called “serials crisis”. Libraries were forced to cancel subscriptions. The reduction in circulation forced publishers to raise subscription prices further to make ends meet, and the resulting cycle of cancellations and price increases led to a fear that the whole system would collapse. If few libraries could afford subscriptions, fewer scholars would be able to read the articles, diminishing the attractiveness of publishing.

The advent of web-based publications in the 90’s led many to believe that the solution to the serials crisis would be a shift of the scholarly publishing industry to so-called “Open Access” business models. Open Access publications are those that can be read at no cost to the reader or the reader’s institution. The traditional model of publishing supported by subscription fees was thus styled as “Toll-Access” publishing. It was hoped that the combined cost reductions from digital distribution and automation would stop the cycle of rising expenditures.

Perhaps the most successful implementation of Open Access has been ArXiv, a database of digital preprints and reprints (“e-prints”) originally focusing on the particle physics community. Originally started by Paul Ginsparg, a physicist at Los Alamos National Labs, ArXiv is now located at Cornell University and hosts more than 670,000 scientific articles in e-print form. Authors deposit articles they’ve written into the repository, and other scholars are free to search, browse and download articles without needing any sort of subscription.

One reason for the success of Open Access archives has been that they have grown up in a parallel coexistence with the traditional academic journals, which have mostly shifted onto the web. In the so-called “Green” model for Open Access, many journals allow versions of accepted articles to be made available via repositories. Authors can thus submit their articles to high-prestige subscription-supported journals without worrying about colleagues’ access, because scholars that need to read their works can always access versions from free sources.

Meanwhile, the shift of traditional journals onto the web has allowed the rise of secondary distribution channels. Most academic libraries today enjoy access to a much broader range of journals compared to 20 years ago because of the availability of article databases that aggregate content from large numbers of journals.

The past decade has also seen the rise of “Gold” Open Access journals. These journals leverage low cost Internet distribution to allow articles to be read universally with no subscription charges. Led by Biomed Central and PLoS, these journals cover expenses by charging publication fees to the submitting author. They build prestige  and avoid becoming “vanity” presses by establishing rigorous review processes.

The success of Open Access journals and articles has for the most part not yet been duplicated in the word of books. There are a number of possible reasons for this. The first is the matter of cost. Publication fees for Open Access journal articles are in the range of $600-$3000; editing and production expenses for a book published by a university press are estimated to range from $10,000 for a book that’s mostly text to much more for a book with figures, photos, equations and cover art. Author-funded publication fees this large are unlikely to be practical, even with significant institutional subsidies.

Another factor holding back Open Access books may be a preference for print books over e-books. Books are much longer than journal articles, and many readers are uncomfortable reading a book on a computer screen. It’s only in the past two years that dedicated reader devices such as the Kindle and tablet computers such as the iPad have improved the e-book experience enough to gain wide consumer acceptance.

The business environment for book publishers is another possible factor. The university publisher loses money on much of its catalog, but compensates for this by having one or two titles that cross over to be successful outside the academic environment. Amazon.com has bolstered this pattern, by providing wide distribution for small print-run titles that would never have been available in bookstores before. In contrast, journal articles almost never cross over into non-professional markets.

Nonetheless, there have been a few notable attempts to publish Open Access e-books. I’ll cover these later in a section on business models for Open Access e-books, but it wouldn’t be right to omit mention of Project Gutenberg at this point. Project Gutenberg (PG) produced not only the first Open Access e-books, it produced the first e-books, period. Started by Michael Hart in 1971, PG aimed to take the text of public domain works and make them available via the Internet. To date, PG has put over 34,000 works into its collection, entirely through the efforts of volunteers.

Distribution of Open Access e-books can be thought of as an enterprise separate from their production, since the costs involved are of a different nature. The scaling laws of Internet distribution favor centralization, and as a result, organizations such as the Internet Archive are able to distribute appropriately licensed e-books on a vast scale; businesses such as Google are able to search and organize them; libraries, blogs, and portal sites are able to select and “curate” them. To some extent, this type of distribution depends on the self-contained nature of the book; it shouldn’t require the context of a specific website to retain and accumulate value.

Open Access for e-books provides many benefits in addition to allowing people to read for free. Access to the full text of books makes for more complete indexing. The utility of Google Books, and the effort Google has put into digitizing books from libraries, even when they are unable to make the books available because of copyright, is testament to the value of indexing the full text. Long-term preservation of our cultural heritage is another public benefit of Open Access to e-books.

next post in series -> 

Article any source

Thursday, January 6, 2011

Fundamental Constant Numerology

My father was obsessed with units of measurement and fundamental constants. He got his engineering degree at the Royal Institute of Technology (Kungliga Tekniska Högskolan) in Stockholm, Sweden. His favorite professor there was Erik Hallén, who was famous for his work on antenna theory and for laying groundwork for the world's most widely used system of measurement, the SI system.

My dad nearly failed Hallén's class, which could be one reason for his lifelong obsession with units. The other reason was that Dad was convinced that he could explain some of physics' deepest questions about the nature of matter by applying the electromagnetic theory he learned in Hallén's class. Dad   explained the structure of the electron by modeling it as a circulating charge wave in a resonant cavity formed by general relativistic warping of space. In his notes, he wrote
To me it looks like all the puzzle is defined and ripe to be put together and the extension to other particles will not be difficult- only time consuming.
Dad used these insights to come up with a relationship between the gravitational constant and the mass and charge of the electron. Here's his equation:
                     G = 6/π 10-44 Z0 c (e/m)2
where G = the gravitational constant (which defines the force that holds the universe together), Z0 is the impedance of free space, c is the speed of light, and e and m are the charge and mass of the electron.

Here's a prettier version of that equation, made using Roger's Equation Editor from the following TeX code:
                      G = \frac{6}{\pi} 10^{-44} Z_0 c (\frac{e}{m})^2
(TeX is the most commonly used formatting language for mathematics.)


A derivation of this equation would easily have earned my dad a Nobel Prize, but without a derivation and explanation of the underlying physics, it was just numerology. If you plug in the numbers, Dad's equation is within 0.025% off of the consensus value for the gravitational constant, whose experimental uncertainty is about 0.01%. Dad understood that his equation was worthless without an explanation, so he spent endless hours studying Bessel equations and all sort of obscure mathematics. He was sure that somehow, somewhere, there existed a solution to Maxwell's equations combined with general relativity to explain the 6/π and confirm his resonant cavity. (He said the 10-44 was just using the right units- I never understood that!)

The Nature of the Physical WorldMy dad was not alone in physics numerology. The fine structure constant, which is very close to 1/137, has attracted all sorts of numerological explanations. (Dick Lipton calls it a "miracle number") Arthur Eddington, one of the most famous physicists of his time, had an explanation for why the fine structure constant should be exactly 1/137 involving the number of protons in the universe. A more modern numerological result is that of James Gilson, whose suggested value for the fine structure constant is only 30 parts per trillion off.

While fundamental constant numerology has deservedly been on the fringes of science, new internet search technologies may soon change that. Last year, scientific publisher Springer introduced a beta service called LaTeX Search that allows researchers to search for LaTeX formatted equations in all of Springer's journals. (LaTex is a dialect of TeX most widely used for scientific publishing.) That's something you can't do with Google, or any other search engine. The ability to connect obscure mathematical discoveries from disparate fields of science could soon be facilitating new avenues of research, perhaps even new methodologies.

For example, I can search for a fragment of my dad's equation and get at least one result that seems relevant. I don't know of any meaningful discoveries that have been made so far with LaTeX search, but if my dad had been able to search all of the mathematical literature to connect his numerological result with a mathematical solution, perhaps he would have explained the gravitational constant and structure of the electron and would have won his Nobel Prize.

He would have been 83 today. Happy Birthday, Dad! We miss you.
Enhanced by Zemanta

Article any source

Thursday, September 16, 2010

Can Libraries Work Together to Acquire eBook Assets?

Libraries like to work together. They also love to form structures to shape this cooperation. There are all sorts of library associations, consortia, coalitions, collectives, cooperatives and federations. One of the big international library conferences is that of the International Federation of Library Associations and Institutions. There is even an International Coalition of Library Consortia. There's not yet been much effort to federate the library consortia coalitions.

When I suggested last month that libraries ought to form an ebook acquisition collective to buy up ebook rights and make the ebooks available on an open-access basis, some readers misunderstood what I was suggesting. They thought that I was suggesting something similar to the  purchasing consortia that many libraries have formed to aggregate their buying power and obtain lower prices and standard terms from publishers. Indeed, a recent report (pdf, 3.4 MB) from the Chief Officers of State Library Agencies (COSLA) has proposed formation of a national ebook buying pool. But this report doesn't contemplate making the ebooks open-access, it assumes the buying pools would negotiate deals with existing ebook providers.

To make it clear what I'm suggesting, I'm going to refer to "ebook asset acquisition" instead of just "ebook acquisition". The asset to be acquired is nothing less that the right to distribute an ebook, to anyone, anywhere, without charge. Here's the way I think of this:
  1. The mission of libraries should be to provide access to books to anyone in their community without charge.
  2. To fulfill that mission, libraries should be trying to acquire all the rights needed to give unfettered access to ebooks to anyone anyone in their community without charge.
  3. Community access to ebooks is in many ways incompatible with business models in which restricted access to ebooks is sold to consumers. It doesn't make a lot of sense to pretend that ebooks have the same usage characteristics as print books.
  4. Rational publishers should be willing to sell off ebook rights outright if the price that exceeds the present value of the ebooks expected revenue stream.
  5. The amount of money libraries currently spend on books of all types would be sufficient to acquire outright many works, particularly those sold primarily to libraries.
  6. The way for libraries to meet the publishers' price for these ebook assets is to organize into a collective that aggregates the total demand and acquires ebook rights.
  7. Therefore, an ebook asset acquisition collective out to exist. Why is this not happening?
To some extent, it IS happening.

Frances Pinter, the publisher at Bloomsbury Academic, gave a keynote at the O'Reilly Tools of Change for Publishing conference this past February on her proposal for an "International Library Coalition
for Open Books".

Pinter sees a funding model through the eyes of a publisher, which I'm not, but the idea is pretty much the same. A mechanism that enables open-access to scholarly monographs while keeping the enterprise of publishing economically viable would be of great benefit to society at large.

After the previous post, I got feedback that publishers would be very hesitant to sell off ebook rights; publishers view their intellectual property rights as "part of the company" and divesting of this rights would be a "liquidation strategy". I was surprised at this view. I would have thought most publishers are more interested in the production and launch of books than in servicing them in perpetuity. From a business point of view, books are like the loans made by banks. You put out money up front, and make that money back as a stream of payments.

The banking industry has advanced far beyond the view that holding assets is their core business. Banks that originate the loans don't service them any more, they sell off the loans so they can make new loans. And, no that's not what led to the banking crisis. Publishers of all stripes should likewise be willing to cash in their book assets and use the proceeds to produce new ones.

Another type of feedback I got was that the logistics of buying and selling ebook rights could be difficult. For example, how would a publisher plan print runs and promotion? If a print run had just been paid for, the publisher would be loathe to sell off ebook rights. If a book was due to be published in February and announced in a Spring catalogue that goes to press in December, when might it be sold to a library collective, and how quickly would the collective be able to decide?

There are a couple of answers. First of all, smaller print runs and more POD minimize the risk of printing too many books. This is happening anyway. Second, I assume that publishers can build in print risk to the price they ask for. If printing is really only 25% of first copy cost, well that's only 25%. Supply chain markup is as much as 50%

I'm a pretty strong believer in markets. I imagine that a library collective would help to create a more efficient market for publishing assets which are  deployed awkwardly in the current system. Nothing would force publishers to accept any disadvantageous price for an asset.  The library collective would weigh the price and quality of each asset before deciding, presumably using automated systems,  which ones to acquire collectively. Building internet-based markets is a technology that is reasonably well understood; librarians are familiar with automated purchasing systems and with collaborative selection. It shouldn't be too hard to put an expiration date on an publisher's offer to sell- we've all heard of eBay.

I expect that there will be a variety of product life-cycles for ebooks. The typical midlist book today spends some time on bookshelves before being returned to the publisher or set out on the discount racks. Its life may be extended on Amazon; used-bookstores enter the channel without contributing any revenue back to the publisher. The typical library-acquired title might have had a "toll-access" run of a year or so; the publisher could capture the most eager purchasers using delivery systems using DRM. Most libraries would wait for the book to be acquired by the collective at a discount and made open-access. This market market segmentation is similar to the hardcover/paperback model which adds efficiency to the current market.

Some publishers might try to sell ebooks even before the book is published; they might even succeed. Some books won't get bought at any price, let alone at the cost to produce them; others would probably be acquired even at an obscene profit margin for the publisher. Some things never change.

The biggest uncertainty with a system that allows libraries to collectively acquire ebook rights (including the rights to give them away!) is size of the revenue lost to free-riders. The free riders would be of two types; libraries that don't participate in the system, and book buyers not associated with libraries. The share of sales made by university presses outside of libraries varies from press to press, but some presses indicate that as much as 70% of their sales occurs outside of libraries. This share has increased over the last decade or so, thanks to new sales channels such as Amazon and tightening library budgets. One possible solution to this issue would be to open the collective to consumers; I will write more about this in the future.

Then there's the issue of libraries that would choose not to participate in funding the acquisition collective but would still benefit from the ebooks liberated by the collective. If an ebook rights acquisition collective comes to pass, we'd really find out how much libraries like to work together!

Article any source

Thursday, June 24, 2010

Inter-Library Loan Reinvented for eBooks and Just-In-Time

My graduate school training was in engineering and in physics. In engineering, you put things together and try to get them to work. In physics, you smash things (the polite term is "perturbation") to help you understand how they had been working. I still use these approaches to help me understand the things I write about. You can learn a lot about a system be noting the bits that squawk when faced with a perturbation of the system.

I got a lot of interesting feedback on my article on patron-driven ebook acquisition. It seems that this perturbation in library processes could have wide ranging effects far outside of libraries and book publishing. Coincidentally, the patron-driven model, along with other changes in the library/publisher ecosystem, was discussed last weekend at a meeting of the American Association of University Publishers (AAUP). Publishers Weekly has a nice report. (See also a report in the Chronicle of Higher Education.

The biggest perturbation being imposed on this system is of course the reduction of library budgets, which has come down quite painfully on university presses and their monograph businesses. Still, speaker Joe Esposito was surprised that the strongest reaction to his talk was to his prediction that libraries would make up a shrinking fraction of the university presses' sales.

It seems that there is worry that a contraction or restructuring of monograph publishing could have repercussions for how scholars obtain tenure in the humanities:
The fact that monograph publishing exists to support tenure and the structure of academic employment is an inconvenient truth that can no longer be glossed by either the Academy its associated University Presses. At some point the Academy is either going to have to stop expecting University Presses to fulfill this need, or find a more honest and transparent way of funding it.
if that's the worst thing that happens, well, what's the big deal?

It won't be a shift to patron-driven acquisition that kills off monograph publishing, however. My reasoning is that from the point of view of economics, patron driven acquisition is roughly isomorphic with the current system of just-in-case purchasing coupled with inter-library loan (ILL).

Here's how things work for print monographs. Suppose a university press published an obscure but brilliant scholarly monograph five years ago. It might have sold 100 copies for $100 apiece, most of them to libraries. At $10,000 gross revenue, it was hard for the press to make much profit, but occasionally they get lucky and make enough to cover the losses on the rest of their catalog. Now here's the problem: Over the five years, there were only about 100 scholars in the entire world that really wanted to read the monograph. Unfortunately, only 50 of them worked at institutions that purchased the book. The libraries of the other 50 didn't purchase the book because the selectors in their libraries weren't omniscient or perfect, and they didn't have mind-reading abilities or the power of divination. Or maybe the libraries used an approval plan that hadn't been crafted with the obscure field of this monograph in mind.

But those 50 others still got to read the monograph, because of inter-library loan. For some libraries, ILL is even a revenue center, because their costs to lend are less than the fees they charge. Although publishers made money from the 50 libraries that bought the book and didn't use it, they don't capture any of the revenue from ILL activity. The libraries that spent money to buy the book right away are partially compensated for that expenditure by ILL revenue or reciprocal loans.

Now let's think about what happens in a future where just-in-time ebook acquisition dominates. The 100 users still get to use the monograph, but none of them need to wait for an ILL transaction to go through. The costs are assigned to the institutions that actually use the work. If we assume that the price of the monograph is unchanged, the publisher's revenue is also unchanged; except it's pushed out to the time of usage, which can be many years, especially in the humanities. The time value of this revenue stream is reduced- it takes longer to make back the money spent on producing the book.

The compensation for the publisher is that the revenue continues for as long as the work is still used. The book doesn't go out of print. In addition, since users can discover the monograph more widely, and obtain it immediately, there is the possibility of making additional sales to users who would never have requested the title via ILL.

In a sense, the patron-driven acquisition model is souped-up ILL, with usage fees accruing to the publisher. The comments of Macmillan's John Sargent earlier this year that publishers would like to see fees for library ebook lending don't seem so controversial when examined under this lens.

It's worth thinking through a publisher's pricing strategy. If libraries persist in their preference to remove price as a factor in the patron's decision to use an ebook, then publishers have no incentive to cut costs and keep prices moderate. If libraries allow automatic purchase of any ebook under $100, then publishers will price all of their products at $99. A similar dynamic in the US health care industry has not worked well for consumers, to say the least. Indeed, one university press publisher writing about patron-driven acquisition and the AAUP meeting has opined that patron selection will lead to higher monograph prices:
What this Patron Driven Access model means to university presses is that our future is likely to include two things—higher prices and fewer titles.

It's clear that there would be winners and losers under a just-in-time acquisition system. Librarians don't always select what their patrons really want to read. Controversial works might do quite well, as should engaging but hard-to-categorize works and works that don't break new ground but are readable and useful. Dry, unreadable, redundant works that sell well today because of the author's fame or because they fit into a "hot" field of research will be losers. A work that today is unread because it's too innovative and ahead of its time will eventually find its time under the just-in-time acquisition.

The huge change for monograph publishers will be in the way they market their products. The emphasis will shift from pre-publication marketing to libraries towards search engine optimization and post-publication marketing directly to users. Famous professors may find themselves awash in free ebooks as monograph publishers jockey for key citations and mentions; social networks and subject specific communities will be prime targets of monograph promotion. Publishers will abandon library convention exhibits like ALA in droves; parties and receptions for librarians will disappear.

I suppose we should have fun with the current system while it lasts, even as there are new and more efficient things to build.
Enhanced by Zemanta

Article any source

Thursday, June 10, 2010

How Electronic Resources Really Get Priced

The recent letter (pdf) from the University of California Digital Library (CDL) about price increases  proposed by Nature Publishing Group (NPG) and the response from NPG have raised a storm of controversy and rebuttal (pdf). To me, the mess is symptomatic of a communications failure between publishers and librarians.

In the interests of promoting better library-publisher understanding, I've decided to reveal some secrets from both sides.

Libraries: Here's how publishers set pricing for electronic resources.

Once upon a time, pricing for library materials had a relation to the cost of their production. Even before the internet came along, this started to change. Printing costs fell, and more and more of the production costs of a good-quality journal were "first-copy" expenses. With electronic materials, the marginal cost of servicing an additional subscription became almost zero. Pricing then became a game whose object was to sustain existing pricing, along with "reasonable" annual increases of a few percent per year or so. (Nature Publishing translates "a few" to "7".)

The game was most difficult for very large or complex institutions. The value of a top medical journal to a US medical school is huge; the value of the same journal to a vo-tech school would be much smaller, but still significant. A medical school in a developing country will also need the journal, but it's not fair to ask them to pay the same as a US school. Differential pricing helps a publisher capture value while still extending access to customers who might otherwise be able to afford the journal. But how, then, to set pricing?

I learned the secret of e-resource pricing through long hours of research (spent mostly in bars). Here's how it works:
  1. Find out how much money the customer has.
  2. Set price somewhat higher than that.
  3. After hard bargaining by customer, offer discount to closely match customer's available funds.
  4. Swear customer to secrecy; you can't give that price to everybody!
In times of budgetary cutbacks, this pricing mechanism works to a library's advantage. A library that needs to cut its electronic resource expenditure in half simply needs to disclose to salespeople the fact that their funding has been cut in half, cancel the subscription...and wait for panic to set in.

The library's leverage will never be greater. It's much more painful for a digital publisher to lose an digital customer than it is to lose a print customer. That's because the publisher has to spend money on digital publishing infrastructure when its customer base grows, but doesn't get anything back when a the customers go away. If anything, the publisher will have to spend more on sales to try replace the customers.

Oh and by the way, libraries, this all works easier if things are quiet- when you agree to give a publisher a higher price than you wanted to, swear them to secrecy- you can't afford to give the same deal to every one of your publishers!

Publishers: Here's how to get libraries to cave on pricing

Many librarians suffer from feelings of powerlessness. They are prisoners of their patrons' needs and desires. They are captive to changing technologies and archaic standards. They are trapped in arbitrary budget gaps. And they are stuck in endless committee meetings.

Libraries are thus willing to spend a great deal on things which offer escape from powerlessness. They dislike monolithic packages that bundle content together, even if they save money. The current reality in libraries is that budgets have been cut. Publishers need to give their library customers options that help them deal with budget cuts. The smartest publishers can figure out ways to help libraries cut costs and free up funds currently spent in other areas.

Cool Hand Luke [Blu-ray]When I was developing an electronic resource management service for libraries, I had this recurring nightmare that e-journal publishers would someday make it as easy for libraries to activate and maintain an e-journal subscription as Apple's iTunes makes it to buy and maintain a song. My software would instantly become worthless. Just kidding- I slept soundly knowing it would never happen in a million years.

Perhaps NPG confused powerlessness with weakness and saw an opportunity to force CDL to act like the other prisoners. Perhaps NPG never saw the movie "Cool Hand Luke".
Article any source

Tuesday, May 4, 2010

Authors are Not People: ORCID and the Challenges of Name Disambiguation

In 1976, Robert E. Casey, the Recorder of Deeds of Cambria County, Pennsylvania, let his bartender talk him into running for State Treasurer. He didn't take the campaign very seriously, in fact, he went on vacation instead. Nonetheless, he easily defeated the party-endorsed candidate in the Democratic Primary and went on to win the general election. It seems that voters thought they were voting for Robert P. Casey, a popular former State Auditor General and future Governor.

Robert P. Casey almost won the Pennsylvania Lieutenant Governor's race in 1978. No, not that Robert P. Casey, this Robert P. Casey was a former teacher and ice cream salesman. Robert P. Casey, Jr., the son of the "real" Robert P. Casey, was elected to the United States Senate in 2006. Name disambiguation turns out to be optional in politics.

That's not to say ambiguous names don't cause real problems. My name is not very common, but still I occasionally get messages meant for another Eric Hellman. A web search on a more common name like "Jim Clark" will return results covering at least eight different Jim Clarks. You can often disambiguate the Jim Clarks based on their jobs or place of residence, but this doesn't always work. Co-authors of scholarly articles with very similar or even identical names are not so uncommon- think of father-son or husband-wife research teams.

The silliest mistake I made in developing an e-journal production system back when I didn't know it was hard was to incorrectly assume that authors were people. My system generated webpages from a database, and each author corresponded to a record in the database with the author's name, affiliations, and a unique key. Each article was linked to the author by unique key, and each article's title page was generated using the name from the author record. I also linked the author table to a database of cited references; authors could add their published papers to the database. Each author name was hyperlinked to a list of all the author's articles.

I was not the first to have this idea. In 1981, Kathryn M. Soukup and Silas E. Hammond of the Chemical Abstracts Service wrote:
If an author could be "registered" in some way, no matter how the author's name appeared in a paper, all papers by the author could automatically be collected in one place in the Author Indexes.

Here's what I did wrong: I supposed that each author should be able to specify how their name should be printed; I always wanted my name on scientific papers to be listed as "E. S. Hellman" so that I could easily look up my papers and citations in the Science Citation Index. I went a bit further, though. I reasoned that people (particularly women) sometimes changed their names, and if they did so, my ejournal publishing system would happily change all instances of their name to the new name. This was a big mistake. Once I realized that printed citations to old papers would break if I retroactively changed an author's name, I made author name immutable for each article, even when the person corresponding to the author changed her name.

Fifteen years later, my dream of a cross-publication author identifier may be coming true. In December, a group of organizations led by Thomson Reuters (owners of the Web of Knowledge service that is the descendent of the Science Citation Index) and the Nature Publishing Group announced (pdf, 15kB) the creation of an effort to create unique identifiers for scientific authors. Named ORCID, for Open Researcher & Contributor ID, the organization will try to turn Thomson Reuters' Researcher ID system into an open, self-sustaining non-profit service for the scholarly publishing, research and education communities.

This may prove to be more challenging than it sounds, both technically and organizationally. First, the technical challenges. There are basically three ways to attack the author name disambiguation problem: algorithmically, manually, and socially.

The algorithmic attack, which has long history, has been exploited on a large scale by Elsevier's SCOPUS service, so the participation of Elsevier in the ORCID project bodes well for its chances of success. Although this approach has gone a long way, algorithms have their limits. They tend to run out of gas when faced with sparse data; it's estimated that almost half of authors have their names appear only once on publications.

The manual approach to name disambiguation turns out not to be as simple as you might think. Thomson Reuters's ISI division has perhaps the longest experience with this problem, and the fact that they're leading the effort to open name disambiguation to their competitors suggests that they've not found any magic bullets. Neil R. Smalheiser and Vetle I. Torvik have published an excellent review of the entire field (Author Name Disambiguation, pdf 179K) which includes this assessment:
... manual disambiguation is a surprisingly hard and uncertain process, even on a small scale, and is entirely infeasible for common names. For example, in a recent study we chose 100 names of MEDLINE authors at random, and then a pair of articles was randomly chosen for each name; these pairs were disambiguated manually, using additional information as necessary and available (e.g., author or institutional homepages, the full-text of the articles, Community of Science profiles (http://www.cos.com), Google searches, etc.). Two different raters did the task separately. In over 1/3 of cases, it was not possible to be sure whether or not the two papers were written by the same individual. In a few cases, one rater said that the two papers were “definitely by different people” and the other said they were “definitely by the same person”!
(Can it be a coincidence that so much research in name disambiguation is authors by researchers with completely unambiguous names?)

The remaining approach to the author name problem is to involve the authoring community, which is the thrust of the ORCID project. Surely authors themselves know best how to disambiguate their names from others! There are difficulties with this approach, not the least of which is to convince a large majority of authors to participate in the system. That's why ORCID is being structured as a non-profit entity with participation from libraries, foundations and other organizations in addition to publishers.

In addition to the challenge of how to gain acceptance, there are innumerable niggling details that will have to be addressed. What privacy expectations will authors demand? How do you address publications by dead authors? How do you deal with fictitious names and pseudonyms? What effect will an author registry have on intellectual property rights? What control will authors have over their data? How do you prevent an author from claiming another's publications to improve their own publication record? How do you prevent phishing attacks? How should you deal with non-roman scripts and transliterations?

Perhaps the greatest unsolved problem for ORCID is its business model. If it is to be self-sustaining, it must have a source of revenue. The group charged with developing ORCID's business model are currently looking at memberships and grants as the most likely source of funds, recognizing that the necessity for broad author participation precludes author fees as a revenue source. ORCID commercial participants hope to use ORCID data to pull costs out of their own processes, to fuel social networks for authors or to drive new or existing information services. Libraries and reserch foundations hope to use ORCID data to improve information access, faculty rankings and grant administration processes. All of these applications will require that restrictions on the use of ORCID data must be minimal, limiting ORCID's ability to offer for-fee services. The business conundrum for ORCID is very similar to that faced by information producers who are considering publication of  Linked Open Data.

ORCID will need to navigate between the conflicting interests of its participants. CrossRef, which I've written about frequently, has frequently be cited as a possible model for the ORCID organization. (CrossRef has folded its Contributor ID project into ORCID.) The initial tensions among CrossRef's founders, which resulted from the differing interests of large and small publishers, primary and second publishers, and commercial and nonprofit publishers, may seem comparatively trivial when libraries, publishers, foundations and government agencies all try to find common purpose in ORCID.

It's worth imagining what an ORCID and Linked Data enabled citation might look like in ten years. In my article on linking architecture, I used this citation as an example:
D. C. Tsui, H. L. Störmer and A. C. Gossard, Phys. Rev. Lett. 48, 1559 (1982).
Ten years from now, that citation should have three embedded ORCID identifiers (and will arrive in a tweet!). My Linked Data enabled web browser will immediately link the ORCID ids to wikipedia identifiers for the three authors (as simulated by the links I've added). I'll be able find all the articles they wrote together or separately, and I'll be able to search all the articles they've written. My browser would immediately see that I'm friends with two of them on Facebook, and will give me a list of articles they've "Liked" in the last month.

You my find that vision to be utopian or nightmarish, but it will happen, ORCID or not.

More ORCID and author ID, and name disambiguation links:
Photo of the "real" Robert P Casey taken by Michael Casey, 1986, licensed under the Creative Commons Attribution 2.5 Generic license.
Reblog this post [with Zemanta]

Article any source

Wednesday, December 9, 2009

Supporting Attendance at Code4Lib

In the middle of a session at the Charleston Conference a month ago, I was in some keynote address about the future of libraries and the role of journals in scientific communication, and I got a bit fed up at a notion that scientists were some sort of exotic creatures that used libraries and information resources in ways that the library community needed to understand better. It occurred to me that a much better way to understand the needs of scholars was to just look around the room at the 300 people learning, communicating and synthesizing ideas with each other.

The Charleston Conference started in 1980 as a regional library acquisitions meeting with 24 attendees. This year was its 29th. It covered the world of scholarly information, library collections, preservation, pricing and archiving and it attracted well over 1000 publishers, vendors, librarians, electronic resource managers, and consultants from around the world. Its success is to a large extent the work of one person- Katina Strauch. Over the years, Katina's empire of hospitality has come to include print publications- Against the Grain and The Charleston Advisor, associated websites, and multiple blogs. The Charleston Conference has established itself as an important venue for many types of communication and learning; you might not call it scholarly communications, but so what?

Scientists and scholars aren't so different from librarians and publishers. They go to conferences, drink coffee and beer and learn in the sessions and in the hallways. They exchange business cards and send each other email. They tell stories about the experiments that failed. They gossip. The conferences provide them programs to take home and help them remember who said what. Occasionally someone mentions an article they found to be interesting, and everyone goes home to read it. The Charleston Conference and associated business properties has grown nicely into the internet age and would be an appropriate model for emulation by the scholarly communication community

Another vision for the future is provided by Code4Lib. Code4Lib started as a mailing list in 2003 as a forum for discussion of
all thing programming code for libraries. This is a place to
discuss particular programming languages such as Java or Python,
but is also provide a place to discuss the issues of programming
in libraries in general.
At first, it grew slowly, but people quickly discovered how useful it was. Today it has almost 1,300 recipients and a very high signal to noise ratio.

In 2006, the first Code4Lib Conference was held at Oregon State University. The conference was inspired to some extend by the success of a similar conference, ACCESS, held every year in Canada. The Code4Lib Conference has always been self-organizing (organizationless, you might say), and has been quite successful. Presentations are selected by vote of potential attendees; participation is strongly encouraged using lightning talks and unconference sessions. The conference has tried to stay small and participatory, and as a result, registrations quickly fill up.

Code4Lib is also instantiated as channels of communication such as an IRC channel and a Journal, and the community never seems to fear trying new things. In many ways, it's still in its infancy; one wonders what it will look like if it ever gets to be as long-established as Charleston.

This February, the fifth Code4Lib Conference will take place in Asheville, North Carolina. I hope to be there. But with the "Global Economic Downturn" and library budgets being slashed, I worry that some people who might have a lot to contribute and the most to gain may be unable to go due to having lost their job or being in a library with horrific budget cuts. So, together with Eric Lease Morgan (who has been involved with Code4Lib from that very first eMail) I'm putting up a bit of money to support the expenses of people who want to go to Code4Lib this year. If other donors can join Eric and myself, that would be wonderful, but so far I'm guessing that together we can support the travel expenses of two relatively frugal people.

If you would like to be considered, please send me an email as soon as possible, and before I wake up on Monday, December 14 at the latest. Please describe your economic hardship, your travel budget, and what you hope to get from the conference. Eric and I will use arbitrary and uncertain methods to decide who to support, and we'll inform you of our decision in time for you to register or not on Wednesday December 16, when registration opens.

If you want to help us with a matching contribution, it's not required to be named Eric.

Update: Michael Giarlo and one other member of the Code4Lib community have agreed to match, so it looks like we have enough to support 3 attendees.
Article any source

Friday, November 13, 2009

The New York Times Gets It Right; Does Linked Data Need a CrossRef or an InfoChimps?

I've been saying this long enough that I don't remember whether I was quoting someone else: whenever the internet disintermediates a middleman, two new intermediaries pop up somewhere else. It's disintermediation whack-a-mole, if you will. The reasons for this are:
  1. The old middlemen became fat on mark-ups an order of magnitude larger than needed by internet-enabled middlemen.
  2. Internet-enabled middlemen add value in ways that the old ones didn't.
My last business functioned as an intermediary that aggregated linking data. We'd get data from publishers, clean it up and add it to our collection, then provide feeds of that data to our customers (libraries and library systems vendors). Our customers got good data and support if was a problem. The companies who provided the data didn't have to deal with hundreds of libraries or system vendors, and they came to understand that we would help their customers link to their content.

Some companies, especially the large ones, were initially uncomfortable with the knowledge that we were selling feeds of data that they were giving out for free. They felt that somehow there was money left on the table. Other companies were fearful of losing control of the information, even though they didn't really have control of it in the first place. Once we explained to them how their data contained mangled character encodings, fictitious identifiers, stray column separators and Catalan month names, they began to see the value we provided.

While my company focused on the data needs of libraries (and did pretty well), a group of the largest academic publishers put up some money and formed a consortium to pool a different type of linking data in a way that let the publishers have more control of the data distribution. This consortium, known as Crossref, just celebrated its 10th anniversary. Crossref has not only paid back the money that its founders invested in it; it has arguably done more to push academic publishing into the 21st century than any other organization on the planet.

As academic publishing companies began to understand the benefits of distributing linking data through Crossref, my company, and others like it, they became more comfortable opening up their content and reaping the financial benefits. Despite the global recession, and despite predictions of its impending collapse, STM publishing has been financially healthy with companies such as Elsevier reporting increased profits. This is rather unlike the newspaper industry, for example.

Before I get to the newspaper industry, I should note yesterday's news that InfoChimps are publishing a collection of token data harvested from Twitter.
Today we are publishing a few items collected from our large scrape of Twitter’s API. The data was collected, cleaned, and packaged over twelve months and contains almost the entire history of Twitter: 35 million users, one billion relationships, and half a billion Tweets, reaching back to March 2006.
InfoChimps is positioning itself as a marketplace to buy, sell, and share data sets of any size, topic or format. Yet another intermediary has popped up!

Two weeks ago, I wrote a somewhat alarmist article about problems in an exciting set of Linked Data being released by the New York Times. I am pleased to be able to be report that the New York Times is now getting it right! The most important thing that they're doing right is that they're listening to the people who want to consume their data. They've started a Google Group based community for the specific purpose of understanding how best to deliver their data. They've also corrected the problems pointed out by myself and others. It's not perfect, but it's not reasonable to expect perfect. The New York Times has set a very hopeful example for other companies that want to start publishing semantic linking information on the open web.

If, as many of us hope, many publishers decide to follow the lead of the Times and make more data collections available, will more intermediaries such as InfoChimps arise to facilitate data distribution, as happened with linking data in scholarly publishing? Will ad hoc groups such as "the Pedantic Web" become key participants in a less centralized data distribution environment? Or maybe large companies will turn off the spigots as "the suits" grow increasingly worried about their ability to control data once it is let out into the web of data.

Perhaps the time is ripe for a set of forward-looking publishers to emulate the nervous-but-smart journal publishers who started Crossref 10 years ago and start a similar consortium for the distribution of Linked Data.
Reblog this post [with Zemanta]

Article any source

Monday, July 13, 2009

Dung Beetle Armament and the Real Threats to Scientific Publishing

To illustrate an article on dung beetle armament, the New York Times Science section published a graphic with a spectacular montage of 35 animals with grotesque armaments, ranging from the Narwhal to the Giraffe Weevil. 13 of them are extinct. The reason that many dung beetles have evolved such elaborate armaments is not so much that they are effective in combat with other dung beetles, but rather that female dung beetles select mates based upon the outward display of combat fitness.

In my last post, I argued that scholarly publishers were not being threatened by imminent disruption by the same factors that have the newspaper publishing industry on the brink. I suggested that a potential vulnerability of the scholarly publishing industry would be the disintegration of the linkage between the industry's activity, publishing scholarly articles, and the industry's main revenue source- library subscriptions. I see two possible ways that this could occur. It could occur through a collapse of library funding; I hope to discuss that in a future post. This post discusses another way this could occur: I think there is a possibility that the adoption of social networking technologies will lead to a collapse of scholarly publishing as it exists today.

If this sounds a bit far-fetched, consider the parallels between scholarly publishing and dung beetle armament. The development of scholarly publishing today is driven by the selections made by authors about where and how to publish articles. The authors ultimate goal is to propagate their work and thus gain tenure, status and funding, just as "the ultimate goal" of the female dung beetle is to gain a safe tunnel to enjoy dung and raise baby dung beetles. The authors do not really know which journals do the best job of propagating their work, but they recognize prestige and the badges of prestige, and they know what sorts of publications will look best to their tenure committees. Authors do not consider the cost of journals any more than female dung beetles consider the energy cost of male armature. The size and form of today's scholarly publishing ecosystem is thus driven to a significant extent by the superficial judgments of tenure committees.

Anything that might change the way tenure committees, and thus authors, perceive journal publication has the potential to reverse the fortunes of journal publishers. To my mind, social networking technologies have that potential as do few other other things on the horizon. The reason is that tenure decisions have used journal publication records as objective measures of a candidate's social status within the scientific community. Publication in a prestigious journal has been an important way for scholars to become known, to gain speaking invitations, and to advance ideas. But publications are only part of this process. Knowing the right people, studying with the right professors, schmoozing at conferences, all of these are probably more important to the advancement of new ideas, but they have been very hard to measure in any objective way.

Social network technologies open new possibilities for the propagation of new ideas and for the assessment the impact of those ideas. Already, we see people using the number of followers they have on Twitter or the number of recommendations they have on LinkedIn as measures of social status, so it's not much of a stretch to imagine that similar measures could be used to evaluate young academics or to award grants. It's beyond dispute that Twitter is already being widely used to propagate links to interesting technical papers and posts on scientific subjects. If targeted development of social network-based evaluation methodologies were pursued by groups such as the library community who wish to re-inject usage and low-cost access into the tenure equation, the competitive environment for scholarly and scientific publishers could change radically.

Every threat is an opportunity, of course, and it's equally possible that social networking technologies could reinforce the scholarly publishing industry- after all, dung beetle armaments evolve to adapt to changing fashion choices among female dung beetles. A potential weakness- for example, the unwillingness of people to post or retweet links to subscriber-only content, could turn into strengths if publishers develop access models that grant special access for re-tweeted links or an author's Facebook friends. Publishers could also try to ward off challengers in the scholar evaluation game by developing improved and more ostentatious badges of honor- best paper prizes, awards for the most forwarded paper, etc.

In fact, I've come up with a mathematical model for how all this will evolve. First, assume a spherical dung beetle...
Article any source

Friday, July 10, 2009

Spherical Livestock and the Alleged Disruption of Scientific Publishing

Physicists have a joke about "spherical cow approximations" referring to their tendency to simplify a problem to make calculations easier, even though such simplifications bring into question the solution's application to reality. My favorite version of the joke, which I first heard directly from Hans Bethe, has Nikita Khrushchev asking his most elite scientists to help the Soviet Union with its difficulty meeting its five year plan for the dairy industry. The biologists and the chemists are completely stumped by the problems of increasing milk production, but the physicists proudly announce they have solved the milk production problem, but only for the case of spherical cows.

In a post entitled "Is scientific publishing about to be disrupted?", quantum information theorist Michael Nielsen describes what he thinks is a general explanation for why businesses and industries fail, and goes on to draw an analogy between the newspaper industry and the scientific publishing industry. Although the post is well written and highly entertaining, (I find his discussion of "immune systems" particularly delicious) I find part of his analysis to be even worse than a spherical cow approximation- he's trying to study milk production by analyzing the spherical chicken! Let me explain.

Nielsens "spherical chicken" is illustrated in this graph from his blog:

In the graph, he plots some sort of measure of success versus some sort of configuration parameter that presumably could be tuned to turn the New York Times into TechCrunch, or vice versa. He goes on to say that
The problem is that your newspaper has an organizational architecture which is, to use the physicists’ phrase, a local optimum. Relatively small changes to that architecture - like firing your photographers - don’t make your situation better, they make it worse. So you’re stuck gazing over at TechCrunch, who is at an even better local optimum, a local optimum that could not have existed twenty years ago
The problem with this analysis is that TechCrunch is completely immaterial to the difficulties that the newspaper industry is undergoing. The financial health of the New York Times and the newspaper industry is not being undermined by news blogs, it's being undermined by non-news sites such as Craigslist, Zillow, and the internet as a whole. Craigslist has focused on classified ads, and only classified ads, and unburdened by the expense of producing the rest of a newspaper, it is able to provide a much more effective solution for the classified advertiser. Zillow has done the same thing in the real estate advertising category. Another big revenue source for newspapers is display advertising to consumers. But nowadays, when someone wants to buy something or find a service, their first thought is to go directly to the internet. Want to find when a movie is playing? You used to pull out a newspaper, now you go to the internet. A company like BestBuy used to communicate with customers through newspaper ads; while they still do so to some extent, the internet allows them to communicate directly with consumers through their web site. None of the newspapers' real competitors are in the news business at all, and there is no configuration parameter of any sort that could be tuned to transform the New York Times into Craigslist.

The news industry's core problem is not, as Nielsen suggests, their inability to adopt disruptive technologies, but rather the disintegration of the linkage between their main activity and their revenue streams. In the past, good news would attract readership, and readership would attract advertisers. The biggest difficulty for newspapers today is not so much the loss of readership, it's that advertisers now have many more ways to connect to that readership. In applying the lessons of the newspaper industry to the evolution of the scientific publishing industry, it's the stability of activity-revenue linkage that needs to be closely examined.

Even a cursory look at the scholarly publishing industry reveals a very different situation from that of the newspaper industry. First of all, there is much more business-model diversity in scholarly publishing. There are huge companies like Elsevier competing with cottage companies which produce a single journal. There are large non-profit societies such as the American Physical Society that produce extremely cost effective journals and who make much of their content available for free. There are journals that have long survived primarily on advertising and journals that have long survived primarily on society member dues. There is also a lot of experimentation with business models going on, including author-paid open access publishers, and mixed "open choice" business models. This business model diversity gives scientific publishing industry robustness against the prospect of any one business model being severely disrupted. In addition, the transition to digital delivery which is giving the newspaper industry such difficulty is to a significant extent already being accomplished in the journal publishing industry.

The scientific publishing industry does have a similar activity-revenue linkage problem that it needs to pay attention to. The people who write the biggest checks to scientific publishers are institutional libraries. But scientific journals, for the most part, do not cater to libraries, they cater to author communities, because the biggest determinant of a scientific journal's success has been the quality and quantity of articles it is able to attract. As long as libraries continue to be attracted to the authorship attracted by journals, and continue to attract the institutional funding they need to support their subscription, the biggest revenue stream for scientific publishers will be secure. But suppose that institutions start deciding to outsource their libraries or begin to require researchers to directly fund their journal subscriptions? Or suppose that libraries are successful in attracting authors directly into open-access institutional repositories?

A better analogy from physics for the scholarly publishing business might be the polaron. A polaron is the combination of a particle and interactions with the environment that it moves in, and the combination has a mass significantly larger that the "bare" particle moving on its own. In the case of the scientific publishing business, the interactions with its environment include the way tenure committees rely on the prestige of a journal that has published a candidates work, or the way accreditation boards require libraries to subscribe to certain numbers of journals. The polaronic industry thus gains mass and inertia, allowing it continue longer than it might otherwise do. Computer operating systems work in the same way- they induce the creation of third party software that interact with the operating system and thus increase its mass and inertia in the market.

Strongly interacting polarons can distort their environments so much that the become trapped by their cloud of interactions- think of a celebrity trying to walk though a crowd of fans. For a business this can be a fatal situation if objectives change, and there is no possibility to adapt.

How's that for a spherical cow?


Article any source