Showing posts with label Book Use. Show all posts
Showing posts with label Book Use. Show all posts

Monday, August 12, 2013

A Rational Framework for Library eBook Licensing

Since the Redigi decision made it clear that there is no right of first sale for digital content in the US, it's been much easier to think up realistic doomsday scenarios for public libraries in the US. Why should a publisher let a public library lend an ebook if Amazon or some other competitor were to offer much better terms? How would our public library system, saddled with difficult-to-use systems and unfavorable contracts, ever hope to compete?

Back when HarperCollins first announced that it would only let libraries lend their ebooks 26 times before they would expire, there was widespread outrage from the library community. Looking back on that, it seems pretty clear that a lack of consultation and poor customer communication fueled the furor. By itself, the lending limit could have terrible long-term consequences for libraries, but as part of a wider, well-thought out framework, it could be useful component.

I've been doing a lot of thinking about this over the last 3 years, and I've decided it's time to float a comprehensive proposal for how libraries and publishers might work together on ebook distribution to benefit the entire reading ecosystem. eBook lending as implemented to date has been founded on a combination of irrational fears and outmoded processes. We deserve better.

Behind this framework is a set of assumptions.
  1. Library ebook distribution must sustain and increase the total population of readers; this is a prerequisite for a healthy book publishing industry.
  2. Patron discovery of ebooks in libraries must connect effectively to ebook sales.
  3. Library distribution must become much more efficient, and overhead must become much smaller for ebooks than it is today for print books and ebooks.
  4. Long term preservation of ebook availability must be a joint undertaking of libraries and publishers.
  5. The economic models used for library ebook distribution must provide incentives for libraries and publishers to promote points 1-4.
I don't pretend that people won't disagree with some or all of these 5 assumptions, but if any of them are false, then, I think there will be NO distribution of ebooks through libraries. I also recognize that not all books are alike; even if library distribution works for some ebooks, it's unlikely that it will work for every ebook.

So the fifth assumption is what this post is really about. Given 1-4, what should an economic framework look like? Here are the features of a model that makes sense to me:
  1. Decoupled pricing. An ebook license that allows for lending makes the ebook more valuable, so why shouldn't it cost more than an individual, non-transferable license? I can't say whether Random House's 300% markup for libraries is excessive, but why not let the marketplace decide? For new, super-popular ebooks, maybe 500% markup makes sense. On the other hand, maybe ebooks that need exposure should have an 80% markdown because libraries might turn them into bestsellers.
  2. Rate limits instead of DRM. Patron license embedding.  I've written about this before. This may take the most convincing, but in thinking about the imperatives of effective discovery, low distribution overhead, and long-term preservation, I've concluded that there are no alternatives to major change in library distribution technology.
  3. Circulation charges after an initial period. Most books are bought in the first year of publication. Today, libraries "deaccession" books to match their declining demand. But there's no reason for a library to deaccession an ebook, so for most books the global supply for any given ebook will eventually exceed global demand. If the library can cut its transaction cost from ~$2 per circulation to $0.20 per circulation it seems fair to reward the publisher with part of the difference for developing books with long term value. 
  4. License transferability/InterLibrary Loan. Libraries rely on interlibrary loan to expand the scope of their collections and meet special needs. But ebook loans can be instantaneous, so digital ILL can compete directly with backlist sales. If the transaction costs (currently ~$10) for ILL can be squeezed down to $1 or so, there's plenty of margin to provide a transaction payment to the rights holder for the privilege of doing so. 
  5. Patron-funded purchases. Libraries are tight on funding even as they need to completely transform what they do. Their biggest asset is a huge reservoir of public goodwill. At this pivotal juncture, their ebook offerings are characterized by long hold queues. Why can't a library patron buy an extra copy for the library and jump to the front of the queue? Why don't publishers offer "Buy for your Library" buttons on their catalog pages? The reasons are complex, but it's mostly a case of "we haven't done that before". But if it doesn't happen I just can't fathom how library discovery can effectively plug into publisher commerce.
  6. License durability. If libraries are expected to "buy" ebooks, it should be pretty much for keeps. If the publisher for some reason has to revoke a license without cause, the library should get a refund of the license price.
  7. Archival copies. Libraries need to do a lot of things with books other than lending. Indexing and archiving are good examples. The saddest thing about the most successful library ebook distributors today is that libraries don't get access to unencrypted ebook files. If libraries are to offer effective discovery and archiving of ebooks, they need access to the files. Seems a no-brainer to me.
There are a bunch of parameters to plug into this framework; here's my guess as to what they should be:
  • Rate limits: One authenticated user per two weeks.
  • Circulation fee: $0 for the first year, after the first year, 2% of purchase price or $1 whichever is greater. 
  • ILL fee (publisher share): 5% of purchase price or $2, whichever is greater. 

A rational ebook lending framework would mean big changes for both the book publishing industry and the library industry. Even if a HarperCollins decided today that this was an attractive way forward, it would be hard-pressed to find a way to implement it, because libraries just don't work that way. So it seems a bit far-fetched at this point. Based on the iBookstore fiasco, it appears to be illegal for big publishers to even talk to each other, let alone drive business model changes. It's good that a library group is still trying to figure it out.

Maybe some small startup company could try some sort of pilot program.


Article any source

Saturday, August 3, 2013

Wattpad Usage is in the Ballpark of US Public Libraries

Billion Reasons Why, a novel
by xXdemolitionloverXx
on Wattpad
While working on another article, I came across this bit of data. Wattpad, the reading and writing community that's sort of a YouTube for stories, claims that its users are spending 3.5 billion minutes per month on the site. That's a number so big that I had no context for it.

So I wondered, how many minutes per month do people spend in their public libraries? There's a lot of data available for US public libraries from IMLS. In 2010, the most recent year for which data is available, 1.57 billion visits were made to US public libraries, or about 131 million visits per month. I have no idea how long an average library visit lasts, but let's say it's a half hour, then the total minutes of "user engagement" by US public libraries would be about 3.9 billion minutes per month. Roughly the same as Wattpad.

Maybe we should also count the time that readers sped at home with a library book, 30 minutes might be a serious underestimate. (see update) Also, Wattpad's usage is spread out internationally- they are the top mobile site in the Phillipines, for example. So its usage within the US is probably quite a bit less than public libraries. But it's also concentrated in certain demographics- teenage girls, for example. And it continues to grow at a solid pace.

Update: Karen Coyle point out in comments that you could also estimate library user engagement by looking at circulations. By that measure, assuming an average of 4 hours of reading per book, you get that US public libraries are about 8 Wattpads of engagement.

Any way you look at it, that's a lot of reading going on.
Enhanced by Zemanta

Article any source

Tuesday, May 14, 2013

Hack the Publishing Hackathon

Why a publishing hackathon?
Book discovery needs innovation. It’s never been easier to get a book into a reader’s hands—just one click. But, with over 10,000 books published each year on every topic imaginable, how do people find out about them? There are fewer bookstores to help readers discover exciting new authors and ideas. There’s currently no digital experience that replicates the serendipity of browsing bookshelves. Recommendation engines are fairly primitive – they know what you bought, but they don’t know why. It’s a disruptive opportunity that hasn’t been explored.
Seriously, the sponsors of this event don't think book discovery has been explored? I guess they were too busy suing Mr. Google to notice that Google Books is a pretty good discovery tool. I suppose they never thought to ask Mr. Wikipedia how many books are published every year.

All in all, I find the description of this hackathon INSULTING to just about every developer that's worked in the general vicinity of the book industry.

Umm. Mr. Steinberger. If you and Perseus really want to promote discovery innovation, then perhaps you have heard of Goodreads? They're driving some decent discovery of books. Maybe it doesn't count if Mr. Amazon is buying them. Perhaps you've heard of Amazon? They popularized the "If you liked this, maybe you'll like..." feature that everyone in the publishing industry tries to copy. If you don't like Goodreads, maybe I can introduce you to LibraryThing, which has been driving valuable book discovery in more ways than I can list here. I know that "library" in their name is a big turnoff for your big 6 colleagues, but libraries are huge book discovery machines. I don't suppose you want them to disrupt anything. And umm DP.LA????

People mostly discover books by word of mouth. Some  innovators promoting social reading include Readmill (who had their own publishing hackathon) and (giving props to the NYC home team) ReadSocial and the stuff Bob Stein has been exploring. And Kobo, Copia and Zola are doing some amazing things to integrate book discovery with ebook selling and reading environments. I've written previously about Jellybooks' fresh approach to discovery.

And some more on libraries. When I was at OCLC, we worked on real simple problems like "how do you discover the other editions of the same book?" and we found that publishers had NO CLUE what they'd published 5 years previous. So yeah, we did our bit.

But I'm coming to the hackathon anyway. because despite the ridiculous framing, this event has some clueful backers. NYPL for one. Small Demons for two. And they're even wasting prize money on a new age library metadata thingy. (I might be wrong about the wasting part.)

I'm hoping that some people will be interested in rethinking ebook front matter. Unglue.it needs books to work better all by themselves. The best discovery instrument for a book is the GDMF book, to my mind. So let the book do some work. With a little javascript. And no more DRM, thank you very much!
Enhanced by Zemanta

Article any source

Friday, December 16, 2011

How to Dig for Book Data Treasure


To me, surest indicator of an impending doom for book publishing is hearing a publisher cite the advertising of Attributor, an anti-piracy solutions company, as if it were science. It's not the attitude towards piracy that bothers me, that's entirely sensible. It's the implied devaluation of honest data that depresses me.

There's hope though. I've gotten to know quite a number of people throughout the reading ecosystem with whom I can use the word "data" as high praise, roughly equivalent to the word "gold". If you're reading this, chances are you're a member of this secret society, and what follows is a sketch of a treasure map.

In a recent post, I promised to suggest ways that we might measure the effects of library ebook lending on book sales. If you think about it, there are many parallels between attempting such a measurement and previous studies that have tried to measure the effect of ebook piracy on book sales. Unfortunately, the only objective study I know of was a small study done by Brian O'Leary, and the effects observed in that study were small and in a direction counter to popular narratives (and thus rarely noted in the sort of presentations that cite Attributor advertising).

In that study, O'Leary looked for time-domain correlations between sales figures for books from two publishers and the appearance of the same books on BitTorrent. A similar study focused on library Lending could be much more compelling, because library circulation data is a much more direct measure of distribution than any sort of torrent tracking, and librarians are much better than pirates at sharing data.

With the cooperation of booksellers, library circulation and holdings could be compared and correlated to store-by-store sales. For example, you could look at a book that's held in a significant fraction of libraries and look for correlations (positive AND negative) between areas where a library is circulating the book and stores where the book is selling. You've have to remove regional and demographic variance, of course, but with enough data, almost anything is possible.

With the cooperation of a large publisher, rigorous experiments could be done. Scientific experiments derive rigor from the use of controls. To prove that lending influences sales, it's not enough to do lending and look for sales. A rigorous experiment would have both a trial where books are lent and an identical trial where the same books are not lent.

One way to control a lending experiment would be to make a random selection of a publisher's catalog available for lending. Imagine if Penguin had worked with the library community on an experimental withholding of a random part of its catalog from Overdrive. The sales could be analyzed for patterns and trends.

It's important that data analysis of this sort be done objectively by researchers with integrity. In any large collection of data, it's possible to focus on data which supports one narrative over another. If lending-sales studies were done, my guess is that some types of books would show correlations very different from others.

I've used the word "cooperation" several times already. I'm not so naïve as to think that data sharing will materialize out of thin air. Perhaps the sort of eco-system wide organization envisaged by the same Brian O'Leary could be the vehicle to make data treasure digging possible. Opportunity in Abundance for the win!

Enhanced by Zemanta

Article any source

Friday, December 9, 2011

Book Lending Ignorance


To what degree does library book lending complement book sales, and to what degree does library lending substitute for book sales? I don't think anyone knows for sure. (Well maybe Amazon, but they're not telling.)

With over 40 billion dollars per year of sales at stake, you would think that the US book publishing industry would want to know as much as possible about how those sales are generated. Since US public libraries circulate more items than US bookstores sell, the industry needs to understand the role of libraries in getting people to read and purchase books. Is it small or big? Does the existence of libraries promote sales or hurt sales? How do the equations change when books become digital?

Publishers do a pretty good job of compiling sales data, and they spend a lot of money to figure out what books are selling and who's buying them. According to BookStats, a cooperative study by the AAP and BISG, Americans bought an average of 7.32 books in 2010.

On the library side, there's a bunch of interesting data. IMLS has been compiling a wealth of data about the footprint of public libraries, which is why I can tell you that the average American borrowed 8.1 items from public libraries in 2009. Library Journal has recently published the first installment of results from a fascinating survey of library patrons. (Aside: this study should be made available in every library!) They find that 46% of respondents use the public library less than 2 times per year.

The LJ Patron Profiles survey shows a strong relationship between library use and book purchasing. For example, over half of survey respondents report buying a book by an author whose works they'd previously borrowed from the library. That's a huge number, considering that 20% of respondent never go to the library, period. At the same time the survey indicates a competition between reading and borrowing. Respondents who report that they've decreased their use of libraries buy 12.18 books per year, while those who've increased their library usage buy only 10.9 books per year. What we can't tell from the data is cause and effect. With the recession having a wide impact, who's to know whether the folks showing up more at libraries might buy even fewer books if the libraries weren't around!

It costs about 11 billion dollars a year to run public libraries in the US, and libraries work hard to demonstrate their value to the communities that support them. They compile data to measure their activity and the community's return on their investment in libraries. These studies assign much of the benefit of library spending to substitutional activity. For example, a survey by Denver Public Library determined in 2009 that it saved its community $105 million based on the cost to use alternative sources of information, and delivered an additional $5 million by avoiding "lost use", activity that wouldn't have occurred if the library did not exist. (See Public Libraries- A Wise Investment (PDF, 1.4 MB) from Library Research Service)

Do libraries really believe that 91% of their circulations would have resulted in purchases if they didn't exist? There's no hard evidence anywhere that that's true. Every librarian can tell you about patrons who loved a book so much they went and bought the whole series, but there are also users who never buy a book they can get in the library. And what about those readers who never go to the library? Surveys are a cheap way to collect data, but they often don't reflect the real behavior of the people surveyed.

So much is unknown, and so much is to be gained by knowing more. What hasn't been done, as far as I know, is to try to compare and correlate hard data on book sales and library lending in any meaningful way. In my next post, I'll describe how a cross-industry cooperative approach to book data collection and analysis might provide some light amid the gloom of the reading industry's winter solstice of understanding.

Article any source

Sunday, March 27, 2011

Statistician Can't Distinguish Library Patrons from Monkeys

If you're a librarian nodding at the title, no, that's not what I mean.

The statistician in question is Carnegie-Mellon Statistics Professor Cosma Shalizi. He's made a habit of debunking claims by physicists, economists, and computer scientists that their data shows power-law behavior in this-or-that system. When I say he can't distinguish library patrons from monkeys, I don't mean that Prof. Shalizi is near-sighted or that he's unfamiliar with the grooming habits of library patrons. I mean that Shalizi is arguing that the distribution of book circulation that I wrote about two weeks ago can be explained by completely random processes.

In his comment on my blog post, Shalizi reanalyzed the circulation data from University of Huddersfield and shows that it can be fit well by a "log-normal" distribution, and that the very high-usage tail of the Huddersfield data is not consistent with a power law (such as the one I gave in my post). I've confirmed  his analysis, which went much farther into the high-usage tail than my first pass. This is done by looking at the cumulative distribution, i.e. plotting the number of books that have circulated less than a certain number of times.

If you want to make the connection to the monkeys in the library, it's important to understand the generating mechanisms that lead to log-normal distributions. These often arise from random growth processes, and are just like the standard "bell-curve", but on a log scale.

Here's how a random growth process could apply to book use. Let's suppose that every day, everyone who has read a book flips a coin. If heads, they do nothing. If tails they try to get someone else to also read the book. The group of people that has read the book thus grows by some percentage. Repeating this process over and over causes the book's usage to grow randomly. If we then measure the  size of these groups, the readership sizes will follow a log-normal distribution.

There's a saying among experimental physicists. "Keep taking data until you have enough to write an article for Physical Review Letters. Then stop taking data." In my previous post on book use, I violated this rule by asking other libraries to share their circulation data for analysis. Ross Riker at Goshen Public Library in Indiana stepped up to the challenge.

Goshen has accumulated circulation data since their automation system was installed in 1996. Riker sent me the number of times each of 144,269 items currently held had been circulated, for a total of 3.04 million circulation events. I've plotted the data on the graph below, alongside the Huddersfield data. It looks somewhat different, doesn't it? I sent the Goshen data to Shalizi, and his analysis was that neither log-normal or power-law distributions could fit the data.

Is book use in an American public library governed by different principles from that in a British academic library? Probably not. I noticed that the maximum number of circulations at Goshen was 251. The standard circulation period at Goshen is 3 weeks, so there's one book that's been checked out for 14.44 years solid, or since late 1996, which is about when Goshen began collecting data.

If we want to look at book use, what we should be plotting is the the rate at which the book is being circulated. That's equal to the number of circulations divided by the time the book is actually on the shelf.

After applying time-on-shelf corrections, the data from both Goshen and Huddersfield are well fit by log-normal distributions. To compare the Huddersfield data to the Goshen data, we need to take into consideration another difference. The Goshen data is listed item by item, so if there were two copies of a book, they count as two items. The Huddersfield data groups the circulation counts for all copies of the same book. To properly compute the time-on-shelf factor, I adjusted the circulation rate based on the number of copies held for each book.

After applying the appropriate corrections, the resulting distributions (below) are amazingly similar for the two libraries, and fit beautifully to log-normal distributions. Both distributions even have a bulge at the very highest circulation rates. At Huddersfield, inspection of the relevant bulge items suggests that they're texts used in particular courses, and have circulation times shorter than the main collection.

You may be disappointed to learn that the distribution of book use can be explained by random processes without reference to metadata quality, selection efficiency, or discovery system details. Nor does it derive from a power law characterizing the structure of user networks or citation graphs. All of the circulation distribution data I've looked at is consistent with there being one driving force in the distribution of book use. The non-technical term for this driving force: word of mouth.

Maybe I should stop taking data.


Notes:
  1. The formula for a log-normal distribution is:
    where μ and σ are the mean and variance of the logarithm of the distribution. If you use Excel, the lognormal distribution is built-in:   LOGNORMAL(f,μ,σ,FALSE). (TRUE gives the cumulative distribution function)
  2. It's not surprising that you get a better fit with a log-normal distribution than with a power-law. The log-normal distribution gives you an extra fitting parameter, after all. But when you include the full high usage tail, the power law predicts a lot more extremely high usage books than is observed.
  3. The time-on-shelf correction has a bit of fudge-factor in it. If the standard circulation period is 3 weeks, that doesn't mean that every user keeps it for three weeks, or that the book gets reshelved immediately after 21 days. My fit uses an average time-off-shelf period of 18 days.
  4. My log-normal fit for Goshen has a mean of 2.95 and a sigma of 0.94. For Huddersfield, I get a mean of 2.22 and a sigma of 0.77 after conversion to item data. The larger mean gives the higher circulation per item at Goshen. Feel free to speculate about the sigmas.
  5. The titles with the highest per-copy circulation rates at Huddersfield are:
    • Music in medieval Europe
    • An introduction to business ethics
    • A guide to the harpsichord
    • On humour : its nature and its place in modern society
    • Japan
    • The BBC and public service broadcasting
    • Authenticity in performance : eighteenth-century case studies
    • Mozart's Requiem : on preparing a new edition.
    • Asia's next giant : South Korea and late industrialization
    • Handel's operas : 1704-1726
    • Cognitive psychology : a student's handbook
  6. Raw data sets are available for Huddersfield and Goshen.
  7. For a readable discussion of generating mechanisms for power laws and log-normal distribution, I recommend "A Brief History of Generative Models for Power Law and Lognormal Distributions" by Michael Mitzenmacher, Internet Mathematics Vol. 1, No. 2: 226-251. [PDF 382KB].
  8. My comments re Harper-Collins are unaffected by this re-analysis, but my quantitative modeling of the budget impact of the new ebook policy will change a bit.
  9. Monkeys aren't really random, but I bet if one started reading a book, there would soon be a crowd of monkeys wanting to read the same book!
Enhanced by Zemanta

Article any source

Tuesday, March 15, 2011

Help Me Study the Physics of Book Use

I am not a librarian. I'm not a bookseller. I'll admit to some librarian tendencies- when I was little, I liked to line up my trucks and sort them from biggest to smallest. But my education and training was in engineering and physics. My approach to the analysis of data is that of a scientist. So when I analyzed the distribution of circulation across the collection of University of Huddersfield, I treated the data as a window into the physics of book use.

"Physics???" you may be thinking to yourself. Yes, physics. Well, maybe it would be economics if I had gotten past Econ 101 in college. But I feel comfortable with physics- I have 76 published articles to fall back on. Physics tries to describe things that happen in terms of simpler phenomena. It aims to connect observables (thing you can measure) to their root causes, and then uses that understanding to predict other observables. It doesn't matter so much whether the basic event is one particle hitting another, or one patron checking out a book, if broad patterns can be observed in these events, then a physicist can measure the patterns and try to deduce the causes.

That's why I was so excited to observe a power-law dependence in book-circulation frequency when I analyzed the data made available by the University of Huddersfield. In 15 years of research into crystal growth and electronic properties of semiconductors and superconductors, I never worked with such a well-behaved set of measurements. And as a physicist, I'm trained to believe that when a measured quantity obeys a mathematical relationship, then there must be a reason for it, even if I don't understand that reason yet.

Right now, I don't know why the book circulation in the Huddersfield library obeys a power law. A physicist would call this power law "phenomenology". Without an understanding of how it arises, I can't say whether it should apply to other libraries. I can't say if it would apply to ebook sales at Amazon, or holdings in Worldcat. It might be an accident. But it would be really cool if it was real, because at the core, it must be connected to how people choose things to read.

What causes people to buy a particular book, or borrow a particular book from a library? You would think that many people might want to know. Publishers and librarians might answer that books are read because they're good. But is there any concrete evidence that book quality has anything to do with sales or circulation? Ask any author if sales are correlated to quality, and they'll tell you about a wonderful book that nobody has bought or read. So maybe other factors  are more important.

A lot of recent discussion has revolved around the unproven hypothesis that library circulation leads to increased sales. The evidence cited, though compelling, is anecdotal and non-quantitative:
Eat, Pray, Love: One Woman's Search for Everything Across Italy, India and IndonesiaPenguin’s runaway hit, Eat, Pray, Love (Viking), was published in February 2006 with an initial run of 30,000 hardcover copies. The title didn’t become a bestseller until March 2007. In the meantime, copies of Eat, Pray, Love changed hands thousands of times through book clubs and libraries, scoring rave reviews from Library Journal and stirring up chatter among leading library blogs such as Memphis Public Library and San Mateo Public Library. Thanks to word-of-mouth marketing and library lending, when the paperback hit newsstands, Eat, Pray, Love sales skyrocketed.

It would be useful to really know how important this factor is.

I'm guessing that the power law I observed has very little to do with distributions of book quality and much more to do with how people are distributed and connected to each other- for example, city sizes are well described by a power law. I think that people pick books to read based mostly on what other people have read. That's what creates a best-seller. By studying the distribution of book usage, we may be able to prove that this is so.

So here's where I need help. We need to have more data sets to look at. If the power-law behavior is universal, it should show up in a wide variety of circulation statistics.

There are also situations where the power-law won't apply. It may seem odd to say this, since we don't understand where the power law comes from in the first place, but there are things it CAN'T do. For example, in the comments on the last post, "miker" reported some circulation numbers from a consortium. He blindly plugged in his numbers to my formulae, and got predicted numbers within a factor of two of the observed numbers, which seemed pretty miraculous to me. He was disappointed. 

Miker's data covers 4 years compared to Huddersfield's 13, and so a book that has circulated 100 times probably has spent little time on a library's shelves. A power law predicts significant numbers of books even at impossibly high usage. For example, the power-law fit to miker's data predicts that over a thousand books would be circulated more than once a day, which isn't possible given normal lending periods. See the notes if you're not scared of math and want to know how to adjust a fit.

The best way to advance the study of this phenomenon is to look at more data. If you have access to library circulation data, you can extract some numbers and publish them. A comment here would be appreciated. It's most helpful to report the number of items that have circulated f times as a function of f. Tab delineated text works great. In addition, analysts need to know the total number of items, total number of circulations, and the number of years covered by the data. An indication of the typical lending period would also be nice.

Along with a better understanding of how book collections get used, a better science of book use will help libraries and publishers formulate ebook circulation models that make sense for everybody who benefits from the reading of books. That's all of us.

Notes:
  1. If you want to fit a power law to circulation data truncated at some lending frequency fmax, you have to adjust the fitting parameters. We still have the same expression for the number of circulations for a given frequency, N(f).
    But the computation of the parameters from collection size and total circulations is more complicated:
    It's easiest to solve these equations numerically for N0 and A from the known C, N and fmax
  2. Please read the follow-up.
Enhanced by Zemanta

Article any source

Friday, March 11, 2011

The Pareto Principle and the True Cunning of HarperCollins

I take it back. I see now that HarperCollin's new strategy for ebooks in libraries is not nearly as senseless as it first seemed to me. In fact, it's a cunning plan worthy of Blackadder.
Black Adder IV - Black Adder Goes Forth
In case you're new to this library and publishing controversy, HarperCollins, one of the "Big 6" US publishers, has decided to require the expiration of the ebooks it offers to libraries after 26 checkouts. A library would have to relicense the ebook after the 26 checkouts if it want to keep the ebook in its circulating collection. Needless to say, librarians and many others were not happy about this.

HarperCollins' strategy puzzled me, because I couldn't figure out how it would make any money for them. I thought any extra sales caused by ebook expirations would likely be offset by poor sales of the limited-durability ebooks.

Libraries struggled to figure out how the new policy would affect them, and started looking at their circulation statistics. For example, Laura Crossett reported that at her library, 23,083 out of the 88,680 circulating books in her library's collection had been checked out more than 26 times over the course of 15 years. 220 books had been checked out more than 100 times. Matt Hamilton reported his numbers: 7566 books from a collection of 288,793 had circulated more than 26 times; 942 items had circulated more than 52 times. Most of the materials in his library are 3-4 years old. On Twitter, West Chester Public Library reported over 10,000 books from its collection of 58,000 had been borrowed more than 26 times over 17 years. Jason Griffey reported stats from his (academic) library: in 10 years, only 126 items from a collection of 409,213 had circulated more than 26 times.

These numbers are a bit all over the map, and I wanted to make some sense of them. According to IMLS data for 2007, US public libraries had collections totaling a bit more than 812 million print volumes. They circulated these items 2.17 billion times in 2007. That works out to an average of 2.6 circs/volume. Of course circulations will be unevenly distributed, but if HarperCollins terms were applied to print, the "average" volume would be expected to last 10 years.

A true understanding of these numbers would come from a better characterization of how circulation is distributed over the collection of a real library. You've probably heard of the "80/20 Rule" which in this case would say that 80% of the borrowing is concentrated on 20% of the collection. This is also known as the "Pareto Principle" which is a consequence of power-law distributions. I wondered if this was a good description of book circulation in libraries. I wanted to see some data.

OCLC's Lorcan Dempsey pointed me to the motherlode. The University of Huddersfield, in England, has released a huge file containing circulation and recommendation data extracted from almost 3 million transactions spanning over 13 years. I set to work analyzing the data.

The result is quite remarkable. The data shows a distribution of circulation frequency following a power law over 3 orders of magnitude, with a R2 of 0.9969! (update: see note 10 below.) Here's the plot of the number of books that have been circulated N times at Huddersfield:

The equation for the circulation is pretty simple:

Here, N(f) is the number of books that have been checked out f times. N0 and A are fitting parameters; I used A=9 in my plot of the Huddersfield data. If I use the total number of circulations and the total size of the collection to fix these two parameters, I get a zero-parameter fit of the data that's still amazingly good, R2 of 0.9760

Using this equation, I can calculate what a limited check-out ebook "should" be worth, but I'll leave that to another post, seeing as even one equation may be too much for this blog post.

What I'll focus on here is the what's been referred to in the library literature as the "vital few" principal that results from this distribution. A large majority of the circulations are taken up by a relatively small fraction of the collection. In the Huddersfield data, roughly 20% of the collection is in fact responsible for roughly 80% of the circulation.

If we think about this in the context of ebook lending models, we see that HarperCollins has played a neat trick. By focusing our attention on the books that are lent many times, supposedly shortchanging the publisher and the author, HarperCollins has gotten us to overlook the 80% of books that don't circulate much at all. Libraries pay full price for those, too, and it's pretty clear that publishers make infinitely more money on books that don't circulate in libraries than on books that don't sell in bookstores!

On balance, the economic effect of libraries, in addition to those I've discussed before, is to shift money from very popular books to those that are less popular. It can be argued that libraries support a breadth of culture that would go away without their support. Guess who publishes those very popular books? The Big 6 publishers, of course. They pay the big advances to authors, the big coop advertising fees to bookstores, they get their authors on talk shows and their books reviewed in the Times. That takes a lot of money, but the expenditure is richly rewarded by a "vital few" or "smash hit" economy.

So here's the cunning. By focusing on popularity-driven revenue mechanisms, HarperCollins is pushing money towards the smash hits and away from the long tail. Libraries may be adversely affected, but they're collateral damage. It's the long tail publishers that HarperCollins is trying to destroy.

All of HarperCollins' strategy is directed  at making hits bigger. The loss of big-box bookstores like Borders has disproportionately hurt  smash-hit publishing houses. They're poorly positioned to take advantage of the internet-induced fattening of the long tail that has been documented by Brynjolfsson, Hu and Smith in their paper on Amazon sales rankings. Rather, Big 6 profitability is improved by selling more copies of fewer books.

I didn't think so, but the HarperCollins strategy really does make sense. It's part of the big push.



Notes:
  1. For a review of what people have written about HarperCollins, Librarian by Day is all over it.
  2. Thanks to Dave Pattern at Huddersfield and the JISC TILE Project for making the release of the circulation data possible.
  3. The Huddersfield data starts at books with 5 circulations. For counts greater than 100, I binned the data in groups of 10 to reduce noise. The data falls off the power law at over 400 circulations/book. This must be close to limit of always being in circulation.
  4. Yes, all you need is the total circulation and the collection size to predict the distribution of the circulation. If you want to model your own circ stats, the formulae for A and N0 are as follows:
    • A = C/2N where C is the total circ and N is the number of items in the collection.
    • N0 = (3/4) (C3/2N)1/2
    Amazing, isn't it? Remember this is an idealized system, so your mileage may vary. Weeding will pull down the small N part of the curve; availability limits will truncate the large N part of the curve.
  5. The "vital few" principle was articulated by JM Juran in 1954. "Universals in management planning and controlling" Manage. Rev. 43(11), 748–61 (1954).
  6. JD Eldridge has a nice discussion of Juran, Pareto, and Trueswell (another scholar of book circulation) in "The vital few meet the trivial many: unexpected use patterns in a monographs collection", Bull. Med. Libr. Assoc. 86(4), 496–503 (1998). http://www.ncbi.nlm.nih.gov/pmc/articles/PMC226441/
  7. Brynjolfsson, Erik, Hu, Yu Jeffrey and Smith, Michael D., "The Longer Tail: The Changing Shape of Amazon’s Sales Distribution Curve" (September 20, 2010). Available at SSRN: http://ssrn.com/abstract=1679991. I plotted the Huddersfield data as done in this paper, and the library curve has the same slope they report for the 2008 Amazon data. Not very straight, though.
  8. Brynjolfsson, Erik, Hu, Yu Jeffrey and Simester, Duncan, Goodbye Pareto Principle, "Hello Long Tail: The Effect of Search Costs on the Concentration of Product Sales" (November 2007). Available at SSRN: http://ssrn.com/abstract=953587. This is a study very relevant to libraries. I wish these guys would show more data, though.
  9. There's a lot of old work (60s and 70s) on library circulation distributions with a whole bunch of theory. It's impressive, because they seem to have collected data by hand, but I fear the theory is too old to be useful. The 80s and 90s were marked by huge advances in the scientific study of self-organizing systems resulting in power laws.
  10. (added March 17) Cosma Shalizi (first commenter on this post) has done a fit of the Huddersfield data to a Log-Normal distribution; I'll try to explain what this means in a subsequent post.
Enhanced by Zemanta

Article any source