Showing posts with label ALA Midwinter. Show all posts
Showing posts with label ALA Midwinter. Show all posts

Wednesday, February 13, 2013

One eBook to Prove Them All

I've not written much about it here, but over the past year I've been participating in the American Library Association's "Digital Content Working Group". DCWG is broken up into smaller groups focusing on specific areas. I've been working on "Business Models". At ALA's Midwinter meeting, DCWG sponsored a jam-packed symposium.

The DCWG's meetings at ALA's mid-winter and annual conferences are open for anyone to attend, and they've been covered by the library press. Our recent meeting in Seattle was covered by Library Journal's Matt Enis, and he highlighted an idea that came out of our subgroup, the "One eBook" program:
The American Library Association’s Digital Content and Libraries Working Group (DCWG) has begun exploring an idea that could help publishers better understand the powerful impact that libraries can have for their authors and their bottom line.
I've finally had a chance to write this up for American Libraries' E-Content Blog:

There’s a lot of data suggesting that exposure to books in libraries increases sales for those books. There’s also a lot of data that suggests that many publishers believe the opposite—namely, that the availability of books in libraries depresses sales, and that if libraries improve the ebook lending process, making it easier for library users to substitute loans for sales, then ebook sales will be hurt even more. 
That word “suggests” is the problem. We don’t have controlled experiments that have really measured the broad effect of the library lending of ebooks on ebook sales. ALA’s Digital Content and Libraries Working Group has been examining the situation, and we had an idea. What if libraries all around the country promoted a single ebook for a month? What if that ebook’s publisher offered a special deal so that for that one month, libraries could lend that ebook to as many patrons in their communities as possible without decimating their acquisition budgets? Once the month was over, that specially promoted library ebook deal would end. What do you think would happen?
There are a lot of details to work out of course, but we've had a lot of positive reactions. It's the practical and technical details I'm thinking about right now. For example, how can we make such a program available to as many libraries as possible, regardless of whether they are currently offering ebooks? How can we make the ebooks work on all sorts of platforms? How do we make the one-ebook ebooks expire after a month?

As if I didn't have enough to do...

Enhanced by Zemanta

Article any source

Monday, February 6, 2012

Libraries Happen

Michael Scotto's Just Flash is a children's book about a strange animal trying to figure out its identity.
Flash is the only animal of his kind, but he wishes he could just belong to a pack like everyone else at the zoo. He first tries to become another animal, in order to fit in. He tries to be a giraffe; he tries to be a gazelle; he even tries to be a zebra. Along the way, though, Flash realizes that it's more rewarding to be himself – no more, no less.
Libraries and librarians crossing over from print to digital must feel a lot like Flash. Even newly born entities like the Digital Public Library of America (DPLA) are finding themselves in an awkward period of not knowing what they should be.

I was a guest at a recent DPLA "Audience and Participation Work Stream" meeting that brought together participants with a variety of perspectives and experience, but joined by a deep concern for the future of libraries. We met at the Dallas Public Library, just after the American Library Association's Midwinter Meeting at the nearby Convention Center. Our charge was to help articulate "why we need a public library when it's all on the web" and suggest how to attract involvement from all types of libraries.

The DPLA has the incredibly difficult task ahead. It must weave together many distinctive strands of activity emerging from a strong but inchoate desire for the missions of libraries to continue in new more powerful ways.  It's encouraging to me that the DPLA leadership is spending a lot of time learning about the needs and desires of diverse communities. But I think that sometimes our understanding of what libraries are today holds us back from seeing what the library movement could be tomorrow. Separating the essential from the institutional manifestation is not always easy to do.

During my time in Dallas, I had the fortune to witness a "library" happening in its purest, most human form.

YALSA is one of the American Library Association's many incomprehensible abbreviations. YALSA stands for the Young Adult Library Services Association. I don't really know much about what YALSA does, but on the Sunday exhibit floor of the ALA meeting, I saw more than one group of teenagers roaming the exhibits wearing YALSA T-shirts. On the way back to my Dallas hotel, I shared a subway train car with one of these groups. I started chatting with two of the teens. It seems that the librarian at their suburban Dallas high school had organized their expedition. The teens had made off with huge bags full of books courtesy of the exhibit vendors.

Sitting on the other side of the train car was a Hispanic family- a young mother, her daughter of perhaps 6 years, and an older man who was probably the girl's grandfather. The little girl was reading what appeared to be a fast food restaurant placemat, but it didn't seem to interest her much. The YALSA teen next to me noticed. She asked the mother, "Are you going a long way on the train? I have some books here she might like." And out of the bag appeared a copy of Just Flash. The girl's eyes lit up. For the next 20 minutes I watched the girl enchanted by an illustrated story that had appeared as if by miracle, chosen specifically for her. As I got off the train, a second book was emerging from the teen's bag.

I don't know what sort of animal libraries will evolve into. Maybe librarians will ride trains with mega-book digital libraries on memory sticks for kids who need them. Maybe restaurant placemats be ultra-cheap reading devices. Whatever stripes it wears or what name it answers to, the simple act of letting a book bring joy and wonderment to a little girl will define what a library must be, no more, no less.
Enhanced by Zemanta

Article any source

Thursday, January 26, 2012

Unglue.it Preview All Systems Go

A preview version of "unglue.it", the crowd-funding site for creative commons ebooks that I've been working on for more than a year, opened last week. Some key features are missing (pledging, campaigns) but the site lets you make a list of books you would support for "ungluing".

You can't really plan for a launch. Things always happen that you don't expect. It helps to have had a good night's rest, but other than that...

Our first unexpected event was that Library Journal ran a piece about our "soft launch" on their Digital Shift website, while we were in the process of deploying the website to production. They didn't link to us, but a few impatient readers typed in the website name and started exercising the site before we were finished testing the deployment. Nothing awful happened. Thanks, dave and gsf! Then Google spidered the site, exposing one or two errors. Thanks, googlebot!

We wanted our mailing list subscribers to be the first to see our work, and we finally sent out the email on Thursday. List readers discovered that our "popular" and "unglued" views were running very very slowly, loading down the site. Raymond studied the problem, and, as seems to happen so often with Django, found the answer hidden in plain sight (the documentation). After moving some nested queries, the pages returned 100x faster. The miracles of EC2 allowed us to spin up a bigger server to help with load. And the high load from the glacial queries helped expose some concurrency problems that we never would have found in a million years of normal operation. Or so says the errant coder- me.

We wanted to open the website when we did so that we could show our work to our many friends at the American Library Association Midwinter meeting in Dallas. So on Friday, I got up at 5AM (after bugfixing till 1AM) to catch a flight. I decided not to have any coffee so I could sleep on the plane. When I arrived at my Dallas hotel, I discovered another unexpected occurrence: I had left my laptop on the plane.

Before reading the next paragraph, check your laptop, your iPad, your kindle, your nook, or whatever. It probably looks plain, like mine (left). If there is no identification on it, go get one of those free address labels you got from the Awful Disease Foundation, and stick it on. Also some stickers from your favorite organizations. When you are done, it should look like @vmbrasseur's (right). Are you done?


Here's what I learned about lost MacBook Pros and airlines. Once the battery runs out, you can't even find a serial number. The typical baggage claim operation does not have geek squad backup. They don't have spare power cords or batteries to help them ID lost laptops. What they DO have is a safe, and that's where errant laptops go to die. If you ever find yourself in my position, go to the airport and ask the friendly lost-luggage attendant to go look in the safe. Otherwise, you will never see your laptop again, even if you have entered its serial number into the web form that has replaced the lost and found phone number that no one helpful ever answers.

In contrast, Jeanette, the DFW Continental Airlines baggage claim professional that I talked to in person, was very helpful. She called back to the guy with the safe, and we hardly had time to joke about the huge bag of dried fish from Africa that was smelling up the lost baggage area before safe-guy came out with MY LAPTOP. Yay!

Raymond had emailed me with a status report from the virtual home office, reproduced here in its entirety:
http://idioms.thefreedictionary.com/All+systems+go
So I was feeling pretty good. I got to the Convention Center and found Andromeda doing an in-person demo of unglue.it. The in-person demos are an invaluable complement to submitted feedback reports because they let you see expectation mismatch as well as outright failures. Andromeda seemed to have the demo drill down to a science. I am thankful for the generosity of our in-person testers, whose insights will soon be incorporated into the site.

Since this post has been accepting digressions, I must note here that Andromeda gives new meanings to the adjective "awesome". At some point over the past few months, most likely due to lack of proper supervision coupled with web development despair, Andromeda has learned to code javascript and CSS. In a subsequent period of inadequate supervision, Andromeda seems to have recruited a squadron of librarians learning to code, which is somehow becoming an official ALA "codeyear" Interest Group. I doubt that we have heard the last of this.

And then we had dinner at Wild Salsa. (I'm skipping a few things here and there.)

On Sunday, Beth Kephart's article on Unglue.it went live at Publishing Perspectives. Beth writes so beautifully that it hurts. Her first book, A Slant of Sun is on my Unglue.it wish list. Her article introduced the Unglue.it concept to hundreds of new book lovers, more than a few rights holders, and generated a bunch of traffic. We've always expected that we'd need some publicity to find significant numbers of rights holders willing to take the plunge for a completely new business model, and the Publishing Perspectives article was a great start. Ed Nowotka's more cautious commentary is spot on, as well.

In the first week, the preview site has signed up 133 users (a conversion rate of about 10%) and we've received numerous suggestions for improvement. Our intrepid ungluing pioneers have added over 7500 works to our database. The most frequent comment is that we need better ways to indicate works that are already "unglued", either by virtue of being in the public domain, or by being already available under creative commons licenses. Raymond is currently working on loading Project Gutenberg titles; there will be more "unglued" books added as we go on, as well as ways of adding them directly. Coming in second were requests to have more selective imports from GoodReads and LibraryThing.

I should mention that we've had some great talent helping the core unglue.it team. Most prominent is the design work of Stefan from Design Anthem. We've had part-time help on systems and software from Ed Summers and Jason Kace. And it's hard to overlook the contribution of the countless developers who contributed to the open source Python and Django projects.

If you haven't tried the site yet, please give it a spin and tell us what you like (or dislike). The more people that sign up, the less skeptical rights holders with interesting books will be about the concept. If you're a rights holder or a rights manager of any kind, please contact Amanda at rights@gluejar.com with your ideas and questions. Follow @unglueit on Twitter, like unglueit on Facebook.

Enhanced by Zemanta

Article any source

Friday, January 21, 2011

Doing Good Things Together

I've spent a lot of time in the past few weeks explaining to everyone I meet why I think ordinary people might be willing to help acquire ebook rights for the public commons. Meanwhile, I kept noticing how people are getting together to do other good things.

At ALA Midwinter, there was tweeting going around about an effort by four twittering librarians, Andromeda Yelton (@ThatAndromeda), Ned Potter (@theREALwikiman), Jan Holmquist (@janholmquist), and Justin Hoenke (@JustinLibrarian) to "buy India a Library". As of Wednesday, they had raised £1384; the fund raising ends today, so hurry on over if you want to participate.

On the mailing list for organizers of the Code4Lib conference, Dan Chudnov was agitating for a way for anybody to become a conference sponsor. There were a number of minor issues to overcome, but Kevin Clarke took up the challenge and created a ChipIn page to collect money from those who wanted to contribute to a sponsorship. This page raised $1,240 from 28 contributors. (Sorry, too late for that!)

I also found out about an ambitious effort by Michael Porter and friends called Library Renewal. They've created a non-profit organization to explore "new content solutions for libraries, while staying true to their larger mission." This is an effort that's still in its formative stages; it's an effort you can join and help shape.

These three projects are in all the library world, but please don't think that good people doing good things aren't everywhere around you. I've been inspired by my college friend Noel Valero. After graduation, he worked as an aerospace engineer and then as an IT consultant, until he began to have trouble with spasms in his arm. He spent a lot of time seeing doctors who were unable to help him until finally, with the help of another classmate, he was diagnosed with dystonia, a little-known but not-so-rare disease that causes progressive loss of motor control.
Dystonia is the 3rd most common movement disorder, with an estimated 500,000 patients diagnosed with primary and secondary forms of the disease and possibly at least another 500,000 others that are undiagnosed or misdiagnosed. Yet Dystonia lags significantly behind in research funding when compared to other neurological disorders.
Many sufferers of dystonia lose hope amid the progression of the disease, partly because of the isolation it forces on people. Simple everyday tasks become huge barriers. Even holding a book to read it can be difficult. One dystonia sufferer that Noel introduced me to reports that she can only manage her graduate school textbooks by chopping off their spines and dividing them into easy-to-hold segments. Driving a car or typing on a computer can become exhausting activities.

With loving support from his family and friends, Noel has climbed out of his initial despair. He started reaching out to other dystonia sufferers on Facebook (his daily joke posting is a resource for non-dystonia-sufferers as well!) and was surprised to find how much it helped for people with dystonia to be able to support each other. In 2009, he took these efforts to the next level by forming the American Dystonia Society.

On February 2, Noel will be on an episode of "Mystery Diagnosis". If you have access to the new "Oprah Winfrey Network", please join me in watching the show (or record it for later viewing). And if you enjoy reading this blog (or if you don't), I would be honored if you made a donation of any size to the American Dystonia Society in appreciation.

In possibly related news, the blog's Amazon Associate revenue statement for 2010 just came in: $10.04.
Enhanced by Zemanta

Article any source

Saturday, January 15, 2011

Why ProQuest Bought ebrary

The New York Times
Take a look at the New York Times homepage. Then take a look at CNN.com or MSNBC. How do you tell which website belongs to a newspaper and which ones belong to a television network? All of them have video. All of them have text. All of them have blogs and forums. As media moves onto the internet, the boundaries between old media genres begin to blur, and new forms emerge, optimized for the purposes they're being used for.
CNN.com

Just as delivery of news is being transformed by the Internet, the needs of students, researchers, and scholars are driving a similar boundary-blurring transformation in libraries. It's also driving a transformation in the companies that serve the library industry.

Marty Kahn, President of ProQuest, used the Times-CNN analogy to explain to me why his company had acquired ebrary, a leader in providing ebooks to academic, corporate, and other libraries. It no longer makes sense for a company to specialize in only journal articles, databases, or eBooks if it wants to be able to provide coherent and evolving solutions.

A look at ProQuest's existing product suite bears that out. With full-text journal databases, newspapers, dissertations, historical archives and government documents (including the CIS division recently acquired from LexisNexis) ProQuest was already able to integrate an impressive array of content. The Summon service from ProQuest's SerialsSolutions unit, which centrally indexes a library's content, has experienced rapid growth, with sales at 200 institutions already. Still, the most common questions that Summon staff were fielding at ALA Midwinter surrounded the integration of ebooks into Summon. With the acquisition of ebrary, ProQuest can now answer that question authoritatively for at least one ebook vendor. (See my previous article focusing on Overdrive.)

Somehow, the topic of EBSCO and their recent acquisition of NetLibrary hardly came up in my talk with Kahn.  We spent a lot more time discussing Google. Between Google Search, Google Scholar and Google Books, Google also has the potential to present a comprehensive information solution for libraries. I often hear librarians expressing the sentiment that they need help from companies like ProQuest to present credible alternatives to Google and free sources available on the internet.

One thing Summon and other library search solutions have lacked is the ability to search the full text of the books in a library's collection. Put next to Google Books' full text plus metadata search, the metadata based search offered by a traditional library catalog can seem rather limited to most users. ebrary will bring with it a huge library of full-text book content for search within Summon.

ebrary was founded by high school friends Christopher Warnock and Kevin Sayar. Libraries were the focus from the very start. Warnock had left a job at Adobe Systems and was working on a project for Stanford University when Stanford University Librarian Mike Keller told him that in order to get paid, he had to incorporate. Warnock called up his friend Sayar, then an attorney at the legendary Silicon Valley law firm of Wilson Sonsini Goodrich & Rosati, and asked if he wanted to act on their high school dreams of starting a company together. The project at Stanford led to the conception of ebrary's initial service for libraries. (I've often heard the misconception that ebrary is somehow an Adobe funded spin-off, because of Warnock's father's role as a Founder of Adobe. In fact, Adobe and the elder Warnock had no role in starting ebrary.)

Warnock has always been passionate about ebrary's mission. "If every library acquired information digitally, all the worlds information would be free to everybody", he told me. He is genuinely excited about what ebrary will be able to do as part of ProQuest. "Being part of ProQuest will allow us to realize our dreams".

Those dreams include the creation of a vast digital library with all kinds of content. ProQuest has "billions" of PDF documents, according to Warnock; ebrary's PDF indexing and search technologies are considered to be unsurpassed anywhere. Although ProQuest is not known for ebook distribution, there's not much difference between a book and a dissertation, if you think about it. ProQuest distributes 70,000 of those every year.

ebrary has also been an innovator in business models as well as in technology. ebrary's initial model was to make ebooks available for free viewing; rights-holders were compensated using a micro-transaction model where subscribers were charged every time they did things such as print pages. Based on customer feedback, they shifted to a model where most content is available for use with on flat subscription. Fee. This year, they've begun to implement a patron-driven acquisition model.

Looking forward, Sayar will be running the ebrary business unit; Warnock will move to ProQuest to work on strategy. Given the ambitious vision outlined by Kahn, he has his work cut out for him.

The ebrary content platform has definitely gained some ardent advocates in libraries. I heard one librarian say "not only do we love ebrary, but our students love ebrary. They really do." At the end of the day, when we ask ourselves how libraries will respond to the dizzying changes in both information and economic landscapes and worry about what will happen, isn't love all we really need?


Article any source

Sunday, January 9, 2011

Bridging the eBook-Library System Divide

Despite what you might have read on the blogs, libraries show no signs of imminent ebook-induced death. The latest data from Overdrive, the dominant provider of eBooks to public libraries, shows staggering growth. Digital checkouts doubled in 2010 to 15 million, looking at Overdrive alone. Based on the buzz at this weekend's American Library Association Midwinter Meeting, Overdrive should blow those numbers away in 2011- It seems that almost every librarian I've talked to here has decide to "take the plunge" into eBooks in a big way in 2011.

The ebook companies focused on academic libraries are experiencing the same growth- Ebook Library told me that for the prior year their monthly sales have been double the prior year. The biggest plunge was taken by Proquest, which announced their acquisition of ebook provider ebrary. (I’ll have a separate story on that later.)

To some extent, most libraries have been only sampling the ebook water, and despite noted usability issues and e-reader device fragmentation, patrons seem to want more and more and librararies are responding to patron demand. But not everyone is happy. One librarian told me, after a few beers, that “Overdrive sucks!” and then went on to use language unsuitable for a family-oriented blog.

As far as I can tell, there are two issues around Overdrive that are troubling libraries. One derives from the DRM system from Adobe that Overdrive uses. Adobe’s system is pretty much the only option for libraries and booksellers other than Amazon and Apple; Overdrive has no choice but to use this system in order to work with reader devices and software from Barnes&Noble, Sony and Kobo. The Internet Archive’s Brewster Kahle, in a panel on Saturday morning, slammed the Adobe system, even though it’s used by the Archives OpenLibrary. In OpenLibrary's experience, users were able to complete a lending transaction in only 43% of their attempts. Overdrive is working to improve the smoothness of these transactions, and is introducing new support methods to make the processs easier.

The second issue was discussed by library system vendor executives at Friday’s RMG President’s Panel. According the Polaris Library Systems President Bill Schickling, many of his customers are worried that their libraries will be marginalized by ebook providers like Overdrive.  Although Overdrive offers extensive customization options for their ebook lending interface, libraries are still upset that patrons have to use separate interfaces for books and ebooks, one provided by Overdrive and the other provided by their ILS vendor. Libraries often think of the library system as their primary "brand extension" on the internet.

It seems a bit odd that this should be an issue. For years, libraries have lived with databases and electronic journals delivered from separate systems. But books are different. Libraries want ebooks and books to live side by side. It makes little sense to force a user who wants to read a Steig Larsson novel  have to check in two places to see print and digital availability.

Overdrive is working overtime to address this second issue, it seems. Overdrive's CEO, Steve Potash, told me that his company is working on opening a set of APIs (application programming interfaces) that will allow system vendors, libraries and other developers to more deeply integrate Overdrive's ebook lending systems into other interfaces. Overdrive has needed these interfaces internally to build reading apps for Android, iPod and iPhone. Overdrive hopes to have an iPad-optimized reading app in Apple's iTunes stare by the end of first quarter 2011, and will be working with selected development partners to work out many of the details. Potash hopes Overdrive will be able to unveil the APIs this summer at the ALA meeting in New Orleans.

The Overdrive APIs and the usability improvement they lead to should come as welcome news to libraries and library patrons everywhere. Library system vendors and developers in libraries will have a lot of work to do over the coming year.

And library patrons will be reading a lot of ebooks.
Article any source

Saturday, January 8, 2011

Inside the Dataculture Industry

wild blueberries
I don't really know how all the food gets to my table. Sure, I've gathered berries, baled hay, picked peas, baked bread and smoked fish, but I've never slaughtered a pig, (successfully) milked a cow or roasted coffee beans. In my grandparents generation, I would have seemed rather ignorant and useless. Agriculture has become an industry as specialized as any other modern industry; increasingly inaccessible to the layperson or small business.

I do know a bit about how data gets to my browser. It gets harvested by data farmers and data miners, it gets spun into databases, and then gets woven into your everyday information diet. Although you've probably heard of the "web of data", you're probably not even aware of being surrounded by data cloth.

The dataculture industry is very diverse, reflecting the diversity of human curiosity and knowledge. Common to all corners of the industry is the structural alchemy that transmutes formless bits into precious nuggets of information.

In many cases, this structuring of information is layered on top of conventional publishing. My favorite example of this is that the publishers of "Entertainment Week" extract facts out of their stories and structure them with an extensive ontology. Their ontologists (yes, EW has ontologists!) have defined an attribute "wasInRehabWith" so that they can generate a starlet's biography and report to you that she attended a drug rehabilitation clinic at the same time as the co-star of her current movie. Inquiring minds want to know!

If you look at location based services such as Facebook's "places", Foursquare, Yelp, Google Maps, etc, they will often present you with information pulled from other services. Often, a description comes from Wikipedia and reviews come from Yelp or Tripadvisor and photos come from Panoramio or Flickr. These services connect users to data using a common metadata backbone of Geotags. Data sets are pulled from source sites in various ways.

Some datasets are produced in data factories. I had a chance to see one of these "factories" on my trip to India last month. Rooms full of data technicians (women do the morning shift, men the evening) sit at internet connected computers and supervise the structuring of data from the internet. Most of the work is semi-automated, software does most of the data extraction. The technicians act as supervisors who step in when the software is too stupid to know when it's mangling things and when human input is really needed.

There's been a lot of discussion lately about how spammers are using data scraped from other websites and ruining the usefulness of Google's search results. There are plenty of companies that offer data scraping services to fuel this trend. Data scraping is the use of software that mimics human web browsing to visit thousands of web pages and capture the data that's on them. This works because large websites are generated dynamically out of databases; when machines assemble web pages, machines can disassemble them.

A look at the variety of data scraping companies reveals a broad spectrum. Scraping is an essential technology for dataculture; as with any technology, it can be used to many ends. One company boasts of their "massive network of stealth scrapers capable of downloading massive amounts of data without ever getting blocked. Some companies, such as Mozenda, offer software to license. Others, such as Xtractly and Addtoit are strictly service offerings.

I spoke to Addtoit's President, Bill Brown, about his industry. Addtoit got its start doing projects for Reuters and other firms in the financial industry; their client base has since become more "balanced". Companies such as Bloomberg, Reuters and D&B get paid premiums to provide environments rich in structured data by customers wanting a leg up on competitors. Brown's view is that the industry will move away from labor intensive operations to being completely automated, and Addtoit has developed accordingly.

A small number of companies, notably Best Buy, have realized that making their data easily available can benefit them by promoting commerce and competition. They have begun to use technologies such as RDFa to make it easy for machines to read data on their web sites; scraping becomes superfluous. RDFa is a method of embedding RDF metadata in HTML web pages; RDF is the general data model standardized by the W3C for use on the semantic web, which has been discussed much on this blog.

This doesn't work for many types of data. Brown sees very slow adoption of RDFa and similar technologies but thinks website data will gradually become easier to get at. Most websites are very simple, and their owners see little need or benefit in investing in newer website technologies. If people who really want the data can hire firms like Addtoit to obtain the data, most of the potential benefits to website owners of making their data available accrue without needing technology shifts.

The library industry is slowly freeing itself from the strictures of "library data" and is broadening its data horizons. For example, many libraries have found that genealogical databases are very popular with patrons. But there is a huge world of data out there waiting to be structured and made useful. One of the most interesting dataculture companies to emerge over the last year is ShipIndex. As you'd expect from the name, ShipIndex is a vast directory of information relating to ships. Just as place information is tied together with geoposition data, ShipIndex ties together the world of information by identifying ships and their occurrence in the world's literature. The URIs in ShipIndex are very suitable for linking from other resources.

The Götheborg
ShipIndex is proof that a "family farm" can still deliver value in the dataculture industry. The process used to build ShipIndex. Nonetheless, in coming years you should expect that technologies developed for the financial industry will see broader application and will lead to the creation of data products that you can scarcely imagine.

The business model for ShipIndex includes free access plus a fee-for-premium-access model. One question I have is how effectively libraries will be able leverage the premium data provided with this model. Imagine for example the value you might get from a connection between ShipIndex and a geneological database bound by passenger manifests. I would be able to discover the famous people who rode the same ship that my parents took to and from the US and Sweden (my mom rode the Stockholm on the crossing before it collided with the Andrea Doria). For now though, libraries struggle to leverage the data they have; better data licensing models are way down on the list of priorities for most libraries.

Peter McCracken
ShipIndex was started by Peter and Mike McCracken, who I've known since 2000. Their previous company (SerialsSolutions) and my previous company (Openly Informatics) both had exhibit tables in the "Small Press" section of the American Library Association exhibit hall, where you'll often find the next generation of innovative companies serving the library industry. They'll be back in the Small Press Section at this weekend's ALA Midwinter meeting. Peter has promised to sing a "shanty" (or was that a scupper?) for anyone who signs up for a free trial. You could probably get Mike to do a break dance if you prefer.

I'll be floating around the meeting too. If you find me and say hello, I promise not to sing anything.
Enhanced by Zemanta

Article any source

Wednesday, February 10, 2010

Branches of Koha

An arborist recently came to look at the white oak tree in my back yard. The tree is about 80 years old and is the biggest in the neighborhood. According to the arborist, our tree was in excellent health because of its large number of leaders, or main branches. Even in the strongest wind, these leaders will bend and a few might even break, but the tree itself is very unlikely to topple. Some neighbors have a very tall, scary tulip tree with only one main branch. I fear it will come down one day very suddenly.

In my discussion of Koha and LibLime (part 1, part 2), I promised to write more about the so-called "forking" of the Koha development process. What happened was that about 10 months ago LibLime stopped participating in the the Koha open development community. According to Josh Ferraro, LibLime's CEO, this happened because LibLime developers were having trouble completing development projects that LibLime had committed to doing for customers. Together with his development partners, Ferraro judged that LibLime's developers were spending too much time providing support to non-customers and that the overhead of the community development process was slowing down development more than it was contributing to LibLime's development objectives.

This judgement is hotly contested by developers advocating a open community development process. (See, for example, this post by Chris Cormack, or Owen Leonard's post on the Koha List.) It's not hard to imagine that an open community might pose difficulties during focused development. Two coders may disagree about how a task is to be done, and depending on personality and skills involved,  such disputes might easily become major time-sinks. Implementing a feature important to US libraries might break a feature important to European libraries, for example, and making both features work at the same time might be a lot of work. On the other hand, a small community such as the one working on Koha can ill afford to fragment into factions and start working at cross purposes.

When one branch of code diverges too far from another, the branches can become incompatible, or forked. When this happens, effort applied to one branch may have to be duplicated for the other branch. Forking is a common occurrence in open source projects, and can be evidence of a project's health. Such a fork in the Linux kernel came to light just last week, as some drivers added to support Google's Android system were deleted from the project's main tree. The downside is that forks increase the maintenance burden. It's often worthwhile for developers to work hard to join their code to a main branch so that others can maintain the contributed code and keep it from breaking.

Open Source projects can be motivated in many ways. Some projects have their origin in proprietary software, when the developers decide their businesses would benefit from wider adoption or support. Or perhaps the emphasis of the developer's business has changed. Etherpad is a recent example- the company was acquired by Google, whose main interest was to improve Google Wave rather than to continue the Etherpad service.

Other Open Source projects arise as "calling cards". That's how IndexData started doing Open Source. Sebastian Hammer, IndexData's Founder and President, told me that when he started, he just wanted his software to be widely used. His business was primarily custom development, and companies who were using his software because it was free began using his company for development because the free software worked well.

Only a small percentage of open source software projects are supported by more then just a few developers, and even fewer survive without an acknowledged lead developer. Koha has been blessed with significant contributions from a number of developers (and it uses free open source components such as Apache, MySQL, Perl, PHP and IndexData's Zebra).

You can imagine that Koha contributors outside LibLime would be very upset at LibLime's withdrawal from community development. Their contributions to the project were made with an understanding of Koha as an inherently community-driven effort for the benefit of all Koha libraries, and LibLime's withdrawal from the community process implicitly minimizes the value of their ongoing contributions. In fact, several contributors within LibLime were upset at the changes, and are now working for competing companies.

The fact of competition among Koha project participants inevitably leads to conflicting incentives. While proprietary software creates incentives for vendors to compete for initial sales with a robust platform and advanced features at the expense of ongoing service, open source software creates incentives for companies to focus on service and custom development. A company that puts a lot of effort into the core software may gain no advantage from that work if it has competitors which instead focus on services. Competitive considerations may certainly have been an important factor in the manner of LibLime's withdrawal from the community process.

It's interesting to see the messaging that LibLime's competitors used to respond to LibLime's withdrawal from the community development process. These ranged from the sunny "Equinox Promise" to a pointed post from BibLibre and a worried post from ByWater. There is clearly a struggle among all these companies to resolve the tension between competition and the need for cooperation that underlies Open Source support businesses. (Note: Equinox supports a different library system, Evergreen, so it's not so directly competitive with LibLime. Update Feb. 11- Equinox announced its entry into the Koha support market.)

Another reason given by LibLime for the change in its development process was that they wanted development customers to be able to test and approve new functionality before it would be released to the world. In the words of a LibLime press release:
"A public software release of each version of LibLime Enterprise Koha will occur periodically, after the sponsoring library and LibLime's customers have had adequate time to ensure that the codebase is of sufficient quality and stability to be contributed back to the Koha Community."
One Koha developer described this rationale to me as "nonsensical" and pointed out that the code quality seemed to be good enough for LibLime production customers. A look through the Koha developer wiki gives the impression that an elaborate QA process has been built by the community; I don't know how well it is followed.

My perspective, as someone who has managed a development project of similar scope, is that testing and quality assurance require a fair amount of disciple and attention to process. Getting developers to comment their code, do proper testing, and keep documentation up to date (i.e. adhere to a documented QA process) is not always easy, even if you're signing their paychecks. So while I have little insight into whether LibLime's new internal development processes are in fact resulting in better or more timely code, I think that the explanation given is at least plausible.

Managing a software development project is really, really hard. A lot of people imagine that their success in managing one project is evidence of superior process or ability, when really they were just lucky to have the right people at the right time. So I'm really skeptical when someone says that "community development" is the best way to build software, or that "agile methodology" is the one true way. In the real world, development managers may have the skills to succeed in one style of development (or group of developers) and be lacking in the skills needed to succeed in another style.  Software development projects only work if they work. In the case of the two branches of Koha, only time will tell whether one branch will wither and die, or whether two branches will end up diverging, both healthy.

While the people in charge at LibLime and PTFS have been in no position to comment on what they will do before their transaction is complete, other Koha stakeholders that I talked to were "hopefully optimistic" that PTFS would ultimately decide to rejoin the community development process and help reunify the Koha code base. PTFS developers have been active with contributions during the period that LibLime has pursued separate development. At ALA Midwinter, PTFS' John Yokley emphasized that a decision as to the extent of PTFS participation in Koha community development had not yet been made. In the meantime, Koha stakeholders other than LibLime have launched a new website to be the "Temporary home of the Koha Community."

 (Update Feb.12 - the acquisition is not happening.) (Update Mar. 16 - the acquisition closed after all.)

You can look at my tree analogy in two ways. You could say that having multiple branches of the Koha code is good for the project, as it is for my oak tree. You could also say that concentrating development of Koha in one company is dangerous, and worry about it as I do about the tulip tree.

Or you could just be happy that spring is coming and buds are already appearing on the trees.

This is the third part of a series. Also see Part 1 and Part 2
Reblog this post [with Zemanta]

Article any source

Tuesday, February 2, 2010

Back to the Future at the Storefront Library

I wish library I can buy book.

I wish we had a permanent library.

I wish to be happy and proud of my accomplishments.

In the window of the Chinatown Storefront Library in Boston stood a Wish Tree. Modeled after Yoko Ono's Wish Tree Project, the tree was meant to allow patrons to pass on a spirit of energy and hope. The instructions were:
Make a wish. Write it down on a piece of paper. Fold it and tie it around a branch of a wish tree. Ask your friend to do the same. Keep wishing until the branches are covered.
The Chinatown Storefront Library closed its doors on January 17, 2010, the Sunday that ALA Midwinter was in town. Always meant to be a temporary library, the Storefront Library was an expression by Boston's Chinatown community of its need and support for a library of its own. The Chinatown neighborhood of Boston has been without a branch of the Boston Public Library since 1956, when the branch was closed and demolished to make way for a highway.

Without a local branch, Chinatown residents needing library services have to go to the main library in Copley Square, which, though a beautiful building, may seem rather imposing and hard to navigate for someone looking for Chinese language materials.

The founders of Chinatown Storefront Library, Sam and Leslie Davol, had been involved in community meetings surrounding the proposed design and construction of a new branch of Boston Public Library, and in that process had gotten to know faculty at Harvard's Graduate School of Design. With a new branch on hold due to budgetary reasons, the Davols decided to take action. A local developer offered to let them use a vacant storefront for free. Design students made some gorgeous, modernistic shelving pieces for the library, enabling it to create an inviting environment in an bare commercial space. Library students from Simmons paired with Cantonese- and Mandarin-speaking community volunteers to staff the facility. Donations of over 5,000 books were solicited, and for twelve weeks, a community library came into existence. The operating budget for the entire project was about $10,000.

The day before the closing, I had a chance to tour the Storefront Library and sit down with Sam Davol. Formerly a legal-aid lawyer in New York, he and his wife moved back to Boston with their two children, partly so that Sam could devote more time to music. The Library project was an outgrowth of their involvement in the community and other cultural programming they've produced.

In just a few short months, the Storefront Library has had a clear impact on its neighborhood. People who used to avoid the block because of its vacant, spooky feel began to feel welcomed by the activity surrounding the library. Cultural activities, language classes and storytimes attracted people from the community and passersby.

Initially, the Storefront library did not plan to circulate books, but in the first week of operation patrons told them that they really wanted to take books home with them. A makeshift paper-based circulation system was implemented, and over 1,374 books were circulated in 11 weeks of operation, over half of them in Chinese. Over 4,000 books were catalogued using LibraryThing.

In talking to librarians in general about the storefront library concept, I've gotten a consistent reaction that small storefront spaces could not provide sufficient room to provide internet access; terminals take up more room than books. At the Storefront Library, the computers tended to be lightly used. When I was there, some older gentemen were reading newspapers, some children were reading books, but no one was using the computers or internet access. This could be because the Library did not subscribe to electronic resources.

I think the most important lesson that can be learned from the Storefront Library experiment is that even small temporary libraries can be powerful agents of community development. In Boston, this role was accentuated by a location in close proximity to people's everyday lives. While I've written that the future of public libraries may be in smaller locations, the Chinatown Storefront Library reminded me that many public libraries began as grassroots efforts to promote knowledge and culture.

Now that the Storefront Library has closed, its books will be going to a new reading room, to local schools, and a few to the Chinese Historical Society of New England. The furniture will be going to local schools and daycare facilities. Information about the project will be published on the storefrontlibrary.org website so that similar projects in other communities can learn from their experiences.

As for Sam Davol, he goes on tour. He plays cello with the indie-pop band "The Magnetic Fields", which has a new CD out, Realism. I just got my tickets for one of the shows at New York's Town Hall in March.
I wish there were more people experimenting with libraries.


Reblog this post [with Zemanta]

Article any source

Friday, January 29, 2010

Who Owns Koha?


In New Zealand, Maori customs are taken seriously. For example, in 2002, the route of a new highway through a swamp had to be altered because it was believed that three taniwha - Karutahi, Waiwai, and Te Iaroa - lived there, and were being disturbed by the road, causing an unusual number of accidents. Taniwha are mythical beings that act as guardian spirits. Many taniwha arrived in New Zealand as guardians of specific ancestral canoes and then took on a protective role over the descendants of the canoe's crew.

Another Maori custom, one that has crossed over into general New Zealand culture, is that of "koha".  "Koha" is often translated as "gift", but according to Chris Cormack, one of the original developers of the Free Open-Source Software (FOSS) Library System with the same name, a more accurate translation would be a "gift with expectations".   Cormack got his B.A. degree in both mathematics and Maori Studies, so he should know. A koha is gift that is offered with an expectation that it will be reciprocated.

In the U.S., (and to a lesser extent, in Europe) it's our lawyers that we seem to take seriously. And so when we let people use software that we've developed, its not enough to offer it as a koha, we have to use a legal license to spell out the terms of the release. The license that has become popular because of the expectations of reciprocity built in to it is the Gnu Public License (GPL).

Software released under the GPL cannot be thought of as an unconditional gift by its developer. GPL software is not in the public domain, it is copyrighted. A copyright owner can exert control over the use of the copyrighted material; the GPL uses that power to require licensees to publish any modifications they make if they want to redistribute the modified work.  When the Horowhenua Library Trust was choosing a license to use for Koha (the library system software), they chose GPL (version 2) because they thought it would prevent Koha from being further developed as non-open software.

Included in the recently announced (but not yet completed) acquisition of LibLime by PTFS were Koha-related assets, including source code copyrights, trademarks and the koha.org website. It may be difficult for the casual observer to understand what value these assets have, especially in light of the GPL license attached to Koha. Could these assets be used to privatize Koha in some way? The short answer is "No", but it gets complicated.

Some FOSS companies use "dual licensing" as a business model. They release their code under GPL, but if another company wants to use the software in a way that would not be allowed under GPL, it has the option to pay for a commercial license. Dual-licensing can only be done by the original copyright owner. Index Data has used this model for years to the great benefit of libraries everywhere. MySQL AB is an example of a company which was extremely successful with this model; it was acquired by Sun Microsystems for about a billion dollars. (See this article for an overview of how GPL licensing fared in Oracle's subsequent acquisition of Sun.)

The GPL makes it very difficult, however, for anyone to change licensing terms after software has been released into the world. That's because of the way copyright law determines the copyright holder of derivative works. In general, if I take a piece of software that you have written, and modify it in such a way that involves creative effort on my part, then I own the copyright to the changes that I've made and the resulting work is a derivative work that both you and I have a copyright interest in. Even though you are the original copyright holder, you would need my permission to release the derivative work under any license other than the GPL.

Unlike Index Data's software, Koha has included significant contributions from many developers, including many who have never worked for LibLime or Katipo Communications (which sold its copyrights to Koha source code to LibLime in 2007). So although LibLime probably owned clear copyright to a majority of Koha at some point, Koha is still a collective work locked into GPL, version 2, and it is unlikely that LibLime or PTFS would be able to distribute Koha under terms other than GPL without doing a thorough rewrite of the software.

Trademarks are a different story. LibLime owns the US trademark for Koha; a European trademark is held by BibLibre. Trademarks are frequently used by open source projects to prevent splintering. The excellent primer on legal issues by the Software Freedom Law Center puts it this way:
FOSS applications develop reputations over time as users come to associate an application’s name with a particular standard of quality or set of features. Trademark law can help protect this relationship of trust and reliance that a project develops with its users; it allows the project to maintain a certain amount of control over the use of its brand.
Since GPL and other FOSS licenses allow anyone to modify and distribute software as long as the license conditions are met, they frequently spawn variants. The owner of a trademark can prevent these variants from using the trademarked name, and thus enforce unity in a project.

In the case of Koha, there are currently two parallel tracks of development being pursued, one inside LibLime, and the other by the community of developers outside LibLime. I will have to postpone a discussion of the issues surrounding  these development tracks to yet another article, but for now, let's just assume there will be two main versions of Koha, LibLime Koha and Community Koha. In the US, LibLime could theoretically prevent anyone from  using the name "Koha" without its authorization, and could strip Community Koha of the right to use "Koha" in its name. In fact, LibLime and BibLibre threatened to use this power a year ago to regulate PTFS's use of Koha trademarks in the marketing of its Koha support services. Liblime could even apply the Koha name to non-Open Source software. Similarly, BibLibre could regulate the use of the Koha name in Europe, preventing LibLime from marketing LibLime Koha there.

Based on discussions I've had with the leaders of almost every open-source library system company, I think it is unlikely that there will be any such "trademark war". Even if the development of Koha continues on separate but related tracks, the success of every Koha-based company is tied to the success Koha as a whole, and vice versa. It would be advantageous for every stakeholder if the two trademark owners develop some sort of "big tent" system of Koha trademark governance. Assuming PTFS's acquisition of LibLime is completed, such governance will need to be acceptable to both PTFS and BibLibre, and will need to accommodate differing styles of software development.

Until a general agreement on the use of Koha trademarks is reached, Koha stakeholders would be well advised to recognize that collective copyrights tie them into the same canoe and that they should avoid disturbing the taniwha that guards and protects them.

This article is the second part of a series. Part 1 is here. Part 3 is here
Reblog this post [with Zemanta]

Article any source

Thursday, January 21, 2010

PTFS to Acquire LibLime and Move to Library Systems Premier League

Update Feb.12 - the acquisition is not happening
Update Mar. 16- the acquisition closed after all.

In 2009, the New York Yankees had a payroll almost ten times that of the Florida Marlins. The reason that baseball lives with that disparity is that the financial interests of many owners do not align with their fans- they take in roughly the same amount of money no matter what the team's performance.

It's different in the English Football Leagues. Teams which fall to the bottom of the standings in the Premier League are relegated to the second division, the equivalent of baseball's minor leagues. At the same time, the best teams in the second division are promoted to the Premier League, giving them a chance to make much more money. There is a clear alignment between the interests of the fans and the owners.

The library industry has likewise been troubled by misalignment of interests between the owners of the companies and their customers. That's why it's important for libraries to pay close attention to the frequent mergers and acquisitions of the companies that serve them. These transactions are often announced just before an ALA meeting, and this past weekend's ALA Midwinter Meeting was no exception.

The big story of the weekend was the pending acquisition of Koha support vendor LibLime by PTFS (Progressive Technology Federal Systems, Inc.). (The acquisition is still in the due diligence phase and is expected to close in early February; terms were not disclosed.) The surprising part of the announcement was the sudden emergence of PTFS, which has had a very low profile in the library industry, into the top tier of integrated library system vendors.

Here I must digress to discuss a bit about business models in the library industry. Libraries have traditionally viewed their catalog system vendors as long term partners; the migration of data from one system to another is a major project, not lightly taken, and preferably not attempted more than once a decade. The choice of a new system touches almost all the library's processes, and thus involves many consultations and lengthy RFPs.

From the vendor's point of view, the sales process is very expensive. Promises to customize the system to address customer peculiarities are common, and these add to the cost of system maintenance. Once the system has been sold, a proprietary system vendor has a guarantee of continuing profits from support contracts. Only the vendor has the system knowledge (and sometimes even the system access) to make even the most trivial changes. It's in the support phase that the vendor and customer interests can become misaligned. The vendor has every incentive to do the least work at the highest price possible. The customer is locked into whatever system they have chosen.

Companies with strong cash flow have been attractive acquisition targets for private equity firms. Once acquired the company's new management focuses on eliminating expenses by cutting support staff and cleaning up the balance sheet by offloading liabilities such as unfinished development, thus making the company very profitable. The company can them be resold at a good mark-up. Customers often become very unhappy during the process. The company they "hired" during their system selection process transforms into something different.

The recent popularity of open source library management systems is in large part a search for business models that better align the interests of vendor and customer during the support phase. If the support vendor doesn't perform to the library's expectations, the library can hire a new support vendor without ditching their automation system. If a library wants to add a new feature to their system, or integrate it with a system from another vendor, they can hire a developer based on qualifications rather than access to source. The important thing to the library is not so much the access to source or the cost of the license, it's the absence of vendor lock-in.

The reason that PTFS is not widely known is that it specializes in an obscure segment of the market- it supports libraries predominantly in the government and the military. Founded in 1995, PTFS has been installing ILS systems, doing conversions and supporting systems in the unique security environment of government systems. John Yokley, a co-founder and the CEO of PTFS, spent 13 years as a Sirsi system administrator and programmer at the U.S. Courts, NASA, and University of Virginia Health Sciences Library has spent 20 years, 5 more than the age of PTFS, working in the library industry. Yokley himself worked in a government library for a short period in the early 90’s designing and building virtual library technology.. The company has experienced steady 20% per year growth and today has 120 employees. PTFS is particularly proud of their development of the US Government Printing Office's Federal Digital System (FDsys) which supports over a thousand libraries, but the company also has a library staffing component and a digitization facility.

Although PTFS has had a strategic partnership with SirsiDynix to market ArchivalWare, a digital content management system that grew out of technology developed for FDsys the Naval Research Laboratories (NRL) TORPEDO project, it found itself hamstrung in supporting its customers because of the lack of access to source code of the proprietary systems it was supporting. About 18 Months ago, PTFS decided that Koha was the Integrated Library System that it could most easily integrate with ArchivalWare, and it began to offer support for Koha. Koha is generally considered to be the first open source integrated library system; it was initially developed in New Zealand by Katipo Communications Ltd. and first deployed in January of 2000 for Horowhenua Library Trust.

LibLime (which is actually a trade name of Columbus, Ohio based Metavore, Inc.) was started in 2005 by Joshua Ferraro, Tina Berger and two others. LibLime has been the hardest-charging and fastest-growing proponent of the Koha Library System in the world. Over the intervening years, LibLime has acquired key Koha-related assets, including the US trademark, copyrights to Koha source code, and the Koha website. The combination of PTFS and LibLime will be supporting 640 installations of Koha under 123 contracts. The combined business will have Koha-related development contracts totaling $1.7 million. Despite the state of the economy, LibLime has actually had an increase in business over the past few months.

Recently, Ferraro and his co-principals at Metavore became very interested and excited by an opportunity outside of the library space. As the LibLime business grew, they recognized that they couldn't pursue both the new opportunity and LibLime, and they began to look for an acquisition partner. PTFS was the first company they went to. Given the reasons for the sale, only Ferraro among the Metavore principals will remain with LibLime; and he will stay only for 18 months to oversee the completion of planned development.

PTFS will keep the LibLime name and fold its own Koha support business into LibLime, which will be run by Patrick Jones. At the press conference held at ALA Midwinter in Boston, PTFS CEO John Yokley indicated that PTFS was committed to the concept of user-driven development and the open source concept, but also emphasized that he was still learning about open source and he was reviewing the LibLime business model; there is much left to be decided about how the LibLime business will move forward.

I spoke with Yokley afterwards. In his conversations with LibLime customers, he has found that their top priority for adopting Koha was to avoid vendor lock-in: their systems should be expandable by LibLime, the library or by another vendor. He sees Koha as a component of a fully capable integrated library system, and vowed that in two years, Koha will be fully capable of running a major academic library. The integration of Koha and ArchivalWare will be only the first phase. Although his team has discussed making ArchivalWare into an open source project, there are issues with third party components used which may prevent that from happening.

Yokley's clarity on avoiding vendor lock-in will be reassuring to customers, particularly with respect to LibLime Enterprise Koha (LLEK), a service announced by LibLime in September of 2009. LLEK is perhaps the most exciting asset being acquired by PTFS, and also the most controversial. The controversy deserves another article entirely, as it represents a break between LibLime and other developers supporting Koha. I plan to write that article in the coming week; please e-mail me if you wish to comment.

LLEK represents the evolution of LibLime's entry into cloud computing (also known as "software-as-a-service". Unlike vendors whose idea of cloud computing is simply to offer fully hosted services, LibLime's implementation of the cloud is more in line with that of modern "lean startups" who don't even own their own servers. By using Amazon EC2, LibLime has access to instantly expandable, low cost computing resources. LibLime is able to provision, configure, and implement a new Koha server in less than an hour. To accomplish this, LibLime has developed sophisticated deployment software (which it does not intend to release).

PTFS is already doing a sort of software-as-a-service, building private clouds for its military customers who don't have the option of going out on the open internet. As Yokley explained to me, "Economies of scale are an interesting thing. We've had a few large customers, but now with LibLime, we can provide services to large numbers of small libraries."

Welcome to the big leagues, PTFS!

This article is the first part of a series. Part 2 is here. Part 3 is here.

Reblog this post [with Zemanta]

Article any source

Monday, January 18, 2010

Google Exposes Book Metadata Privates at ALA Forum

At the hospital, nudity is no big deal. Doctors and nurses see bodies all the time, including ones that look like yours, and ones that look a lot worse. You get a gown, but its coverage is more psychological than physical!

Today, Google made an unprecedented display of its book metadata private parts, but the audience was a group of metadata doctors and nurses, and believe me, they've seen MUCH worse. Kurt Groetsch, a Collections Specialist in the Google Books Project presented details of how Google processes book metadata from libraries, publishers, and others to the Association for Library Collections and Technical Services Forum during the American Library Association's Midwinter Meeting.

The Forum, entitled "Mix and Match: Mashups of Bibliographic Data", began with a presentation from OCLC's Renée Register, who described how book metadata gets created and flows though the supply chain. Her blob diagram conveyed the complexity of data flow, and she bemoaned the fact that library data was largely walled off from publisher data by incompatible formats and cataloging practice. OCLC is working to connect these data silos.

Next came friend-of-the-blog Karen Coyle, who's been a consultant (or "bibliographic informant") to the Open Library project. She described the violent collision of library metadata with internet database programmers. Coyle's role in the project is not to provide direction, but to help the programmers decode arcane library-only syntax such as "ill. (some col)". The one instance where she tried to provide direction turned out to be something of a mistake. She insisted that, to allow proper sorting, the incoming data stream should try to keep track of the end of leading articles in title strings. So for example, "The Hobbit" should be stored as "(The )Hobbit". This proved to be very cumbersome. Eventually the team tried to figure out when alphabetical sorting was really required, and the answer turned out to be "never".

Open Library does not use data records at all, instead, every piece of data is typed with a URI. This architecture aligns with W3C web standards for the semantic web, and allows much more flexible searching and data mining than would be possible with a MARC record.

Finally, Groetsch reported on Google's metadata processing. They have over 100 bibliographic data sources, including libraries, publishers, retailers and aggregators of review and jacket covers. The library data includes MARC records, anonymized circulation data and authority files. The publisher and retailer data is mostly ONIX formatted XML data. They have amassed over 800 million bibliographic records containing over a trillion fields of data.

Incoming records are parsed into simple data structures which looked similar to Open Library's, but without the URI-ness. These structures are than transformed in various ways for Googles use. The raw metadata structures are stored in an SQL-like database for easy querying.

Groetsch then talked about the nitty-gritty details of data. For example, the listing of an author on a MARC record can only be used as an "indication" of the authors name, because MARC gives weak indications of the contributor role. ONIX is much better in this respect. Similarly, "identifiers" such as ISBN, OCLC number, LCCN, and library barcode number are used as key strings but are only identity indicators with varying strengths. One ISBN with a chinese publisher prefix was found on records for over 24,000 different books; ISBN reuse is not at all uncommon. One librarian had mentioned to Groetsch that in her country, ISBNs are pasted onto a book to give it a greater appearance of legitimacy.

Echoing comments from Coyle, Groetsch spoke with pride of the progress the Google Books metadata team has made in capturing series and group data. Such information is typically recorded in mushy text fields with inconsistent syntax, even in records from the same library.

The most difficult problem faced by the Google Books team is garbage data. Last year, Google came under harsh criticism for the quality of its metadata, most notably from Geoffrey Nunberg. (I wrote an article about the controversy.) The most hilarious errors came from garbage records. For example, certain Onix records describing Gulliver's Travels carried an author description of the wrong Jonathan Swift. Most of these errors come from garbage records, and when one of these is found, almost always, the same problems can be found in other metadata sources. Google would like to find a way to get corrected records back into the library data ecosystem so that they don't have to fix them again, but that there have been issues with data licensing agreements that still need to be worked out. Article like Nunberg's have been quite helpful to the Google team. Every indication is that Google is in the metadata slog for the long term.

One questioner asked the panel what the library community should be doing to prevent "metadata trainwrecks" from happening in the future. Groetsch said without hesitation "Move away from MARC". There was nodding and murmuring in the audience (the librarian equivalent of an uproar). He elaborated that the worst parts of MARC records were the free text data, and normalization of data would be beneficial whereever possible.

One of the Google engineers working on record parsing, Leonid Taycher, added that the first thing he had had to learn about MARC records was that the "Machine Readable" part of the MARC acronym was a lie. (MARC stands for MAchine Readable Cataloging) The audience was amused.

The last question from the audience was about the future role of libraries in production of metadata. Given the resources being brought to bear on the book metadata by OCLC, Google and others, should libraries be doing cataloguing at all? Karen Coyle's answer was that libraries should concentrate their attention on the rare and unique material in their collections- without their work, these materials would continue to be almost completely invisible.
Reblog this post [with Zemanta]

Article any source