Showing posts with label ALA Annual. Show all posts
Showing posts with label ALA Annual. Show all posts

Sunday, July 1, 2012

Secret Desert Meeting Report: It Was Hot


Don't think that a relative absence of blog posts means I haven't been working hard.


Unglue.it was a first-time exhibitor at the American Library Association Annual Meeting. We talked to a lot of people. We gave out stickers. I gave a talk. I met a mermaid. Here are my observations:

  • I'll bet we talked to 200 people. And of those 200, there were about 5 who weren't excited to learn about what Unglue.it is doing. That's a pretty good batting average.
  • Of the 200, half said they'd heard about us or read about us somewhere. About 5 of those had actually supported one of the campaigns. And those 5 were bringing Andromeda cookies to keep her from expiring. That's hitting a lot of singles but not scoring a lot of runs. But I'm happy to report that the unglue.it team survived.
  • The total attendance at ALA was around 20,000. That means we barely reached 1% of attendees. But they were a very good 1%. The number of people who have signed up to be ungluers passed 1,400. Ungluers pledged over $1500 to the campaign for Joe Nassise's Riverwatch. But Joe will make much more than that by leaving the ebook on Amazon, so it won't be our next book to unglue. We need to multiply our numbers by 10. And we will. Our 195 converts will talk to 195 more and so on. We have enough stickers.
  • The mermaid was also excited about Unglue.it.
  • We had some discussions that will lead to amazing things. 

After the meeting was over, The unglue.it team converged on an undisclosed desert location to figure out what to do next. We decided to get some sleep.

It gets hot in the desert. 100°F by noontime. But that didn't stop us from our mission. We had electronics. We had internet.  We had expense reimbursement policy to discuss. The words "Gold Lamé" were mentioned. World domination was considered inevitable, given enough time.

And speaking of a good batting average, Carlos Ruiz of the Philadelphia Phillies was named to the National League All-Star team today, and the honor has never been more deserved. Major-league catchers crouch and stand about 200 times a game, catching 90 mile-per-hour fast balls and curves that make the air hum. Or not, and then they have to block the ball with their bodies. They are routine smash targets for runners trying to score through their bodies. On top of this, they're expected to throw out runners, manage their pitchers,  and know the weaknesses of the opposing hitters. Oh, and hit occasionally. This year, Ruiz is leading the league in batting average, and to use a techniccal term, mashing.

There was a time when Ruiz was considered too old to be a prospect. He just needed some time.

Awesomeness takes its own time.
Article any source

Tuesday, June 12, 2012

Digital Content Working Group at ALA


I've been very busy with unglue.it and I'm working on something I'll soon share here. We're thrilled that the unglue.it campaign for Ruth Finnegan's Oral Literature in Africa is at 53% but there are just 9 days to go, so if you haven't yet casted your vote for crowd-funded Creative Commons ebook relicensing, now is the time you most do so.

Since early this year, I've been serving on ALA’s Digital Content and Libraries Working Group (DCWG) commisioned by American Library Association (ALA) President Molly Raphael. Our role has been to advise the ALA leadership about the changes in libraries and publishing associated with the transition to digital content, and we've worked to articulate the concerns of libraries in ways that can be acted upon by the entire digital content ecosystem.

Later this month, I'll be speaking on a panel organized by DCWG at the ALA Annual Meeting in Anaheim.
Access to Digital Content: Diverse ApproachesSunday, June 24, 1:30–3:30 p.m., Anaheim Hilton, California B2012 ALA Annual Conference
              As digital content continues to grow in diversity and importance, libraries must make use of multiple strategies to support access for their users. ALA’s Digital Content and Libraries Working Group has been exploring issues of business models, advocacy, education, accessibility, privacy, and libraries as providers of content.  But innovative approaches to making digital content available are taking place in many arenas.
              Come hear about the latest developments, including an update from Working Group co-chairs Sari Feldman (Cuyahoga County Public Library) and Robert Wolven (Columbia University). Leaders and innovators from the library community will discuss some other major initiatives and developments related to digital content:  Peter Brantley (Internet Archive), Maura Marx (Digital Public Library of America), and Eric Hellman (Unglue.it). Lee Rainie will offer some perspectives based on his work with the Pew Research Center’s Internet and American Life Project, and Robert Wolven will offer final remarks to set the stage for the question and answer period.
              Copies of the new American Libraries publication “E-Content: The Digital Dialogue” will be available.
Also, Unglue.it is going to have a table at the exhibits. You can meet me, Andromeda and/or Amanda in person. We'll be in the "small press and new exhibitors" section, which is usually in the far reaches of the exhibit hall, but there will be unglue.it bookmarks and stickers as your reward if you can find us.

Enhanced by Zemanta

Article any source

Sunday, July 31, 2011

Library Data Beyond the Like Button

"Aren't you supposed to be working on your new business? That ungluing ebooks thing? Instead you keep writing about library data, whatever that is. What's going on?"

No, really, it all fits together in the end. But to explain, I need to talk you beyond the "Like Button".

Earlier this month, I attended a lecture at the New York Public Library. The topic was Linked Open Data, and the speaker was Jon Voss, who's been applying this technology to historical maps. It was striking to see how many people from many institutions turned out, and how enthusiastically Jon's talk was received. The interest in Linked Data was similarly high at the American Library Association Meeting in New Orleans, where my session (presented with Ross Singer of Talis) was only one of several Linked Data sessions that packed meeting rooms and forced attendees to listen from hallways.

I think it's important to convert this level of interest into action. The question is, what can be done now to get closer to the vision of ubiquitous interoperable data? My last three posts have explored what libraries might do to better position their presence in search engines and in social networks using schema.org vocabulary and Open Graph Protocol. In these applications, library data enables users to do very specific things on the web- find a library page in a search engine or "Like" a library page in a Facebook. But there's so much more that could be done with the data.

I think that library data should be handled as if it was made of gold, not of diamond.

Perhaps the most amazing property of gold is its malleability. Gold can be pounded into a sheet so thin that it's transparent to light. An ounce of gold can be made into leaf that will cover 25 square meters.

There is a natural tendency to treat library data as a gem that needs skillful cutting and polishing. The resulting jewel will be so valuable that users will beat down library websites to get at the gems. Yeah.

The reality is that  library data in much more valuable as a thin layer that covers huge swaths of material. When data is spread thinly, it has a better chance of connecting with data from other libraries and with other sorts of institutions: Museums, archives, businesses, and communities. By contrast, deep data, the sort that focuses on a specific problem space, is unlikely to cross domains or applications without a lot of custom programming and data tweaking.

Here's the example that's driven my interest in opening up library linked data: At Gluejar, we're building a website that will ask people to go beyond "liking" books. We believe that books are so important to people that they will want to give them to the world; to do that we'll need to raise money. If lots of people join together around a book, it will be easy to raise the money we need, just as public radio stations find enough supporters to make the radio free to everyone.

We don't want our website to be a book discovery website, or a social network of readers, or a library catalog; other sites to that just fine. What we need is for users to click "support this book" buttons on all sorts of websites, including library catalogs. And our software needs to pull just a bit of data off of a webpage to allow us to figure out which book the user wants to support. It doesn't sound so difficult. But we can only support to or three different interfaces to that data. If library websites all put a little more structured data in their HTML, we could do some amazing things. But they don't, and we have to settle for "sort of works most of the time".

Real books get used in all sorts of ways. People annotate them, they suggest them to friends, they give them away, they quote them, and they cite them. People make "TBR" piles next to their beds. Sometimes, they even read and remember them as long as they live. The ability to do these same things on the web would be pure gold.

Article any source

Monday, July 11, 2011

Spoonfeeding Library Data to Search Engines

CC-NC-BY rocketship
When you talk to a search engine, you need to realize that it's just a humongous baby. You can't expect it to understand complicated things. You would never try to teach language to a human baby by reading it Nietzsche, and you shouldn't expect a baby google to learn bibliographic data by feeding it MARC (or RDA or METS or MODS, or even ONIX).

When a baby says "goo-goo" to you, you don't criticize its misuse of the subjunctive. You say "goo-goo" back. When Google tells you that that it wants to hear "schema.org" microdata, you don't try to tell it about the first indicator of the 856 ‡u subfield. You give it schema.org microdata, no matter how babyish that seems.

It's important to build up a baby's self-confidence. When baby google expresses interest in the number of pages of a book, you don't really want to be specifying that there are ix pages numbered with roman numerals and 153 pages with arabic numerals in shorthand code. When baby google wants to know whether a book is "family friendly" you don't want to tell it about 521 special audience characteristics, you just want to tell it whether or not it's porn.

If you haven't looked at the schema.org model for books, now's a good time. Don't expect to find a brilliant model for book metadata, expect to find out what a bibliographic neophyte machine thinks it can use a billion times a day. Schema.org was designed by engineers from Google, Yahoo, and Bing. Remember, their goal in designing it was not to describe things well, it was to make their search results better and easier to use.

The thing is, it's not such a big deal to include this sort of data in a page that comes from an library OPAC (online catalog). An OPAC that publishes unstructured data produces HTML that looks something like this:
<div> 
<h1>Avatar (Mysteries of Septagram, #2)</h1>
<span>Author: Paul Bryers (born 1945)</span>
<span>Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>
The first step is to mark something as the root object. You do that with the itemscope attribute:
<div itemscope> 
<h1>Avatar</h1>
<span>Author: Paul Bryers (born 1945)</span>
<span>Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>

A microdata-aware search engine looking at this will start building a model. So far, the model has one object, which I'll denote with a red box.


The second step, using microdata and Schema.org, is to give the object a type. You do that with the itemtype attribute:
<div itemscope itemtype="http://schema.org/Book"> 
<h1>Avatar (Mysteries of Septagram, #2)</h1>
<span>Author: Paul Bryers (born 1945)</span>
<span>Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>

Now the object in the model has acquired the type "Book" (or more precisely, the type "http://schema.org/Book".

Next, we give the Book object some properties:
<div itemscope itemtype="http://schema.org/Book"> 
<h1 itemprop="name">Avatar (Mysteries of Septagram, #2)</h1>
<span>Author: 
<span itemprop="author">Paul Bryers (born 1945)</span></span> 
<span itemprop="genre">Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg">
</div>

Note that while the library record for this book attempts to convey the title complexity: "245 10 $aAvatar /$cPaul Bryers.$", the search engine doesn't care yet. The book is part of a series: 490 1 $aThe mysteries of the Septagram$, and the search engines don't want to know about that either. Eventually, they'll learn.
The model built by the search engine looks like this:

So far, all the property values have been simple text strings. We can also add properties that are links:
<div itemscope itemtype="http://schema.org/Book"> 
<h1 itemprop="name">Avatar (Mysteries of Septagram, #2)</h1>
<span>Author: 
<span itemprop="author">Paul Bryers (born 1945)</span></span> 
<span itemprop="genre">Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg" 
itemprop="image">
</div>
The model grows.

Finally, we want to say that the author, Paul Bryers, is an object in his own right. In fact, we have to, because the value of an author property has to be a Person or an Organization in Schema.org. So we add another itemscope attribute, and give him some properties:
<div itemscope itemtype="http://schema.org/Book"> 
<h1 itemprop="name">Avatar (Mysteries of Septagram, #2)</h1>
<div itemprop="author" itemscope itemtype="http://schema.org.Person">
Author:  <span itemprop="name">Paul Bryers</span> 
(born <span itemprop="birthDate">1945</span>)
</div>
<span itemprop="genre">Science fiction</span>
<img src="http://coverart.oclc.org/ImageWebSvc/oclc/+-+703315758_140.jpg" 
itemprop="image">
</div>

That wasn't so hard. Baby has this picture in his tyrannical little head:

Which it can easily turn into a "rich snippet" that looks like this:

Though you know all it really cares about is milk.

Here's a quick overview of the properties a Schema.org/Book can have (the values in parentheses indicate a type for the property value):

Properties from http://schema.org/Thing
  • description
  • image(URL)
  • name
  • url(URL)
Properties from http://schema.org/CreativeWork
Properties from http://schema.org/Book
This post is the second derived from my talk at ALA in New Orleans. The first post discussed the changing role of digital surragates in a fully digital world. The next will discuss "Like" buttons.
Enhanced by Zemanta

Article any source

Friday, July 8, 2011

Library Data: Why Bother?

When face recognition came out in iPhoto, I was amused when it found faces in shrubbery and asked me whether they were friends of mine. iPhoto, you have such a sense of humor!

But then iPhoto looked at this picture of a wall of stone faces in Baoding, China. It highlighted one of the faces and asked me "Is this Jane?" I was taken aback, because the stone depicted Jane's father. iPhoto was not as stupid as I thought it was- it could even see family resemblances.

Facial recognition software is getting better and better, which is one reason people are so worried about the privacy implications of Facebook's autotagging of pictures. Imagine what computers will be able to do with photos in 10 years! They'll be able to recognize pictures of bananas, boats, beetles and books. I'm thinking it's probably not worth it to fill in a lot of iPhoto metadata.

I wish I had thought about facial recognition when I was preparing my talk for the American Library Association Conference in New Orleans. I wanted my talk to motivate applications for Linked Open Data in libraries, and in thinking about why libraries should be charting a path towards Linked Data, I realized that I needed to examine first of all the motivation for libraries to be in the bibliographic data business in the first place.

Originally, libraries invested in bibliographic data to help people find things. Libraries are big and have a lot of books. It's impractical for library users to find books solely by walking the stacks, unless the object of the search has been anticipated by the ordering of books on the shelves. The paper cards in the card catalog could be easily duplicated to enable many types of search in one compact location. The cards served as surrogates for the physical books.

When library catalogs became digital, much more powerful searches could be done. The books acquired digital surrogates that could be searched with incredible speed. These surrogates could be used for a lot of things, including various library management tasks, but finding things was still the biggest motivation for the catalog data.

We're now in the midst of a transition where books are turning into digital things, but cataloging data hasn't changed a whole lot. Libraries still need their digital surrogates because most publishers don't trust them with the full text of books. But without full text, libraries are unable to provide the full featured discovery that a a search engine with access to both the full text and metadata (Google, Overdrive, etc.) can provide.

At the same time, digital content files are being packed with more and more metadata from the source. Photographs now contain metadata about where, when and how they were taken; for a dramatic example of how this data might be used, take a look at this study from the online dating site OKCupid. Book publishers are paying increased attention to title-level metadata, and metadata is being built into new standards such as EPUB3. To some extent, this metadata is competing for the world's attention with library-sourced metadata.

Libraries have two paths to deal with this situation. One alternative is to insist on getting the full text for everything they offer. (Unglued ebooks offer that, that's what we're working on at Gluejar.)

The other alternative for libraries is to feed their bibliographic data to search engines so that library users can discover books in libraries. Outside libraries, this process is known as "Search Engine Optimization". When I said during my talk that this should be the number one purpose of library data looking forward, one tweeter said it was "bumming her out". If the term "Search Engine Optimization" doesn't work for you, just think of it as "helping people find things".

Library produced data is still important, but it's not essential in the way that it used to be. The most incisive question during my talk pointed out that the sort of cataloging that libraries do is still absolutely essential for things like photographs and other digital archival material. That's very true, but only because automated analysis of photographs and other materials is computationally hard. In ten years, that might not be true. iPhoto might even be enough.

In the big picture, very little will change: libraries will need to be in the data business to help people find things. In the close-up view, everything is changing- the materials and players are different, the machines are different, and the technologies can do things that were hard to imagine even 20 years ago.

In a following post, I'll describe ways that libraries can start publishing linked data, feeding search engines, and keep on helping people find stuff. The slides from my talk (minus some copyrighted photos) are available as PDF (4.8MB) and PPTX (3.5MB).

Enhanced by Zemanta

Article any source

Wednesday, June 29, 2011

3M's eBook Cloud Library Didn't Come Out of Nowhere!

When the Douglas County Libraries in Colorado installed self check-in stations a while ago, they realized that hey had an opportunity to restructure their space. The circulation desk that dominated the main entrance was no longer needed. It seemed obvious to Library Director Jamie LaRue what to put in its place. Libraries need to greet their visitors with displays of books available for immediate checkout. 80% of Douglas County's adult circulation is generated by visual displays of books, so the best way to entice visitors to read is to show them great books to read.

When Douglas County began investigating how to put ebooks into county resident's computers, they wanted to do something similar. A user looking for ebooks should be greeted with a virtual bookshelf of books waiting to be checked out. LaRue was not satisfied with the offering of industry leader Overdrive because he couldn't do such a simple thing.

Public libraries that offer ebooks are frequently faced with problems posed by the strong demand for ebooks. Their users are frequently disappointed that the ebooks they want are always checked out. Overdrive has not yet implemented an programming interface that would allow library catalogs to check on an ebook's availability before showing it to a user, so the process of finding an available ebook can involve a lot of tedious clicks.

To address these needs, Overdrive has announced the "Overdrive WIN" service, which will address better integration with library automation software along with a host of other improvements and service innovations.

I spoke with a number of library automation vendors at this past weekend's American Library Association meeting in New Orleans. eBook integration is high on the list of their customers' wish lists, but I couldn't find any that could tell me when they would be implementing better Overdrive integration, though many of them were in "discussions".

A new vendor worth mentioning was Toronto-based BiblioCommons, whose EC2-cloud-based OPAC service has been implemented by Seattle Public Library and is in beta with New York Public Library. I'd been hearing about BiblioCommons for long enough that I'd had my doubts as their reality. At ALA, they demoed a clean, modern web interface with plenty of social features- go take a look at Seattle Public. Given NYPL's status as a prominent Overdrive customer and Bibliocommons' actively developing codebase, I had hoped to see some preview glimpses of Overdrive WIN in BiblioCommons, but had no such luck.

Back in Douglas County, Jamie LaRue wasn't satisfied with the available options, so around the end of 2010, he had his team approach their auto-check-in vendor, 3M, to see if they could do something about ebooks. As luck would have it, they could. And they did.

Although 3M's entrance into the library ebook platform business came as a complete surprise to many in libraries and publishing, it seems obvious in retrospect. 3M's RFID tag, self-checkout/checkin, and detection businesses were already integrated with library automation systems, so much of the code needed to integrate to library systems was already written. 3M licensed ebook reader and DRM systems from Adobe, and in the space of six months, with the advice and help of customers such as Douglas County, was able to assemble a strong set of services it is branding as the "3M Cloud Library". These include reader software for iOS and Android, as well as spiffy "3M Discovery Terminals", electronic kiosks "with an intuitive touch-based interface". (pictured) 3M is even going to sell "white-label" eReader devices with software tweaked to meet the needs of libraries that want to lend devices.

While 3M is arguably breaking new ground in integration of ebooks with library systems, 3M is far behind Overdrive in the area of publisher relations, which can't just be switched on in a mere 6 months. Overdrive has announced expansions of its offerings in the school and academic markets. Meanwhile, 3M is going in publishers' back doors as it helps the State of Kansas withdraw from an awkwardly drafted Overdrive contract, which Kansas says allows them to move purchased content from Overdrive to other platforms. It's in publishers' interests to have a library ebook channel that competes with Overdrive, but they do SO like to be asked permission first.

For his part, LaRue just wants to be able to tailor his library service to the needs of his community. "I want to provide a quality, integrated experience with a local focus" is what he told me. That doesn't seem to be asking so much.

Update 6/30/11: At The Digital Reader, Nate Hoffelder reported in May that a lot of 3M's reading platform was sourced from txtr, a German start-up they'd invested in. I wasn't able to confirm this at ALA, but have since done so. The Adobe DRM implementation, reading software, apps, presentation interfaces all originated in txtr. I'm also told by multiple sources that 3M has been talking to publishers since at least December 2010.
Enhanced by Zemanta

Article any source

Monday, June 27, 2011

Four Times Around the Library World


The New Orleans Convention Center is sandwiched between the warehouse district and some railroad tracks, and as a result, it's a kilometer long, end to end. This past weekend, it has hosted the American Library Association Annual Meeting. I've walked the length of the convention center about 10 times over the past 4 days. It's another kilometer from my hotel to it's near end, so add another 8 km to my total. Bourbon Street is 1.3 km down and back; so add 3 km there. There were 2.5 km of exhibits on the show floor at ALA; I make it a point to look at every one, at least briefly. So my ALA pedometer racked up about 25 km (over 15 miles, for the metrically challenged).

It's not over yet, but the ALA conference twitter feed says this week's attendance is over 20,000, including exhibitors. (Update: the final totals are 14,969 attendees and 5,217 exhibitors.) Their mileage may vary, but my estimate is that on average, an ALA attendee walked about 5 miles in total. So the grand total of walking at ALA should be about 160,000 km. That's 4 times the circumference of the earth.

All that walking is good for us. I replenished many of those calories at Cochon, where I hosted some lunches to tell librarians about Gluejar. But the Buttermilk Pecan Tart I had on Friday was worth the whole trip to New Orleans. The pleasure capital of Louisiana has moved a mile south as far as I'm concerned!

Cochon brought back memories of 5 years ago, when ALA was the first big convention to come to ALA after Hurricane Katrina. Cochon had opened just a week before, and I raved to friends after having oven-roasted oysters there. By the end of ALA 2006, the place was packed.

The meeting five years ago was a special one; the city was far from having being repaired or rebuilt, and many workers had been bussed in and bunked in temporary housing just so we could come. Everyone was just so happy to see us, it brings tears to my eyes just thinking about it. In New Orleans they still remember the weekend that librarians brought the city of New Orleans back to life.

Article any source

Wednesday, July 14, 2010

What IS an eBook, Anyway?

One of my secret pleasures at American Library Association meetings is going to Standards sessions. Now before you think I have a completely hopeless case of nerdiness, let me explain myself.

There's never just one Standards session at ALA, there are at least two and often three or more. I'm not sure why, but I think it's because librarians feel that standards are Important, and because there are so many Standards in the library world that people forget which ones were the subject of a Standards session at the last meeting. Its not that librarians are interested in Standards, it's just that they have lots of data problems that might magically go away, if only there were a Standard. Or not.

Because there are so many session on Standards, each one tends to be sparsely attended. That's why I like them. You can go and sit in a room with some really smart and influential people (the panelists), ask them bizarre Standards questions, have some other really smart audience member join in the discussion, and feel like you're a member of some hidden clique of powerful numerologists.

One of the things that has the Standards people concerned this year is the way ISBNs are being applied to ebooks. At the session I went to, Brian Green, the Executive Director of the International ISBN Agency, was giving his standard ISBN Standards update. Brian has been doing this long enough that he expects and parries my pestering questions with aplomb.

So here's this year's burning question: How many ISBN's should be issued when ebooks are published in different formats? Should the ebook have the same ISBN as the print book? If a different ISBN, should different file formats get separate ISBNs?

And here's the burning answer from ISBN International: each ebook file format for a book should get its own ISBN:
Do different formats of an electronic or digital publication (e.g., .pdf, .html) need separate ISBNs?

Different formats of an electronic or digital publication are regarded as different editions and therefore need different ISBNs in each instance when they are made separately available
And here's the language of the Standard itself, (ISO 2108:2005) adopted through the international standards process in 2005:
Each different format of an electronic publication (e.g. ".lit", ".pdf", ".html", ".pdb") that is published and made separately available shall be given a separate ISBN.
So forgive me for having been confused in March, when I read that the “E-book ISBN Mess Needs Sorting Out,” Say UK Publishers. Why are the publishers still talking about this, more than ten years after the question was raised and thoroughly discussed? Why are we having panels at ALA to learn about this? Has the numeracy of the world's book industry been entirely depleted during ISBN's switch to 13 digits???

ISBN stands alone in the world of identifiers because of its widespread pre-internet adoption and success. Even the Internet Engineering Task Force set aside some URI space for it back in the days before "HTTP" became a religious invocation. But most people outside the book industry have had no idea of what it really identified- they usually think it identifies a book or perhaps a book version.

If the book industry had a Facebook profile, it would list its relationship with ISBN as it's complicated.  Consider ISBN 978-1593967574. It is a "Year 5 Harry Potter Bust" manufactured by Diamond Comics. It has no author, pages or even words; it is not a book in any sense. Yet it is well-behaved in the ISBN world, because it is (or was) an item distributed by the world's book supply chain to bookstores and ultimately consumers.

In the print world, it is more or less understood that a paperback has a different ISBN from the hardcover, which has a different ISBN from the library-bound version, and may have a different set of ISBNs when issued in a different country. At the deepest level, the ISBN is just a solution to a problem: "How does an item get tracked through the book supply chain?"

If you see a book on the shelves of a bookstore, you can be pretty sure that it got there through the "supply chain". Book publishers don't sell books to book stores, they mostly sell to distributors such as Ingram and Baker & Taylor. Bookstores use ISBNs to order books, and the distributors use the ISBN to report sales back to the publishers. When books don't sell, they get shipped back to warehouses, which track them using...ISBN.

When there's a question about whether a different ISBN should or should not be issued, the overriding principle is "a product needs a separate identifier if the supply chain needs to separately identify it." This clarity about the function of an ISBN is what has resulted in its overwhelming success. When people try to use the ISBN for other things, it's less successful. Supplemental services such as xISBN (which I helped put into production at OCLC), thingISBN, and emerging identifiers such as ISTC are useful for filling in the gaps between what ISBN really is and what people would like it to be.

Let's look at ebooks with the prism of the supply chain. If an ebook is issued in print, PDF and EPUB formats, it's important to the publisher to know how many of each are sold, thus the separate ISBN's. Similarly, if different DRM wrapping is used by two different channels, in many cases the publisher will need to track sales or manage the product separately. Although in many cases the DRM could be tracked by retailer, and thus wouldn't need a separate ISBN, the ISBN Standard says to give it a different ISBN. As Green has written previously,
Where publishers are selling e-books exclusively from their own websites or through another single channel and do not wish to have them listed in books in print databases then [...] publishers may not wish to bother with ISBNs. However, publishers should beware of taking a short-term view that makes them reliant on a single channel.
Unfortunately some publishers have obstinately refused to give separate ISBNs to ebooks in different formats. The US division of Random House is perhaps the most prominent example. There are excellent arguments for the "single ISBN" approach, but the worst possible situation for the emerging supply chain is for each publisher to use their own inconsistent rules for applying ISBN to ebooks. However strong the argument is for "single ISBN", its inconsistent application negates the advantages and threatens the ISBN system as a whole.

The ultimate problem with ISBN and ebooks is that ebooks are sufficiently adaptable that they expose  ambiguities and limitations of the ISBN identification architecture. For example, suppose you're in the business of selling customized digital coursebooks. You allow professors to choose 10 chapters from 100 available. That means there are exactly 17,310,309,456,440 different ebooks that you could sell. That's about 9,000 times more books than can be identified by all the ISBNs in the galaxy. But you don't need to give them ISBNs, because you sell direct and the ebooks never touch the supply chain. The chapters themselves may need to be tracked so you can pay author royalties, but you need only 100 ISBNs to do that.

How about if a retailer changes (or eliminates) the DRM wrapping an ebook? Do the ISBN's of the ebooks on a consumer's ebook reader magically change? (Transubstantiation is one of my favorite words!) The answer is no, and that's because the the supply chain is not involved.

Are there enough ISBNs for the ebooks that could be sold? The EPUB format is actually an archive file format that uses a dialect of XHTML for its insides, so you might imagine that any website or portion thereof can be packaged as an ebook. In fact, BookGlutton has a tool that (sort of) does this. As of May 2009, over 100 million websites operated, so you can easily imagine that ebooks could use up all available ISBN's almost overnight.

The "supply chain" for ebooks is rapidly mutating. The adoption of an "agency model" is an example of a change that has put new demands on ISBN; "agency" requires a retailer to identify an item's publisher before the moment of sale so that the correct sales tax can be applied. The agency model shift won't be the last or biggest change to the ebook supply chain, either. As one example, I've previously written about ebook pay-per-view and demand-driven acquisition. Another huge change would occur if  a substantial advertising revenue stream for ebooks, such as Apple's iAd system, emerges. Advertising would put new demands on reporting systems and thus on the ISBNs that enable them.

What is an ebook anyway? Ten years ago, a committee of the American Association of Publishers came up with this not-so-useful definition:
An ebook is a literary work in the form of a digital object consisting of one or more standard unique identifiers, metadata, and a monographic body of content, intended to be published and accessed electronically.
I'll bet you never realized that blog posts were really ebooks!

The truth is that we really have no idea what an ebook is or what it will become. There are certainly e-things that correspond to print books, and these are easy to recognize as ebooks. But don't be surprised if there comes a flood of things to read on our connected devices that are too long to be called "articles" or "posts". For these, "eBook" may be the best label we can come up with.

Unless of course they get shackled by a supply chain.
Enhanced by Zemanta

Article any source

Sunday, June 27, 2010

Global Warming of Linked Data in Libraries

Libraries are unusual social institutions in many respects; perhaps the most bizarre is their reverence for metadata and its evangelism. What other institution considers the production, protection and promulgation of metadata to be part of its public purpose?

The W3C's Linked Data activity shares this unusual mission. For the past decade, W3C has been developing a technology stack and methodology designed to support the publication and reuse of metadata; adoption of these technologies has been slow and steady, but the impact of this work has fallen short of its stated ambitions.

I've been at the American Library Association's Annual Meeting this weekend. Given the common purpose of libraries and Linked Data, you would think that Linked Data would be a hot topic of discussion. The weather here has been much hotter than Linked Data, which I would describe as "globally warming". I've attended two sessions covering Linked Data, each attended by between 50 and 100 delegates. These followed a day long, sold-out  preconference. John Phipps, one of the leaders in the effort to make library metadata compatible with the semantic web, remarked to me that these meeting would not have been possible even a year ago. Still, this attendance reflects only a tiny fraction of metadata workers at the conference; Linked Data has quite a ways to come. It's only a few months ago that the W3C formed a Library Linked Data Incubator Group.

On Friday morning, there was an "un-conference" organized by Corey Harper from NYU and Karen Coyle, a well-known consultant. I participated in a subgroup looking at use cases for library Linked Data. It took a while for us to get around to use cases though, as participants described that usage was occurring, but they weren't sure what for. Reports from OCLC (VIAF) and Library of Congress (id.loc.gov) both indicated significant usage but little feedback. The VIVO project was described as one with a solid use case (giving faculty members a public web presence), but no one from VIVO was in attendance.

On Sunday morning, a meeting of the Association for Library Collections and Technical Services (ALCTS), Rebecca Guenther, Library of Congress, discussed id.loc.gov, a service that enables both humans and machines to programatically access authority data at the Library of Congress. Perhaps the most significant thing about id.loc.gov is not what it does but who is doing it. The Library of Congress provides leadership for the world of library cataloguing; what LC does is often slavishly imitated in libraries throughout the US and the rest of the world.  id.loc.gov started out as a research project but is now officually supported.

Sara Russell-Gonzalez of the University of Florida then presented the VIVO which has won a big chunk of funding from the National Center for Research Resources, a branch of NIH. The goal of VIVO is to build an "interdisciplinary national network enabling collaboration and discovery between scientists across all disciplines." VIVO started at Cornell and has garnered strong institutional support there, as evidenced by an impressive web site. If VIVO is able to gain similar support nationally and internationally, it could become an important component of an international research infrastructure. This is a big "if". I asked if VIVO had figured out how to handle cases where researchers change institutional affiliations; the answer was "No". My question was intentionally difficult; Ian Davis has written cogently about the difficulties RDF has in treating time-dependent relationships. It turns out that there are political issues as well. Cornell has had to deal with a case where an academic department wanted to expunge affiliation data for a researcher who left under cloudy circumstances.

At the un-conference, I urged my breakout group to consider linked data as a way to expose library resources outside of the library world as well as a model for use inside libraries. It's striking to me that libraries seem so focused on efforts such as RDA, which aim to move library data models into Semantic Web compatible formats. What they aren't doing is to make library data easily available in models understandable outside the library.

The two most significant applications of Linked Data technologies so far are Google's Rich Snippets and Facebook's Open Graph Protocol (whose user interface, the "Like" button, is perhaps the semantic webs most elegant and intuitive). Why aren't libraries paying more attention to making their OPAC results compatable with these application by embedding RDFa annotations in their web-facing systems? It seems to me that the entire point of metadata in libraries is to make collections accessible. How better to do this than to weave this metadata into peoples lives via Facebook and Google? Doing this will require the dumbing-down of library metadata and some hard swallowing, but it's access, not metadata quality, that's core to the reason that libraries exist.



Enhanced by Zemanta

Article any source

Friday, June 25, 2010

Introducing the Totebag for eBooks

As we hurtle towards a future where books come on Kindles and iPads and Nooks, we tend to overlook the loss of many products and services attached to the print book ecosystem. Tens of thousands of people whose livelihoods depend on books will suffer tragic dislocations in their lives. While many bemoan the plight of bookstore workers, librarians, editors, and authors, there are other small industry segments no one ever thinks of.

Totebag manufacturing is just one of these overlooked industries. Half of the world's  novelty totebags for books are manufactured in a single town in China called Shu Bao (书包). Shu Bao is located in an inland area of China that has concentrated on book related products; neighboring towns specialize in bookmarks, dust covers and those little alphabet labels used in dictionary manufacture. At this weekend's American Library Association (ALA) meeting in Washington DC, I had a chance to speak with Shu Bau's mayor, Yi Rui-Da, who doubles as a sort of totebag ambassador and salesman to the world. Yi was in town to start getting the word out about digital book totebags.

Yi told me that the central committee of his town has been closely watching the shift to eReading for at least 10 years. They've seen one of the neighboring towns become quite wealthy by shifting their manufacturing to iPad covers, and hope to make a similar transition themselves. The lesson of what happened to buggy-whip manufacturers after the introduction of the Model T is known to the committee. Some committee members thought the town was in the luggage business, and preferred to stay in the luggage business. Other committee members, aware of the specialized fibers that must be added to their totebag fabrics, argued that the town was really in the information portability business; these voices prevailed.

To make the transition to transporting eBooks, the town had to nurture its programming talent, of which it has an abundance. Totebags are made in factories that employ hundreds of teenage girls. But it's not like the old days, when the girl were virtual slaves, sewing everything by hand. In a modern totebag factory, the girls program automated sewing robots using specialized smartphone apps. Over the past 5 years, the top sewing machine programmers have gone on to advanced operating system hacking; before, they would get bored with programming and get married.

The culmination of this program of training and development is the digital book totebag. I got a demo of this widget in a private suite at one of the conference hotels, but was not permitted to photograph it. The prototype looks nothing like a canvas totebag of course- it's more a mess of wires and connectors. The functionality is quite impressive, however. I was easily able to download an eBook from a Kindle to the "totebag" using a red suction-cup connector that came with some sort of special grease. I then attached an iPad using a USB connector and viewed the book in iBooks. I was also able to connect the totebag to my Google Books account and use the Kindle book there. Yi had a number of other devices to try; each of them had its own quirks, but more or less worked.

I asked Yi how this seeming magic had been accomplished; the most I could get out of him was that any book is "just another sewing pattern". I also asked him if standards for content and DRM would make ebooks portability possible without his digital totebag widget. We had a good long laugh at that one.
Enhanced by Zemanta

Article any source