Showing posts with label EPUB. Show all posts
Showing posts with label EPUB. Show all posts

Sunday, May 19, 2013

Publishing Hackathon Pretty Much Ignores eBooks

The "First Annual" Publishing Hackathon was this weekend. As advertised, I participated and worked on an EPUB backmatter project. My awesome team consisted of me, Javascript/Ruby developer Max Jacobson (who's going to be even more highly sought-after when he finishes Rails school this summer), and TLC librarian Dianne Coan.

Here's our demo video:

 

Here's how we described the project:

Book Discovery INSIDE the eBook

When is a reader most receptive to reading suggestions? Right when they’ve finished a book of course! That’s why printed books have information about other books by the same author, the first chapter of the next book in the series and similar material at the end as part of the back matter.

Back matter has existed pretty much as long as books have. This includes the appendix, glossary, index, and bibliography. Back matter for digital books needs to be optimized to serve the needs of the digital reader. An informal survey by @suw indicates the most popular endmatter desires were other books by the same author and some information about the author.

Digital back matter for ebooks is not constrained by having to proceed the publication; unlike print, digital back matter can be kept up to date with the release of new content. For instance, if an author publishes a sequel, that title could be included in previously published ebooks.

It’s easy to insert a page listing an author’s other books at the end of an ebook, but how do you keep that list up-to-date? What if you’ve developed a great recommendation system to do “if you liked Pride and Prejudice, you’ll like X”? (or maybe “if you hated...”!)

The answer is to make use of the javascript capability of emerging ebook environments. Our project explores means of connecting to APIs from within an EPUB for the purpose of suggesting the user’s next read.

An existence proof is the “widget” capability of the iBooks iAuthor platform. It allows the insertion of html snippets into extended EPUB. Unfortunately, the javascript capability of ebook reading platforms, like the future, is unevenly distributed.

For this demo, we tested three reading EPUB environments, Readium, Readmill, and iBooks. We modified the Project Gutenberg EPUB version of Pride and Prejudice to include hooks and data to other books by Jane Austen.

Readium, which has been built as an EPUB3 reference environment, is the most capable for our purposes. It supports both javascript and connections to external web resources. In Readium, our EPUB displays the set of books by Jane Austen returned by the ReadMill API.

Apple iBooks has full javascript capability, but doesn’t allow connections to external resources (except perhaps via iBooks Author hooks- this deserves further investigation.) In iBooks, our EPUB displays a result page that we generated and embedded based on Jane Austen works published in 1813, when Pride and Prejudice released. We imagine that such embedded resources could be inserted at download time in a future production bookstore or library environment.

The Readmill environment does not support javascript at all at this time, so ironically, we’re not able to display the Readmill API results, or the iframe embedded resource.

Offline reading in Readium displays the resource embedded in the EPUB, similar to the iBooks version.
There were 30 projects in total presented at the end. Here's the list, along with my one sentence summary.
Banned Books in America
Website that maps book banning incidents and links them to Openlibrary
Book Discoverability: A Graphical Solution
Concept for browsing books as nodes on a graph.
Book Discovery INSIDE the eBook
This was us! Our demo crashed and burned. The popup screens from the wifi messed up the ebook reader display of embedded dynamic content.
BookCity Finalist!
Website that recommends books by connecting them to cities.
BookieGoer
Website that helps you lend the books you've borrowed from the library.
Booklvrs: Read. Discover. Meet.
App that advertises the ebook you're reading to the people around you.
bookmatchup
Website that multi-factor-matches you to books.
BookMob
Website that aggregates book recommendations from your twitter followers.
bookshelf.me
Website that displays books as if they were on a bookshelf. I'm pretty sure there was more to it.
Publy.io
Website that recommends books to users based on books they've liked.
Captiv Finalist!
App and Website that uses machine learning algorithms and your tweet about last night's party to combat the short attention span of Today's Readers. I may not have understood this one.
Coverlist Finalist!
Website that believes in judging books by their cover.
Evoke Finalist and clear judging favorite!
Pinteresty website that recommends books based on emotions categorization.
Happy Chapter
App that recommends books based on tags you click.
I read your Brain
Brain-sensing rabbit ears that wiggle depending on your response to a book from a website.
IGNITE
Website that lets users rate romance novels for steaminess.
KooBrowser Finalist!
Browser plugin that analyses what you read to better sell you books.
Library Atlas Finalist!
Mobile app that sends you geographically appropriate quotes depending on where you are. My favorite.
Literary Trinket with Book Wish
3D printed QR-ish code baubles. Cooler than it sounds.
Meadows
Website that turns reading into a game where you earn points.
Meme a book
Website that turns books into lolcats. (I may not have described this accurately.)
MovieReader
Website that recommends books connected to the movie you just saw.
NYPL Reinvent
Analysis of NYPL metadata advocating a divorce of the library from its classification system.
OkLetsRead!
Website offering crowd-funded serial fiction (ebooks).
Quiply
Website that recommends books based on a user's video viewing.
Reading Tollbooth: A Gateway to Book Discovery
Website to match kids to books.
Something2Read
Website that recommends books based on tags you click.
Valerie's Baby App
App that promotes literacy to a girl named Valerie by making sliding block puzzles and defining words at her.
Visibrary
Website that uses library data to make graphical book circles.
Vookstore
Website that turns ex-bookstore owners into book curation engines.
Interestingly, only 3 of the 30 projects addressed ebooks at all, which seems a bit odd to me, considering the industry's ongoing transition from print to digital. The emphasis on apps (7) and websites (21) is partly due to Hackathon's theme of book discovery, but it also says something about the tech industry. Apps and websites are what the NY tech industry is doing in 2013, not ebooks. Clearly, the publishing community developing ebooks and ebook standards needs to do more outreach to developers; the hackathon was a good first step.

It's also worth noting the growing importance of geo-tagging and other non-traditional metadata. In the new world of publishing discovery, readers want books that fit their mode right where they want to be. Neither MARC nor ONIX know enough to help.

My library friends should rest assured that the hackers did not at all ignore libraries. Although $1000 prize from NYPL was a factor, the ease of connecting to NYPL and OpenLibrary helped a lot. The RDA prize, it should be noted, went unclaimed.

Update: Sorry, Coverlist, I omitted your finalist status. Corrected!
Enhanced by Zemanta

Article any source

Wednesday, February 6, 2013

Calligra 2.6 e Krita 2.6 rilasciati con nuove funzionalità


Il team di Calligra ha quest'oggi rilasciato Calligra Suite 2.6 assieme a Krita 2.6

Scopriamo insieme le novità

Anche in questo caso grandi novità. La prima è l'arrivo di Calligra Author, il nuovo programma della suite indirizzato sia ai novelli scrittori che a quelli navigati. 

Words, il programma di elaborazione testi, beneficia ora di un layout migliorato e di una serie di migliorie e correzioni di bug minori come nel caso del controllo ortografico. Oltre a questo eredita da Author alcune funzionalità come le statistiche sul testo migliorate e l'esportazione nei formati ePub.

Sheets, l'applicazione per i fogli di calcolo, porta con se la nuova funzionalità Solver.

Stage, l'applicazione per le presentazioni, ha un nuovo framework per le animazioni che consente agli utenti di creare e modificare le animazioni nelle slide.

Flow, il programma per la creazione dei diagrammi, porta miglioramenti nella gestione delle connessioni.

Plan, l'applicazione per il project management, porta miglioramenti nella gestione del progetto, nei grafici e tante altre piccolo correzioni.

Kexi, il gestore di database, ha un nuovo supporto per la memorizzazione dei dati utenti, migliora l'importazione e l'esportazione degli CSV. Una lista completa delle novità di Kexi la trovate qui.

Sotto il profilo formati supportati Calligra può ora esportare i documenti nel formato ePub2. A partire dalla versione 2.6.1 aggiungerà inoltre il supporto al formato MOBI. Naturalmente è stato migliorato altresì il supporto ai formati proprietari di MS Office, in particolare con i formati aperti XML di Office 2007.

Krita 2.6. Questa nuova versione del noto programma di editing immagini può ora beneficiare del supporto nativo del sistema di gestione colori OpenColorIO  che è uno standard usato nella creazione di film e di VFX studio. Questa mossa, a detta degli autori di Krita, rende il programma una scelta naturale per tutti coloro che lavorano nel campo della pittura digitale 2D per i film e le vfx pipeline.
L'altra grande novità di Krita 2.6 è il pieno supporto al formato PSD: questo significa che ora Krita può non solo aprire i file ma anche scrivere nel formato PSD di Photoshop.


Miglioramenti infine anche per Calligra Active, la versione di Calligra destinata a tablet e smartphone.

Il codice sorgente di Calligra Suite 2.6 è disponibile per il download a questo indirizzo: calligra-2.6.0.tar.bz2

Sul post di presentazione presente sul sito di Calligra troverete le istruzioni su come installare fin da subito Calligra 2.6 nelle distribuzioni più note.
Any source

Tuesday, June 19, 2012

"Open Access eBooks" eBook is on GitHub


When I try to explain to book industry people why ebooks can and should be free, I often get a look that says "What planet are you from?" In contrast, many of my software developer friends take it as dogma that ebooks should not only be free, but also "Free". And so I seem to spend a lot of time explaining one point of view to the other.

What we're trying to do with unglue.it is to skip over the theory, and just show everyone that it works.

First, an explanation for the 99% of real people who haven't encountered the Free vs. free distinction. An ebook that's Free means more than just not having to pay for the ebook, it means that the ebook is not locked up in any way. You can do things with it without needing permission. Copy it, distribute it, convert it, print it, slice it and dice it. Extract it, analyze it, translate it compute it, archive it. In the software world, that's the essence of Free Open Source Software (FOSS). But the 99% just wants to read the book. So why should it bother with Free?

The book that is on the brink of having a successful ungluing campaign at Unglue.it, Oral Literature in Africa, has a Free license proposed for it, CC BY (Creative Commons Attribution). We can't be certain how much the Free license has been responsible for the success of the campaign (you HAVE pledged, haven't you?), but it certainly adds to the appeal. A successful conclusion to the campaign will do more than just let people read the book. It will allow scholars of African culture to add to the book, to use chapters as course material, to use large excerpts in their own work, to make corrections and translations. And the media handling capabilities of new ebook formats will allow the addition of audio to a work about material that deserves to be audible.

What frustrates me, though, is how difficult it is to actually do all the things that you would want to do with a not-locked-up ebook, even the things that don't require it to be Free. Something as simple as correcting a typo is hard for 99.9% of the public. It shouldn't be that way. There should be tools that make this easy. If I want to add my voice into Oral Literature in Africa, there should be an application that allows me to click and speak.

The software world has developed a wealth of tools that allow distributed teams of developers to work together on free software. Source control systems help to track and manage changes in software. We need the same sort of tools that work for books. Wikis do part of the job, but we need more.

So as a first step, I'm putting the short book I've written using this blog, Open Access eBooks, on GitHub, the service we use to track and manage the software behind Unglue.it. All the book's source code is there, mistakes and all. Its CC BY license allows you to take it, branch it, fix it, translate or modify it, redesign and recode it, whatever. You can send me a pull request if you want to merge your changes with my branch. Maybe you want to update the references or add a chapter. Maybe you want to embed metadata or improve accessibility. Maybe you want to fuse it with Moby Dick for some bizarre art project. Whatever. The future of books is all of ours to create.

Notes


  1. Other factors contributing to the imminent success of the Oral Literature in Africa Campaign have been its academic nature, its modest ungluing fee, and its inherent coolness. What, you haven't contributed yet?
  2. Among the Creative Commons Licenses usable at Unglue.it, CC BY and CC BY-SA are considered by Free Culture advocates to be "Free". The Public Domain Dedication (CC0) is not a license, but is another way to make a work "Free". The SA (Share Alike) restriction is a form of "copyleft" which requires derivative works to be similarly made available.
  3. Other CC licenses may add conditions including NC (Non-Commercial) and ND (No Derivatives).
  4. Wikipedia is a good example of a site that won't allow posting of NC or ND licensed content.
  5. The license used for an unglue.it campaign is specified by the rightsholder who may be constrained by  publishing contracts and byzantine international licensing regimes.
  6. I was disappointed by the lack of good tools to create ebooks. I did everything by hand and was surprised at the mess of shifting standards, conflicting ereader implementations and insular documentation.
  7. Whenever I encounter a roadblock in python or django, Google sends me to StackOverflow for the answer. With ebook production, I always end up at MobileRead, ThreePress or Liz Castro's blog. These are wonderful resources, but they're not StackOverflow.
  8. Despite my struggles with EPUB, Amazon's MOBI tools were painless. I felt so naughty!
  9. Because Git is line oriented, I put every sentence in the content file on its own line. Hope that makes sense!
  10. For an example of another ebook with source on GitHub, check out Structure and Interpretation of Computer Programs, Second Edition (SICP). It's not Free, though.
  11. If you want a really nice "free" dinner next Saturday in Anaheim California, make a $100 unglue.it pledge and ask me for an invite. Space is limited!


Enhanced by Zemanta

Article any source

Wednesday, June 22, 2011

EPUB 3 Beefs Up Metadata, but Omits Semantic Enrichment

Ironic amusement fills me when I hear book industry people say things like "metadata has become cool", or "context is everything". Welcome to the 20th century and all that. Meanwhile, in the library industry, metadata has been cool long enough to coat everything with a thick rind of freezer burn.

There's good news and notsogood news for ebook metadata. The revision to the EPUB standard, published just a month ago, includes metadata tools that could eventually lead to a new era of metadata cooperation between publishers and the entire book supply chain, including libraries. At the same time, the revision fails to take advantage of ready-made vehicles for semantic enrichment of content, a move that could still provide new types of revenue for publishers while giving libraries new opportunities to remain relevant as books become digital.

Since I'm incurably optimistic, I'll start with the half-full glass: Publication-level metadata. EPUB 3 includes a whole bunch of ways to include publication-level metadata in an EPUB container. As an example, imagine an EPUB3 for "Emma" with this mark-up in its package document (essentially the navigation directory for the book):
<metadata>
...
<meta property="dcterms:identifier"
id="pub-id">urn:uuid:A1B0D67E-2E81-4DF5-9E67-A64CBE366809</meta>
<link rel="marc21xml-record" href="http://www.archive.org/download/cihm_29722/cihm_29722_marc.xml" />
<link rel="marc21xml-record"
href="/cihm_29722_marc.xml" />
<link rel="foaf:homepage" href="http://openlibrary.org/books/OL24234129M/Emma" />
...
</metadata>

In this example, the first link element points to a MARC 21 xml record (MARC 21 is a blattarian standard for library metadata (look it up)) at the Internet Archive. The second link element points to the same record included in the EPUB container itself. There is also built-in vocabulary that allows the link element to point to ONIX, MODS, and XMP metadata records.

The example also shows that other vocabularies (such as FOAF) can be added for use in metadata elements. So, if you're a believer in RDA, you can put that in an EPUB file as well.

The meta element can also be used in the EPUB package document's metadata block. It's defined quite differently from HTML5's empty meta element, with an about attribute and allowed text content. In principle, it can be used to encode arbitrary RDF triples, thanks to a prefix extension mechanism borrowed from RDFa which allows EPUB authors to add vocabularies to their documents.

These capabilities, on their own, could support major changes in the way that books are produced, delivered and accessed. In a publisher workflow, the EPUB file could serve as the carrier for all the components and versions of a book, even bits that today might be left out or lost in the caverns of so-called "content management systems". A distributor would no longer need to match up content files with records in a separate metadata feed. EPUB books for libraries could be preloaded with cataloging and enrichment data, greatly simplifying the process of making the ebooks accessible in libraries.

Given the great advances for "package-level" metadata, it's a bit disappointing that semantic mark-up of content documents missed the EPUB 3 boat. The story is a bit complicated, and it's far from over. Imagine that you want to add mark-up to a book's citations- perhaps you want to embed identifiers to support library linking systems. Or perhaps you're a medical publisher and you want to embed machine readable statements about drugs and diseases in a pharmaceutical textbook. Or perhaps you want to publish a travel guide and you want search engines to pick out the places you're describing. These applications are not really supported by the current version of EPUB 3.

EPUB content documents have a feature that you might think would do the trick, but doesn't really. The epub:type attribute supports "semantic inflection" of elements. This attribute can be used to mark a paragraph as a bibliographic citation, for example, and supports many of the requirements imposed by conversion of content from legacy or specialized formats into the HTML5 dialect used by EPUB. It's an important feature, but not enough to support semantic enrichment.

Part of the problem is EPUB 3's dependence on HTML5, which is not yet a stable spec and is enmeshed in some surprisingly raw W3C politics. W3C has been the home of HTML standards development since the very early stages of the web, and has also been the home of semantic web standards development. HTML5 started outside of W3C in the WHATWG, an initiative to develop HTML in a way that would be backwards compatible with good-old fashioned non-XML HTML. W3C was convinced to fold WHATWG into its development efforts because of WHATWG's corporate backing. Even so, the WHATWG version of the HTML5 spec drips with sarcasm towards W3C HTML Working Group decisions.

During part of the development of EPUB 3, the HTML5 draft included "Microdata", a method of embedding semantic mark-up in HTML. RDFa, a standard that competes with Microdata, was developed by W3C channels, and within W3C, it was decided in February of 2010 to move Microdata out of the HTML spec so as to give it equal footing with RDFa. Some participants in the EPUB working group wanted to include RDFa in the standard; others thought this would impose too much of a complexity burden on publisher-implementers. The EPUB draft ended up being released without either RDFa or Microdata.

The recent endorsement of Microdata by the Google-Yahoo-Bing cooperation has changed the competitive landscape for embedded semantics. It's now apparent that Microdata will get priority implementation in HTML development tools, leaving RDFa as a niche technology. For most use cases of EPUB semantic markup, the differences between RDFa and Microdata are small compared to the advantages of piggybacking on the technology investment supporting website creation.

According to members of the EPUB working group, it is expected that a dot release will follow relatively quickly behind EPUB 3.0. It seems to me that picking a semantic markup technology for content documents should now not be so hard. If you work for a publishing company that has ever mentioned semantic markup in a product plan, you should probably be making sure that the EPUB working group is aware of your needs. If you are a librarian who can imagine the possibilities of a semantically enriched EPUB collection, you should similarly be making your concerns known.

Although the EPUB working group includes representatives from tools vendors that might conceivably benefit from the adoption of EPUB-only constructs, the group's track record for adopting wider web standards has been very encouraging. By adopting HTML5 as a stack component, the group has ensured that cheap or free tools to produce and author EPUB 3 content will be readily available.

Once semantic enrichment of ebooks becomes routine, libraries will play a vital role in their use. Libraries provide a copyright-friendly DRM-free community commons in which users can access and build on the information contained in licensed content. (Of course, I see "unglued" books as playing an equally important role in the library commons.)

The EPUB metadata glass is half full, and there's more wine in the bottle!

Note: This is one thing I'll be talking about on Saturday at the American Library Association meeting in New Orleans. (The program is somewhat inaccurate; the program will end at 10:30 AM at the latest. Ross Singer from Talis will lead off with an overview of semantic web technologies in libraries; I'll follow with discussions of RDFa, the Facebook "Like" button and of course, EPUB.
Enhanced by Zemanta

Article any source

Thursday, June 2, 2011

EPUB Really IS a Container

"It's OK for libraries to put things in their EPUB books." That's what Bill Kasdorf, a member of the EPUB Working Group, told me last week at the IDPF Digital Book 2011 Meeting. He checked with EPUB Revision Co-Editor Markus Gylling to make sure. I had been curious if libraries could put all their cataloging information inside an EPUB file instead of siloing it in their catalog system.

It may seem an odd question if you don't know a few things about EPUB. EPUB is a standard format for ebooks. It's used by Apple, Barnes and Noble, Kobo, Overdrive and many others not named Amazon. EPUB is near the end of a revision process that will result in EPUB 3.0.

The EPUB specs define a lot more than just a file format. Both EPUB 2 and EPUB 3 define a container format (in EPUB 3 it's called the EPUB Open Container Format (OCF) 3.0, and then go on to define a number of file formats for files that go inside this container. These files are the resources- texts, graphics, etc. that make up the ebook.

OCF uses the ubiquitous ZIP format to wrap up all a book's resource files into a neat, transportable package. That's pretty much standard these days. Java ".jar" and ".war" files use the same mechanism, as do MacOS' ".app" files.  As a consequence, you can use any unzip utility to look inside an EPUB file and manipulate its contents.

There's even a reserved name for a file to contain book level metadata in OCF: META-INF/metadata.xml, as well as another file for rights information, META-INF/rights.xml. Another file, META-INF/signatures.xml can be used to prove who made parts of the file and determine whether anyone has mucked with them. When Gluejar issues Creative Commons editions of newly relicensed works, we'll use the rights.xml file to make sure the CC declaration is explicit.

The new EPUB revision is coming fast. Last Monday, Bill McCoy, Executive Director of the International Digital Publishing Forum (IDPF) announced the release of the full EPUB 3 proposed specification. My guess is that when we look back on this event 10 years hence, we'll recognize this as the moment EPUB began to revolutionize the world of information, and with it, the book industry.

Although Amazon still uses the aging MOBI format on its kindle devices, it seems only a matter of time before the infrastructure accumulating behind EPUB pushes them into the embrace of the IDPF. Already, most of the content flowing into the Amazon system is being produced in EPUB and converted to MOBI. Don't expect this shift to happen soon though; in his IDPF presentation, Joshua Tallent of eBook Architects described rumors that this would happen soon as "bunk"- but it will happen sometime.

EPUB 3 comes with lots of goodies. The revision adds several modules of sorely needed capability. It includes MathML, SVG and JavaScript over a substrate of HTML5 and CSS2.1. While MathML and SVG are essential for education and technical markets, JavaScript has been somewhat controversial because of the difficulty of making sure things work securely and without connections. Most of the reading systems inherit javascript capability from the WebKit rendering engine they're based on, so a lot of javascript functionality will work in ebook readers regardless.

(left) Autography Founder and Author  T. J. Waters
All this capability will remain latent unless people find compelling uses for it. I'm not worried. As the BookExpo itself got started, I met two different companies who were manipulating ebook files to solve the same problem: how can an author sign a book when the book is digital? Both companies, Autography and InScribed Media, create personalized experiences that leave artifacts of an author-consumer interaction inside ebook container files. Both of these companies have compelling solutions; they differ in their business models. Autography is structured as an author focused bookstore; InScribed is developing partnerships with existing bookstores.

InScribed Media Founder and Author Alivia Tagliaferri
To some extent, InScribed and Autography are forced to be a bit convoluted in the way they deliver their product because they need to live inside DRM green zones; users don't have access to the files inside books without cracking the DRM (which is rather easy, by the way!). It's unfortunate, because personalization of ebooks could be a good way to encourage responsible use. I certainly don't want that picture of me torrenting around the world!

Libraries face a similar dilemma. The insides of an EPUB file could be greatly enriched by  libraries, which have every motivation to enhance discovery both of the book and the information inside of it. But DRM gives the publisher and its delivery agents the exclusive ability to build context inside ebook containers. Libraries and readers are locked out. I think that for DRM systems to survive they will need to accommodate a more diverse set of user manipulations; author signatures are just the tip of the iceberg.

Coming soon, I'll report on EPUB 3 metadata.
Enhanced by Zemanta

Article any source

Wednesday, May 18, 2011

The Object-Oriented Book

To most people, objects are things you can touch, see, maybe even smell. They have existence on their own. Software developers talk about objects as well. Although they're more abstract, software objects can also be touched- programs can interact with them, and they exist on their own as packages of code and data.

In some recent conversations about books and content containers, I've been hit in the face with the fact that most people in publishing haven't been steeped in Object-Oriented Programming (OOP) the way I once was, and as a result, some of the things I've written about the evolution of the book into digital form have sounded a bit strange to many people. So I've decided to write a bit here about how books are becoming software objects, and why it matters.

Object orientation is a style of programming that models problems as spaces of objects from various classes. The programmer solves problems by manipulating objects; the objects communicate among themselves by passing messages. The messages that objects pass are governed by interfaces; every class of objects is defined by the interfaces it supports. If that doesn't make sense to you, don't worry, I'll have some examples.

Let's think about how we might model the book as a software object. With a physical book, you know how to get the title and name of the author. You open up the book to the title page, and there you find the title, probably the words in the largest type size, and the author's name, probably printed below the title, perhaps with a designator word such as "by".

In the prehistory of programming before OOP, a book program might define data structures containing tables of book titles and author names. The program would look in these tables for the book data. An object-oriented program would instead send the book-object messages saying "what is your name?" and "What person was your author?" An object-oriented approach binds the code and data together, so that objects of the book class know what their title is, how many chapters they have, and what the 20th word of the 32nd paragraph of their 3rd chapter is. The set of messages that an object can respond to defines its class. A programmer knows that any object in the Book class will be able to tell you its title.

Another key concept in object orientation is inheritance. A cookbook is a book and inherits from the Book class the ability to tell you its title. But you expect a Cookbook to have recipes, and you should be able to ask it how many recipes it contains.

The reason I think this is important for non-coders to understand is that very soon, the book industry will become focused on producing lots and lots of these software objects. And I'm not talking about some far-fetched digital utopia.

The third revision of the EPUB standard is very soon to become a reality, and I believe its use will quickly become pervasive in the book industry. It would be a mistake to think of EPUB3 as yet another document format. With the adoption of EPUB3, the book industry will, for the first time ever, have standardized a software object model for the book. This comes along with EPUB3's use of HTML5 as a foundational layer.

An object model became associated with HTML documents very early in its evolution. Called the DOM, or Document Object Model, it was developed by programmers working with HTML documents, and it quickly became the basis for most software that works with HTML documents. With the development of Javascript, HTML documents delivered over the web could bind to code that accesses and manipulates their data via the DOM. It's only with HTML5, however, that the DOM is officially becoming part of the HTML standard.

With HTML5 as its basis, EPUB3 becomes a very capable "container" of content. The whole discussion of how containers limit the ways in which content can interact with consumers becomes completely moot, and a bit silly. EPUB3 binds a complete "API" (application programming interface) onto the content, and provide many mechanisms for the extension of that interface. The "API" and the "container" are one and the same.

If we look at the immense infrastructure that arose around the book as a physical object, from book bags and compact shelving, to printing plants, warehouses, libraries and used bookstores, we can get an inkling of the infrastructure that will grow up around the book as a software object. In the coming weeks, I'll try to write about some of the implications of EPUB3 for the industry as a whole.
Enhanced by Zemanta

Article any source

Tuesday, February 15, 2011

How Apple May Inadvertently Boost eBook Linking

"The net interprets censorship as damage and routes around it." John Gilmore, 1993
The official word from Apple finally came out today, in their press release announcing in-app subscriptions.
In addition, publishers may no longer provide links in their apps (to a web site, for example) which allow the customer to purchase content or subscriptions outside of the app.
Assuming that this limitation will be applied to the Kindle App, it means that the "Shop in Kindle Store" button will disappear, and similar features in other ebook reader software, such as Nook, Sony, and Kobo will disappear as well.

You can go elsewhere if you want to read apocalyptic whining about Apple's imperious ways. What I want to focus on is how the net will route around this damage. The net will route around this damage by making more links.

At O'Reilly's Tools of Change for Publishing Conference, I was able to spend some time with Keith Fahlgren, a partner at ThreePress Consulting. He's part of a group that has worked on the improvement of linking capability in EPUB 3. A Public Draft of the specification was released today by IDPF.

Since EPUB3 is based on HTML5, all the outbound linking that you would expect from a web page is already built into EPUB3 (as well as earlier versions of EPUB). Ebook reader apps available on iOS and Android use the "Webkit" webpage renderer for ebooks in EPUB. (Kindle devices use Webkit to render web pages and WebKit is used by Amazon to render Kindle ebooks (in mobi format) on  hardware other than their own.) So it's clear to me, at least, that even if ebook reader apps can't have "Kindle Store" buttons, the apps will be able to present "Kindle Store" links inside the ebook content. I'll bet you anything that Amazon is loading up ebook content with Kindle Store links: "If you like this book, perhaps you'd like this one". They'll even have specialized shop-books containing Kindle store links available for free. Ditto the others.

Publishers aren't going to like Apple's power-play. But neither will they like having their content getting hijacked to promote individual ebook stores. There will therefore be a great deal of pressure for the creation of vendor-neutral, customer friendly ways to link to ebooks from within ebooks, one that Apple can't ban because doing so would break Safari.

Here's where it gets tricky. If a customer has already purchased the linked-to book, it's pointless to send them out to a ebook store, they should connect their copy of the ebook. But figuring out whether a consumer already has the book is messy, given the state of ebook identification. There are many other use cases for linking to a specific chapter or paragraph inside an ebook.

Unfortunately, doing this sort of linking is not a solved problem. EPUB3 adds one tool that will help. A new required metadata property, dcterms:modified,  will help identify the epub in the case where it has been modified- in the past it was poorly specified what should happen to the epub identifier if the file was modified. With EPUB3, it's now clear that EPUB documents are identified internally at a level above the ISBN (different DRM wrappings of the same EPUB file often require different ISBNs) but below the "work".

There's still a lot of apparatus that will need to be built, both inside and outside of EPUB, for linking to work the way it should. Being able to decide which ebook to target will require external mechanisms. Perhaps some linking organization along the lines of Crossref be formed; perhaps a more wikipedia-ish database collaboration will suffice. In any case something like xISBN supercharged for ebooks will be needed. Fahlgren told me that without a strong use case to drive the solution, the EPUB group has had a hard time going very far in their linking development.

A true ebook linking solution would need to include Amazon, of course, and since they've not been using EPUB, it seems to me that ebook linking won't get done by the EPUB group itself. Amazon hasn't had much use for EPUB in the past, but now Apple may have handed the ebook technology community a giant use case for interoperable ebook linking.

Happy Day-After-Valentines-Day, EPUB!

Article any source