Showing posts with label magic. Show all posts
Showing posts with label magic. Show all posts

Tuesday, June 25, 2013

Magic Rights Management for eBooks

The Fraunhaufer Institute in Germany is apparently marketing some "new" technology they're calling SiDiM which embeds digital data into texts by changing word in the text. They're telling publishers that it will fight piracy by making it easy to track files uploaded to torrents and file lockers back to their reprehensible sources.

It sounds kinda dumb, doesn't it?

The internet has over-reacted of course, perhaps everyone is hypersensitive because of the revelations about the NSA and its data collection practices. Nick Harkaway jumped the shark a bit and called it "surveillance". The idea of changing words in books is easy to ridicule and deserves to die, but let's please take a deep breath.

There are lots of ways to put information in an ebook file, and license information is no different. For example, I've been advocating that Creative Commons licensed books should embed a digitally signed license so that the license can be relied upon. When you buy an ebook, an embedded license could protect you from accusations of infringement. Digital signatures can also tell you that a books content hasn't been tampered with.

When you buy a Harry Potter ebook direct from Pottermore, your identifying information gets digitally stamped into the ebook. According to the The Digital Reader, Pottermore uses watermarking technology from Booxtream, and I've been evaluating this technology myself for an Unglue.it project. So far, I'm impressed.

The rationale behind Pottermore's watermarking is that it prevents people from sharing the book beyond what their license allows. If the book gets on a public filesharing site, it can be traced back to the purchaser, and consequences could ensue.

Booxtream claims to be using 9 different watermarking techniques to make the embedded data hard to remove. For example, Booxtream adds digital codes to the names of the content files inside the EPUB, and adds data into image files. Although it's straightforward to strip some of the embedded info, Booxtream needs only to make it uncertain that stripping has been complete to retain some deterrence value.

For the user, the bottom line is that nothing the purchaser does or would want to do is impeded by the Booxtream watermarking. Nothing visible to the user is altered except for an ex libris page that tells the user that the copy has been personally licensed to him or her- it's customizable by the vendor.

A close analogy to ebook watermarking is the bullet serialization that's been proposed as an alternative to gun control. If every bullet was traceable to a purchaser, investigation of weapons related crime would be reduced to finding the bullet and looking it up in a database. Law abiding gun owners shouldn't notice the difference. Or maybe it would be de facto ban on ammunition. YMMV.

The argument against digital watermarking is that there will always be ways to remove the embedded data, no matter how clever you are at hiding it. Someone will make a one-click watermark stripper, and the value in watermerking will be diluted. But almost two years after Pottermore launched their digitally watermarked ebooks, it's quite hard to find watermark stripping tools. But why would anyone bother? There's nothing that the vast majority of ebook purchasers want to do that's impeded by the watermarking. Contrast that with the ease of finding tools to strip PDF watermarks, which are annoying.

You might wonder why SiDiM would be selling their technology with such a scary-dumb sounding marketing pitch. It's because publishers are the customers. I've been talking to a lot of publishers, and they're very clear that they want DRM. Or at least they THINK they want DRM. What they really want is magic. The want their ebooks to come with a magic bullet that stops piracy and over-sharing dead in its tracks. They don't understand the technology behind DRM, but many of them swallow the story that comes with it- that nobody would pay for digital files if they can get them for free from piracy sites.

The truth is that if there's magic in the kind of DRM that comes with Adobe, Apple and Kindle, it's of the variety that Voldemort would use. If there's magic in the watermarking techniques used by Pottermore, it's of the Dumbledore variety. If there's magic in SiDiM, it's like Neville Longbottom's Switching Spell that put ears on a cactus.

I'm here to tell you that magic is real. There's real magic in the stories that authors tell. There's real magic in communities and in relationships between people, between authors and readers. There's real magic in libraries. It's that real magic that will stop piracy and help authors earn a good living in the digital future.

Dumbledore's fictional magic can help make the real magic manifest, and thats what we should work towards
Enhanced by Zemanta

Article any source

Tuesday, October 5, 2010

Aggregating Deep Discount Readers of eBooks

The book publishing industry should be terrified of readers like me. Over the last year, I have purchased a grand total of one new book. Why only one? I have a huge stack of books, both print and digital, in my aspirational reading queue. I read plenty of books, but there are many more books that I would like to read; so many in fact that I see no reason to spend $30 on a book when there are plenty of minimal-outlay books already waiting to fill my hours.

The books on this stack come from a variety of sources. I read ebooks from my wife's Kindle account. The print books are almost all purchased at used book sales, typically for $1 or $2 each. The publisher's revenue from all of this book enjoyment is $0, and of course the author gets only a small fraction of that in royalties.

No matter how terrifying the idea is, discount readers like me represent a big opportunity for book publishers as they move there properties onto digital platforms. Discount readers come in many forms; it's safe to say that the billions of books lent by libraries went to people unwilling to pay full retail price for books. Libraries contribute modestly to the income streams of publishers and authors; used book sellers not at all. In their print businesses, publishers have learned to segment their markets by offering paperback versions as well as remaindered books, but they have largely neglected the deep discount end of the demand curve.

There's a huge amount of value to society in deep discount demand, and it not just in the benefit to readers like me. Libraries include the preservation of our written culture in their mission; this activity wouldn't happen if the first-sale and fair use doctrines didn't limit the control that publishers and authors could exert over the use of their works. We need to think about how to do preservation as we translate the book business into a digital industry.

I've been thinking a lot about business models for ebook publishing, and a lot of my thinking has surrounded market segmentation methods. I've been looking for ways ways that discount readers like me can be aggregated into sustainable revenue streams to sustain institutions such as libraries.

One obvious model to serve the discount reader is to offer subscription packages. I've come to the conclusion that ebook subscription packages have many structural problems. Subscription packages inevitably cannibalize sales of the items they contain, and there's a lot of incentive for the package to exclude items that readers would really want.

In the course of studying academic publishing models, I think I've found a  way for the book business to serve deep discount readers, to reinvigorate libraries, and to create a new, sustainable revenue stream for publishing: public acquisitions of ebook rights.

Here's how it might work for me. There are lots of books I aspire to read, many more than I have time for. There are also lots of books that I'd like to have on my reading devices, because I've read them once in print. I want to use these books in many ways, on many devices, at any time in the future. I want to be able to search them, and have others read them. I don't want to have to mess with DRM. And I want them to be preserved and available forever in public libraries.  I also like the one new book I've purchased (Clay Shirky's Cognitive Surplus: Creativity and Generosity in a Connected Age) enough that I would also pay something towards letting you read it too! Imagine that I could offer $1 for each title on my list to have this magic occur. (Cognitive Surplus is a thought-provoking meditation on the things that happen when the barriers to collective action are lowered, among other things. You should read it!)

OK, here's a stretch, imagine that millions of other people feel the same way!

If millions of people feel this way, there's absolutely no reason this magic can't happen. I have advocated that libraries should work together to collectively acquire ebook assets. The same mechanisms that would allow libraries to act collectively could be used by individuals to act collectively on behalf of books that they care about.

 If a hundred thousand people offered a dollar to Clay Shirky (and Penguin, his publisher) for Cognitive Surplus to be released as a creative commons licensed ebook, certainly at some point they would examine their prospects for future sales and figure out how to say "yes". Once a book is liberated in this way, all the magic just happens.

 I'm not expecting J.K. Rowling to cash in her Harry Potter rights anytime soon, but I think there are many types of works and many types of authors who would find it financially advantageous to monetize their work in this way if it became popular. I also think that it's very common for readers to be passionate about the books they read in ways that transcend their narrow financial self interest.

 If you agree with me that mechanisms for public ebook acquisition by readers should be developed, I would very much like to to hear from you, either privately or in the comments!

Article any source

Friday, June 25, 2010

Introducing the Totebag for eBooks

As we hurtle towards a future where books come on Kindles and iPads and Nooks, we tend to overlook the loss of many products and services attached to the print book ecosystem. Tens of thousands of people whose livelihoods depend on books will suffer tragic dislocations in their lives. While many bemoan the plight of bookstore workers, librarians, editors, and authors, there are other small industry segments no one ever thinks of.

Totebag manufacturing is just one of these overlooked industries. Half of the world's  novelty totebags for books are manufactured in a single town in China called Shu Bao (书包). Shu Bao is located in an inland area of China that has concentrated on book related products; neighboring towns specialize in bookmarks, dust covers and those little alphabet labels used in dictionary manufacture. At this weekend's American Library Association (ALA) meeting in Washington DC, I had a chance to speak with Shu Bau's mayor, Yi Rui-Da, who doubles as a sort of totebag ambassador and salesman to the world. Yi was in town to start getting the word out about digital book totebags.

Yi told me that the central committee of his town has been closely watching the shift to eReading for at least 10 years. They've seen one of the neighboring towns become quite wealthy by shifting their manufacturing to iPad covers, and hope to make a similar transition themselves. The lesson of what happened to buggy-whip manufacturers after the introduction of the Model T is known to the committee. Some committee members thought the town was in the luggage business, and preferred to stay in the luggage business. Other committee members, aware of the specialized fibers that must be added to their totebag fabrics, argued that the town was really in the information portability business; these voices prevailed.

To make the transition to transporting eBooks, the town had to nurture its programming talent, of which it has an abundance. Totebags are made in factories that employ hundreds of teenage girls. But it's not like the old days, when the girl were virtual slaves, sewing everything by hand. In a modern totebag factory, the girls program automated sewing robots using specialized smartphone apps. Over the past 5 years, the top sewing machine programmers have gone on to advanced operating system hacking; before, they would get bored with programming and get married.

The culmination of this program of training and development is the digital book totebag. I got a demo of this widget in a private suite at one of the conference hotels, but was not permitted to photograph it. The prototype looks nothing like a canvas totebag of course- it's more a mess of wires and connectors. The functionality is quite impressive, however. I was easily able to download an eBook from a Kindle to the "totebag" using a red suction-cup connector that came with some sort of special grease. I then attached an iPad using a USB connector and viewed the book in iBooks. I was also able to connect the totebag to my Google Books account and use the Kindle book there. Yi had a number of other devices to try; each of them had its own quirks, but more or less worked.

I asked Yi how this seeming magic had been accomplished; the most I could get out of him was that any book is "just another sewing pattern". I also asked him if standards for content and DRM would make ebooks portability possible without his digital totebag widget. We had a good long laugh at that one.
Enhanced by Zemanta

Article any source

Thursday, June 24, 2010

Inter-Library Loan Reinvented for eBooks and Just-In-Time

My graduate school training was in engineering and in physics. In engineering, you put things together and try to get them to work. In physics, you smash things (the polite term is "perturbation") to help you understand how they had been working. I still use these approaches to help me understand the things I write about. You can learn a lot about a system be noting the bits that squawk when faced with a perturbation of the system.

I got a lot of interesting feedback on my article on patron-driven ebook acquisition. It seems that this perturbation in library processes could have wide ranging effects far outside of libraries and book publishing. Coincidentally, the patron-driven model, along with other changes in the library/publisher ecosystem, was discussed last weekend at a meeting of the American Association of University Publishers (AAUP). Publishers Weekly has a nice report. (See also a report in the Chronicle of Higher Education.

The biggest perturbation being imposed on this system is of course the reduction of library budgets, which has come down quite painfully on university presses and their monograph businesses. Still, speaker Joe Esposito was surprised that the strongest reaction to his talk was to his prediction that libraries would make up a shrinking fraction of the university presses' sales.

It seems that there is worry that a contraction or restructuring of monograph publishing could have repercussions for how scholars obtain tenure in the humanities:
The fact that monograph publishing exists to support tenure and the structure of academic employment is an inconvenient truth that can no longer be glossed by either the Academy its associated University Presses. At some point the Academy is either going to have to stop expecting University Presses to fulfill this need, or find a more honest and transparent way of funding it.
if that's the worst thing that happens, well, what's the big deal?

It won't be a shift to patron-driven acquisition that kills off monograph publishing, however. My reasoning is that from the point of view of economics, patron driven acquisition is roughly isomorphic with the current system of just-in-case purchasing coupled with inter-library loan (ILL).

Here's how things work for print monographs. Suppose a university press published an obscure but brilliant scholarly monograph five years ago. It might have sold 100 copies for $100 apiece, most of them to libraries. At $10,000 gross revenue, it was hard for the press to make much profit, but occasionally they get lucky and make enough to cover the losses on the rest of their catalog. Now here's the problem: Over the five years, there were only about 100 scholars in the entire world that really wanted to read the monograph. Unfortunately, only 50 of them worked at institutions that purchased the book. The libraries of the other 50 didn't purchase the book because the selectors in their libraries weren't omniscient or perfect, and they didn't have mind-reading abilities or the power of divination. Or maybe the libraries used an approval plan that hadn't been crafted with the obscure field of this monograph in mind.

But those 50 others still got to read the monograph, because of inter-library loan. For some libraries, ILL is even a revenue center, because their costs to lend are less than the fees they charge. Although publishers made money from the 50 libraries that bought the book and didn't use it, they don't capture any of the revenue from ILL activity. The libraries that spent money to buy the book right away are partially compensated for that expenditure by ILL revenue or reciprocal loans.

Now let's think about what happens in a future where just-in-time ebook acquisition dominates. The 100 users still get to use the monograph, but none of them need to wait for an ILL transaction to go through. The costs are assigned to the institutions that actually use the work. If we assume that the price of the monograph is unchanged, the publisher's revenue is also unchanged; except it's pushed out to the time of usage, which can be many years, especially in the humanities. The time value of this revenue stream is reduced- it takes longer to make back the money spent on producing the book.

The compensation for the publisher is that the revenue continues for as long as the work is still used. The book doesn't go out of print. In addition, since users can discover the monograph more widely, and obtain it immediately, there is the possibility of making additional sales to users who would never have requested the title via ILL.

In a sense, the patron-driven acquisition model is souped-up ILL, with usage fees accruing to the publisher. The comments of Macmillan's John Sargent earlier this year that publishers would like to see fees for library ebook lending don't seem so controversial when examined under this lens.

It's worth thinking through a publisher's pricing strategy. If libraries persist in their preference to remove price as a factor in the patron's decision to use an ebook, then publishers have no incentive to cut costs and keep prices moderate. If libraries allow automatic purchase of any ebook under $100, then publishers will price all of their products at $99. A similar dynamic in the US health care industry has not worked well for consumers, to say the least. Indeed, one university press publisher writing about patron-driven acquisition and the AAUP meeting has opined that patron selection will lead to higher monograph prices:
What this Patron Driven Access model means to university presses is that our future is likely to include two things—higher prices and fewer titles.

It's clear that there would be winners and losers under a just-in-time acquisition system. Librarians don't always select what their patrons really want to read. Controversial works might do quite well, as should engaging but hard-to-categorize works and works that don't break new ground but are readable and useful. Dry, unreadable, redundant works that sell well today because of the author's fame or because they fit into a "hot" field of research will be losers. A work that today is unread because it's too innovative and ahead of its time will eventually find its time under the just-in-time acquisition.

The huge change for monograph publishers will be in the way they market their products. The emphasis will shift from pre-publication marketing to libraries towards search engine optimization and post-publication marketing directly to users. Famous professors may find themselves awash in free ebooks as monograph publishers jockey for key citations and mentions; social networks and subject specific communities will be prime targets of monograph promotion. Publishers will abandon library convention exhibits like ALA in droves; parties and receptions for librarians will disappear.

I suppose we should have fun with the current system while it lasts, even as there are new and more efficient things to build.
Enhanced by Zemanta

Article any source

Wednesday, March 10, 2010

eBooks in Libraries a Thorny Problem, Says Macmillan CEO

John Sargent, CEO of Macmillan, one of the US's "Big Six" publishers, is not afraid of new business models. Over the past year, Macmillan has been trying to figure out how to push ebook pricing above the $9.99 level that Amazon had set as a standard on the Kindle. They had explored "enhanced" ebooks- ebooks that come with extra content- and were about to implement "windowing" (holding back ebook release to protect hardcover pricing, something that Sargent felt was "completely stupid").

Instead, Sargent decided to take advantage of Apple's announced entrance into the ebook distribution game to force a change of Macmillan's business relationship with Amazon. Instead of using the same discount model for both ebooks and print books, Macmillan wanted Amazon to change to a "agency model" where pricing would be controlled by Macmillan and Amazon would take a percentage. Amazon (which is Macmillan's 2nd largest customer) balked, and stopped selling Macmillan books entirely. But two days later, Amazon gave in. As a result, Sargent has been called publishing's "new hero".

Sargent spoke with the "Publishing Point" Meetup Group today in New York City, and I got to participate in the questioning. Michael Healy, Executive Director Designate of the Book Rights Registry, did a great job of leading the conversation. I was very impressed with Sargent, who dressed in jeans and had a casual, down-to-earth manner that matched. Sargent clearly understands all the challenges his industry faces- disintermediation, shifting distribution, the need to develop technology expertise, but at the same time he's very optimistic about publishing's prospects. He understands the assets at his disposal, in his words, "a lot of extremely good people who know how to obtain manuscripts and who know what people want to read", and who know how to gather enthusiasm around a piece of writing, a process that's "magic".

The most amusing comments by Sargent came in response to Healy's questions about whether the large, generalist, publishing houses would continue to be viable. Sargent seemed to think that in the near term (5-10 years) the big 6 would likely remain intact. (HarperStudio's Robert Miller has predicted the Big 6 could shrink to 3) His reason was not what I expected. The Big 6 are in no danger of implosion- they survived a very hard economic stretch quite well, but no private equity firm or bank would go near them because of "disastrous" balance sheets. They "suck cash, and have terrible profits." "We're disastrous but stable" quipped Sargent.

When my turn came to ask a question, I asked Sargent if he had thought about the role of libraries, and particularly public libraries, in ebook distribution.  His answer indicated that just as he was not afraid of changing the relationship with Amazon, Sargent is not afraid of changing the publisher's relationship with libraries. In fact, change may well be required.

"That is a very thorny problem", said Sargent. In the past, getting a book from libraries has had a tremendous amount of friction. You have to go to the library, maybe the book has been checked out and you have to come back another time. If it's a popular book, maybe it gets lent ten times, there's a lot of wear and tear, and the library will then put in a reorder. With ebooks, you sit on your couch in your living room and go to the library website, see if the library has it, maybe you check libraries in three other states. You get the book, read it, return it and get another, all without paying a thing. "It's like Netflix, but you don't pay for it. How is that a good model for us?"

"If there's a model where the publisher gets a piece of the action every time the book is borrowed, that's an interesting model."

Sargent has clearly thought about libraries, but perhaps he's not talked much to them. His points are valid- the existing business relationship between publishers and libraries won't work for ebooks the way it has worked for print books and the "frictions" that exist for print materials could disappear for ebooks. But he has gaps in his knowledge of libraries. The patron-on-the-couch scenario wouldn't work for libraries either- why would a town support its library's ebook purchasing if everyone could get the ebook from a library 3 states away? The fee-per-circulation model would be a disaster for most libraries, which have fixed annual budgets, and can't just close in September if they've spent their circ budget.

On the other side, the models preferred by libraries are not necessarily going to work for publishers. While the subscription model will probably work for academic institutions, it would turn public libraries into unnecessary intermediaries. The "perpetual access" model would be suicide for publishers if applied to their most profitable top-line books.

Now is the time for publishers and libraries to sit down together and develop new models for working together in the ebook economy. Executives like John Sargent are not afraid of change, but they need to better understand the ways that they can benefit from working with libraries on ebook business models. Libraries need to recognize the need for change and work with publishers to build mutually beneficial business models that don't pretend that ebooks are the same as print.

Enhanced by Zemanta

Article any source

Monday, December 28, 2009

The Case Against Using Spoofed e-Books to Battle Piracy

I've known since I was four years old the difference between the Swedish Santa Claus and the American Santa Claus. The Swedish Santa Claus (the one who comes to our house) uses goats instead of reindeer and enters by the front door instead of the chimney. And instead of milk and cookies, the Swedish Santa Claus (a.k.a. Jul Tomten) always insists on a glass of glögg.

The glögg in our house was particularly good this year (used Cooks Illustrated recipe), so Jul Tomten stayed a bit longer than usual. I had a chance to ask him some questions.

"You're looking pretty relaxed this year, what's up?" I asked.

"It's this internet, you know. What with all the downloaded games, and music and e-books, my sleigh route takes only half the time it used to!"

"Really, that's amazing! I've read about the popularity of Kindle e-books, but I never imagined it might affect you! Are you worried that the sleigh and goat distribution channel will survive?"

"Oh not at all, Eric, remember, Christmas isn't about the goats, it's about the spirit! And even if all the presents could be distributed digitally, someone's got to go and drink the glögg, don't you think?"

"One thing I've been wondering, that list of yours, you know, the naughty and nice list... It must be very different now- do you look at people's Facebook profiles?"

"Ho ho ho ho. At the North Pole, your privacy is important to us, as the saying goes. Well, I'm going to let you in on a little secret. 'Naughty and Nice' is a bit of a misnomer. We never put coal in anyone's stocking. The way we look at it, there's goodness in each and every person."

"I guess I never thought of it that way."

"Just imagine how a child would feel if they woke up Christmas morning to find a lump of coal in their stocking! Even if the child was very naughty, do you a holiday disappointment would suddenly turn the child nice?"

"Besides, if we really wanted to put something useless in a stocking these days, it would be a VCR tape or an encyclopedia volume, not coal."

I've had a chance to reflect a bit on my chat with Santa, particularly about putting coal in naughty people's stockings. I've recently been studying how piracy might effect the emerging e-book market, and I've made suggestions about how to reinforce the practice of paying for e-books. But one respected book industry consultant and visionary, Mike Shatzkin, has made a suggestion that the book industry should take the coal-in-the-stocking approach to pirated e-books.

In an article entitled Fighting piracy: our 3-point program, Shatzkin proposes as point #1:
Flood the sources of pirate ebooks with “frustrating” files. Publishers can use all sorts of sophisticated tricks to find pirated ebooks, like searching for particular strings of words in the text. (You’d be shocked at how few words it takes to uniquely identify a file!) But people looking for a file to read will probably search by title and author. So publishers can find the sources of pirated files most likely to be used by searching the same way, the simple way.

But, then, when publishers find those illicit files, instead of take-down notices, which is the antidote du jour, we’d suggest uploading 10 or 20 or 50 files for every one you find, except each of them should be deficient in a way that will be obvious if you try to read them but not if you just take a quick look. Repeat Chapter One four times before you go directly to Chapter Six. Give us a chapter or two with the words in alphabetical order. Just keep the file size the same as the “real” ebook would be.
Points 2 and 3 of Shatzkin's "program" are reasonably good ideas. But this point 1 is a real clunker.

I'll admit, when I first read Shatzkin's proposal for publishers to put "sludge" on file sharing sites, I thought it an idea worth considering. After having studied the issue, however, I think that acting on the idea would be a foolish and shameful.

First of all, the idea is not original. The tactic of spoofing media files was deployed by the music industry in its battle against the file sharing networks that became popular after the demise of Napster. This tactic was promoted by MediaDefender, a company that also used questionable tactics such as denial of service attacks to shut down suspected pirate sites. Although the tactic was at first a somewhat effective nuisance for file sharers, the file sharing networks developed sophisticated defenses against this sort of attack. They adopted peer-review and reputation-rating systems so that deficient files and disreputable sharers could easily be discriminated. They instituted social peering networks so that untrusted file sharers could be excluded from the network of sharers. The culture of "may the downloader beware" has carried over for e-books. On one site I noted quite a bit of discussion of the true "last word" of Harry Potter and the Deathly Hallows along with chatter about file quality and the like. After seeing all this, Shatztkin's suggested point 1 seems quaint, to put it kindly.

The e-book-coal-in-the-stocking idea could also be dangerous if acted on. The tactic of disguising unwanted matter as attractive content has been widely adopted by attackers going back millennia to the builders of the Trojan Horse. The sludge could be as innocuous as a Amazon "buy-me" link with an embedded affiliate code, or it could be as malicious as a virus that lets a botnet take control of your computer if you open the file. When this really happened happened for video files, it was widely asserted, without any substantiation that the viruses were planted by the film industry operatives themselves. Thus, what began as a modest attempt to harass Napster file sharers ended up resulting in a smeared reputation for the film industry.

Obviously, Shatzkin is not advocating spoofing e-book files with harmful content on file sharing sites. But publishers who are tempted to follow his point #1 should consider the possibility that emitting large amounts of e-book sludge could provide ideal cover for scammers, spammers, phishers, and other cybercriminals. Then they should talk to their lawyers about "attractive nuisances" and "joint and several liability".

Go ahead and accuse me of believing in Santa Claus. I firmly believe that no matter what business you're in, not everybody gets corrupted. You have to have a little faith in people.

Reblog this post [with Zemanta]

Article any source

Thursday, October 15, 2009

Normal and Inverse Network Effects for Linked Data

The human brain has an amazing capacity to recognize familiar patterns in unfamiliar environments. One manifestation of this is pareidolia, the phenomenon of seeing an image of the Virgin Mary in a grilled cheese sandwich or a mesa on Mars that looks like a face. (picture) Another manifestation is our tendency to apply newly popularized or trendy concepts to to totally inappropriate circumstances. For example, once Clayton Christensen popularized "disruptive innovation", any situation where technology brought about change was all of a sudden being labeled as "disruptive".

My latest peeve is what I perceive to be pareidolic use of the term "network effect" to describe almost any example of positive feedback in markets. For example, here's an example that Tim O'Reilly thinks is a network effect:
Google is better at spidering that network than their competitors. They thus benefit more powerfully from the network that we are all collectively building via our web publishing and cross-linking.
While there is definitely a network that enables Google's spidering, it's not a "network effect" that makes Google a good spiderer. Economies of scale are what make Google a good spiderer, even if that scale has resulted in part from network effects.

The originator of the term "network effect" was Bob Metcalfe, the co-inventer of Ethernet. He used it to refer to a mathematical description of how the value of a network scaled with the number of nodes it connected. His reasoning was that the value of each networked node is proportional to the number of other nodes it can connect with, so that the total value of the network scales with the square of the number of nodes.

It's pretty silly to expect that a scaling rule that works for small networks would continue to apply for large networks, and Andrew Odlyzko and Benjamin Tilley have pointed out that more modest scaling laws are a much better fit to market valuations of networks. Still, their suggestion that inappropriate application of Metcalfe's law was to blame for the internet bubble and its subsequent collapse bears reflection.

Recently there's been some discussion of how to apply Metcalfe's law for the network effect to Linked Data and the Semantic Web. Linked Data is information published using standards so that machines can understand its meaning and make inferences from the totality of data that has been collected. One argument says that the value of any set of Linked Data increases in value with every new bit of linked data is added to the world wide cloud of Linked Data. So how does the value of this Linked Data "Network" really scale with the links it contains?

Since I don't know of any way to value an arbitrary bit of Linked Data, I'll pick a simple system where I can compute utility. I'll focus on the direct effects and benefits of linking data together, and ignore for now indirect benefits such as those which result from the use of standards.

Suppose we have two sets of Linked Data entities, Movies and Actors. Let's also assume that both of these sets are essentially complete. We'll then consider the effect on the system utility of adding random "actedIn" links between Actors and Movies. In our utility computation, we'll assume that answering questions about which actors acted in which movies is the primary utility of our set of links, and the number of these questions the set can answer will be the utility measure.

For the questions "What movies did X act in?" and "who acted in the movie Y?" the value of the link collection scales linearly with the number of "actedIn" links. There's no network effect at all for these questions, because the fact that Marlon Brando acted In On the Waterfront adds no utility to the fact that Humphrey Bogart acted in Casablanca.

For the question "Who else acted in movies that X acted in?" the result is different. For this question, the ability of our collection of links to answer usefully scales as the square of the number of links, just as in the classic network case. For this question, there clearly is a network effect.

For the "Kevin Bacon" question, ("how many acted-with degreees of separation are there between X and Kevin Bacon?") the Network effect is even stronger, with the network value scaling as a higher power of the number of actedIn links. Notice that for Linked Data, the network effect is not inherent in the data, but rather is implicit in the types of queries that are made on the data.

What we really wanted to know was the total value of the set of links. We might guess that the value is proportional to the total number of questions that the set of links can answer. That number grows exponentially with the number of links. We can see that an exponentially increasing value doesn't make sense, however, by considering the total value as a power series. We've already discussed the first two terms of that series, but we've not considered the relative coefficient. It's really hard to argue that the value of a system that can answer only the acted-with questions is hugely more valuable than the one that answers only the acted-in questions, despite the fact that it answers N2 questions compared to only N for the acted-in answering system. It seems to me that the Kevin Bacon answering system is less valuable than the other two systems, despite an even larger number of questions (about N12) that it would be able to answer; they're just really stupid questions.

Even if network effects are not inherent in Linked Data, threshold effects can result in the existence of a "critical mass", above which positive reinforcement kicks in to drive the entire system. In our toy system, we can easily imagine that a collection of links might be worthless until there were a sufficient number of links to exceed critical mass. A system that can tell me 90% of the movies that someone has acted in is a lot more than nine times as valuable as a system that can tell me only 10%. That's because an acted-in answering system is worthless unless it's better than the random guy sitting next to me at the bar! So this is kinda-sorta a network effect, but really it's a threshold effect.

It's rather easy to confuse threshold effects for network effects. I'll put it this way: it's not a network effect that causes me to avoid doing my laundry until I have a full load, it's a threshold effect! Never mind that it's really three loads.

Rod Beckstrom, currently CEO of ICANN has described the "inverse" network effect, which occurs in situations where the addition of nodes reduces the value of a network to each participant. Golf clubs are cited as examples- they have an optimum size of about 500 members because additional members make it more difficult for existing members to get playing time. I see two types of inverse network effects at play in the Linked Data world. The first is the cost of expanding a database; the second is the law of diminishing returns.

In an optimally designed database, the cost of accessing any given record is proportional to the logarithm of the number of records. This is a slowly increasing function- if you have this scaling, you can increase from a million records to 10 million records and only increase your cost by 18%. Alas, optimal design is rarely achieved, and to get to that optimum, you have costs that scale much less gently. The result has been that most practical applications of Linked Data use only the most relevant subsets of available data. If data network effects were pervasive and stronger than inverse network effects, this would generally not happen.

In my movie and actor example, I made the assumption that the value of any particular query was equal in value to any other query. In practice this is not true. Most people would agree that there's more utility in knowing that Humphrey Bogart acted in Casablanca, than in knowing that Michael Ripper acted in The Reptile. In many data sets, the 80/20 rule applies (also known as the Pareto Principle) - 80% of the real-world queries would exercise only 20% of the links. The least useful data is typically the most expensive data to acquire, so if we start by adding the most valuable actedIn links, then every additional link reduces the average value of the links in the collection. This mimics an inverse network effect, as the the total value of the collection grows more slowly with every added link, rather than growing more quickly.

The main take-away from this is that you can't look at Linked Data objectively and conclude that it exhibits strong network effects without taking into account the application it's being used for. Some applications will exhibit strong, even exponential network effects, others may exhibit inverse network effects. And sometimes a grilled cheese sandwich is just a sandwich.

Reblog this post [with Zemanta]

Article any source

Friday, October 2, 2009

Google's Wave is a Tsunami of Conversation

Human communication is a magical thing, and over the last decade, we've acquired some super-powers. We've had the telephone for a hundred and thirty three years and it is still developing. Nowadays you can hardly set your twelve-year-old loose in the mall without packing a cell phone, and people are dying because they just have to send texts while driving. Just think about all the "conversation" tools we've added since the birth of the internet: e-mail, instant messaging, chat rooms, mailing lists, bulletin boards, blogs, various and sundry social networks, Facebook, Skype, Twitter. Our use of these tools will continue to evolve, but at some point there's going to be a consolidation. Do we really need both Twitter and Facebook?

Yesterday, I was lucky enough to get an invitation to the Google Wave Preview. I have no idea whether Wave will be the next big thing, but I can tell you it won't fail for lack of ambition. Unlike Twitter, which started with a concept so simple that it sounded really stupid, Wave starts by imagining what email would look like if its designers were to start from scratch, knowing what we know today. The result is a daunting attempt to roll all of our communication superpowers into a single user interface. At first, I couldn't really figure out how to get Wave to do what I wanted it to do. Part of the problem was that some of the features mentioned in the help videos had not been activated in the Preview version, though they are still working in the "Sandbox" version which has been available to developers for a few months now. Other things just don't work yet- the contacts module seems very buggy, which is a big problem because you need that to connect with other "Wavers". Luckily, my multitool communications mischmash is still working, and I was able to find some friends (via Twitter and Facebook) who were in the same situation as myself, still stumbling around a dark room of functionality, looking for someone to connect with.

With some help here and there, I gradually figured things out. Wave introduces several new (to me) user-interface widgets; it struck me that web applications rarely introduce new interface widgets, unlike applications like Excel or Photoshop that put powerful capabilities in inteface objects. I still haven't figured out the sliders. The threaded discussions have little colored boxes that indicate who is typing something. The overall effect is that gMail has been tricked out and turbocharged. After a day of playing with it, I'm not sure I like the user interface, but I can't really think af any way to make it better given the scope of what Wave is trying to do.

The focus of Wave is editable, threaded conversations, i.e. waves. (There's already a convention to use lower case w for the conversations and upper case W for the platform as a whole.) These are quite well done. They support hierarchical threading, styled text, auto-linking, distributed editing, history, history playback, insertion of images, videos, and software objects known as extensions (there is a poll widget pre-installed). The one thing that's missing is undo. I can think of all sorts of uses for these capabilities; I intend to try a few over the next week or so.

The intro videos plug a tool called Bloggy that allows you to publish a wave to your blog. Bloggy is a robot that you add as a participant in your conversation. I spent about an hour trying to figure out how to get Bloggy to work, until I discovered (via Twitter) that Bloggy had not been activated on the Preview version yet.

Having failed with Bloggy, I started thinking "Is that all there is?" and envying Google's white-hot overhype machine. But then @jillmwo pointed me to the way to search on public waves (put "with:public" before a search). Immediately the little ripply waves I was seeing turned into a tsunami, and I began to see the enormous possibilities of Wave. To make a wave public, you add a participant called "Public" (public@a.gwave.com) to the wave; "Bloggy" does the same thing.

The ability to search public waves (and to make your own public wave) is a feature qualitatively different from anything I've seen before. A public wave is sort of the bastard offspring of a Twitter hashtag mated with a Wikipedia topic page. When it matures, this creature will surely be a monstrous beast; what no one can tell yet is whether the beast will be a tame workhorse or whether it will be a velociraptor requiring a good strong cage. New technology is always easy to create and control compared to new social practice.

If it was me trying to invent email from scratch, I would spend 80% of my effort on spam prevention. It looks to me as though many of Wave's feaures have been hidden in the Preview because the functions needed for spam prevention are not ready yet. This is not to say that the Wave team is not spending 80% of its effort on spam prevention- these are incredibly hard things to get right. As an example, it's currently not possible to remove a participant from a wave, and any public wave participant can see the addresses of all the other participants. The problems with the contacts module are also probably related- waves from people you might not know just appear in your inbox, even though it looks like the intent is for them to first appear in the "Requests" folder in the "Navigation" module. Groups are also not yet implemented. It's difficult to know how this will play out until we see everything working.

Since the Preview roll-out, there have been quite a number of negative reviews by people pointing to all the problems in Google Wave. I think these are missing the point to some extent. Sure, Wave could end up being a complete failure, but even if that happens, Wave is giving us today a glimpse of what the future could look like. If not Wave, then surely there will be a SuperTwitter or a SpaceBook or maybe even a MagicForce that will contend for the consolidation of our one hundred flowers of conversation.

Note: this post will also be published as a public wave.

Reblog this post [with Zemanta]

Article any source

Friday, September 25, 2009

A Reading Miracle. It May Be Legal, but Don't Ask the Grox

Over the summer I witnessed a miracle.

Do you know which book was the first you ever read on your own? I'm not sure about mine, but my younger brother's first was definitely Go, Dog. Go! by P. D. Eastman. He was under 4 years old when he start reading it. If you haven't read it, Go, Dog. Go! is a 62 page multiculturalist masterpiece with engaging illustrations first published in 1961. Here is the complete text of pages 3-9:
Dog. Big dog. Little dog. Big dogs and little dogs. Black dogs and white dogs. "Hello!" "Hello!" "Do you like my hat?" "I do not." "Good-by!" Good-by!"
This summer, I witnessed my own son "reading" his first "book". It wasn't written by a single author and it wasn't published by Random House. It wasn't printed on paper, and it wasn't even what we might call an "e-book". It was a website devoted to the game "Spore" that currently consists of 3,819 articles written by website users, and over the course of the summer, my son read a majority of those articles. Here is a sample passage:
The Grox are a sentient species of cyborg aliens generally considered to be the most evil and hostile in the galaxy. They are most notable for their evil and hostility, but are also notable for their asymmetric, weak impish appearance.
Needless to say, my son's outlook on the world and his ability to explore it have dramatically changed.

This miracle was made possible by text-to-speech (TTS) software. You see, my son has a disability that makes normal reading excruciatingly difficult for him. Through a great deal of work, and some considerable courage, he is now able to read, with great effort, printed sentences and short paragraphs on his own. But as a bright 11-year-old sixth grader, books like Go, Dog. Go! and others that pose little reading difficulty hold little interest for him, and so he won't read the words in printed books on his own. He likes computers, though. At the beginning of the summer, I showed my son how to activate the text-to-speech features of his Mac. Mac OS X has text-to-speech capabilities built in, and because Mac applications are built using standardized text display objects with hooks that allow access to the system TTS services, there's a uniform, cross-application way to have text spoken. (In contrast, TTS on Windows Vista is almost useless!) Similarly, Wiki-based websites present content in uniform ways that made it easy for my son to interact with text.

I was amazed by the way my son began to devour the content that interested him. Every day after coming home from camp, he would spend hours staring at the screen and listening to the Mac's robotic voice speak the text to him. Then he would watch some YouTube videos and play some Spore. I realized that TTS had given my son a way to fully satisfy, for the first time in his life, his hunger for information.

People who see miracles tend to develop intense beliefs. I am no exception. I am no longer an objective observer of digital copyright issues when they relate to access by the reading disabled. When I want to feel some anger (it helps me run faster) I think about people and institutions who try to use copyright law in ways that prevent people like my son from being able to read what they want to read.

After having moral imperatives made clear to me, I've spent some time learning about the relevant technology and laws, and I find that these include many of the issues I've been working on and learning about. For example, last year, before I started paying attention, Amazon faced criticism from authors and publishers who argued that text to speech on the Kindle DX constituted a performance that Amazon did not have the rights to deliver. Could publishers similarly enjoin Apple from allowing my son to use its TTS on copyrighted material? With my new perspective, I cannot talk about this without fuming at the blatant immorality of some of the arguments being made.

When Amazon relented, Random House (publisher of Go, Dog. Go!) asked Amazon to turn off text-to-speech on the Kindle DX for its books, which sparked considerable controversy. This led the National Federation of the Blind and American Council of the Blind to file a discrimination lawsuit against Arizona State University which intended to test the Kindle DX as a means of distributing textbooks. The basis of this lawsuit is the Americans with Disabilities Act (ADA), which bars discrimination against people with disabilities in any public accommodation, a term which would include libraries and bookstores. The ADA has been used to force e-commerce websites to make their websites accessible to people with disabilities.

Unfortunately, the laws on accommodating disabled users have not kept up with changing technology. In 1996, the "Chafee Amendment" changed US Copyright law to allow "authorized entities" to make reproductions of previously published nondramatic literary works for the purpose of producing formats used exclusively by the disabled. Unfortunately, the possibility that all the worlds books might someday be digitized and thus made available to those with reading disabilities was remote at that time. As a result the ambiguity of the amendment's language is enough that the American Association of Publishers was able to argue that the Chafee Amendment could not be used by libraries to help them comply with the ADA. Luckily, organizations like Benetech and its BookShare website are working with publishers to get around this sort of conflict. I hope my son will be able to read the books he needs to read through BookShare.

It's my considered opinion that the Google Book Search digitization project has created the potential for a direct collision between book publishers and the ADA, and that this prospect has played a significant role in shaping the controversial aggrement to settle the publishers' and authors' lawsuit against Google, but that's a topic for another article.

Now that I know a bit more about the potential legal obstacles to my son's reading, I'm wondering what I should be doing to make sure those obstacles disappear. I'm still hoping to see more reading miracles. "Good-by!"
Reblog this post [with Zemanta]

Article any source

Saturday, September 5, 2009

RDF Properties on Magic Shelves

Book authors and politicians who go on talk shows, whether it's the Daily Show, Charlie Rose, Fresh Air, Oprah, Letterman, whatever, seem to preface almost every answer with the phrase "That's a really good question, (Jon|Teri|Stephen|Conan)". The Guest never says why it's a good question because real meaning of that phrase is "Thanks for letting me hit one out of the ballpark." Talk shows have so little in common with baseball games or even tennis matches. On the rare occasion when a guest doesn't adhere to form, the video goes viral.

I've been promising to come back to my discussion of Martha Yee's questions on putting bibibliographic data on the semantic web. Karen Coyle has managed to discuss all of them at least a little bit, so I'm picking and choosing just the ones that interest me. In this post, I want to talk about Martha's question #11:
Can a property have a property in RDF?
The rest of my post is divided into two parts. First, I will answer the question, then in the second part, I will discuss some of the reasons that it's a really good question.

Yes, a property can have a property in RDF. In the W3C Recommentation entitled RDF Semantics, it states: "RDF does not impose any logical restrictions on the domains and ranges of properties; in particular, a property may be applied to itself." So not only can a property have a property in RDF, it can even use itself as a property!

OK, that's done with. Not only is the answer yes, but it's yes almost to the point of absurdity. Why would you ever want a property to be applied to itself? How can a hasColor property have a hasColor property? If you read and enjoyed Gödel, Escher, Bach, you're probably thinking that the only use for such a construct is to define a self-referential demonstration of Gödel's Incompleteness Theorem. But there actually are uses for properties which can be applied to themselves. For example, if you want to use RDF properties to define a schema, you probably want to have a "documentation" property, and certainly the documentation property should have its own documentation.

If you're starting to feel queasy about properties having properties, then you're starting to understand why Yee question 11 is a good one. Just when you think you understand the RDF model as being blobby entities connected by arcs, you find out that the arcs can have arcs. Our next question to consider is whether properties that have properties accomplish what someone with a library metadata background intends them to accomplish, and even if they do so, is it the right way to accomplish it?

In my previous post on the Yee questions, I pointed out that ontology development is a sort of programming. One of most confusing concepts that beginning programmers have to burn into their brains is the difference between a class and an class instance. In the library world, there are some very similar concepts that have been folded up into a neat hierarchy in the FRBR model. Librarians are familiar with expressions of works that can be instantiated in multiple manifestations, each of which can be instantiated in multiple items. Each layer of this model is an example of the class/instance relationship that is so important for programmers to understand. This sort of thinking needs to be applied to our property-of-a-property question. Are we trying to apply an property to an instance of a property, or do we want to apply properties to property "classes"?

Here we need to start looking at examples, or else we will get hopelessly lost in abstraction-land. Martha's first example is a model where the dateOfPublication is a property of a publishedBy relationship. In this case, what we really want is a property instance from the class of publishedBy properties that we modify with a dateOfPublication property. Remember, there is a URI associated with the property piece of any RDF triple. If we were to simply hang a dateOfPublication on a globally defined publishedBy we would have made that modification for every item in our database using the publishedBy attribute. That's not what we want. Instead, for each publishedBy relation we wanted to assert, we need to create a new property, with a new URI, related to publishedBy using the RDF Schema property subPropertyOf.

Let's look at Martha's other example. She wants to attach a type to her variantTitle property to denote spine title, key title, etc. In this case, what we want to do is create global properties that retain variantTitleness while making the meaning of the metadata more specific. Ideally, we would create all our variant title properties ahead of time in our schema or ontology. As new cataloguing data entered our knowledgebase, our RDF reasoning machine would use that schema to infer that spineTitle is a variantTitle so that a search on variantTitle would automatically pick up the spineTitles.

Is making new properties by adding a property to a subproperty the right way to do things? In the second example, I would say yes. The new properties composed from other properties make the model more powerful, and allow the data expression to be simpler. In the first example, where a new property is composed for every assertion, I would say no. A better approach might be to make the publication event a subject entity with properties including dateOfPublication, publishedBy, publishedWhat, etc. The resulting model is simpler, flatter, and more clearly separates the model from the data.

We can contrast the RDF approach of allowing new properties to be created and modified by other properties to that of MARC. MARC makes you to put data in fields and subfields and subfields with modifiers, but the effect is sort of like having lots of dividers on lots shelves on a bookcase- there's one place for each and every bit of data- unless there's no place. RDF is more like a magic shelf that allows things to be in several places at once and can expand to hold any number of things you want to put there.

"Thanks for having me, Martha, it's been a real pleasure."
Reblog this post [with Zemanta]

Article any source

Friday, July 31, 2009

Ignition Timing for Semantic Web Library Automation Engines

Last weekend, I had a chance to learn how to drive a 1915 Model-T Ford. It's not hard, but a Model-T driver needs to know a bit more about his engine and drivetrain than the driver of a modern automobile. There is a clutch pedal that puts the engine into low gear when you press it- high gear is when the pedal is up and neutral is somewhere in between. The brake is sort of a stop gear, and you need to make sure the clutch is in neutral before you step on the brake. The third pedal is reverse.

There are a lot more engine controls than on a modern car. In addition to the throttle and a choke, there is another lever that controls the ignition timing. A modern Model-T driver doesn't have to worry much about the timing once the engine has started, because modern fuel has much higher octane than fuel had in 1915. I would not have understood this except that I recently got a new car whose manual says you should use only premium fuel, and so I did some wikipedia research to find out what octane had to do with automobile engines. But I could have lived blissfully in ignorance. Believe it or not, I have opened the hood of my new car only once since I got it in December.

It occurs to me that in many ways, the library automation industry is still in the Model-T era, particularly in regards to the relationship of the technology to its managers. Libraries still need to keep a few code mechanics on staff, and the librarians who use library automation to deliver services still need to know a lot more about their data engines than I know about my automobile engine. The industry as a whole is trying to evaluate changes roughly analogous to the automobile industry switching to diesel engines.

I've been reading Martha Yee's paper entitled "Can Bibliographic Data Be Put Directly Onto the Semantic Web?" and Karen Coyle's commentary on this paper. I greatly admire Martha Yee's courage to say, essentially, "I don't understand this as well as I need to, here are some questions I would really appreciate help with". When I worked at Bell Labs, I noticed that the people who asked questions like that were the people who had won or would later win Nobel prizes. Karen has done a great job with Martha's queries, but also expresses a fair amount of uncertainty.

I was going to launch into a few posts to help fill in some gaps, but I find that I have difficulty knowing which things are important to explain. Somehow I don't think that Model-T drivers really needed to know about the relationship between octane and ignition timing, for example. But I think that people running trucking companies need to know some of the differences between Diesel engines and internal combustion engines as they built their trucking fleets, just as community leaders like Martha Yee and Karen Coyle probably need to know the important differences between RDF tuple-stores and relational databases. But the more I think about it, the less I'm sure about which of the differences are the important ones for people looking to apply them in libraries.

Another article I've been reading has been Greg Boutin's article "Linked Data, a Brand with Big Problems and no Brand Management", which suggests that the technical community that is pushing RDF Linked Data has not been doing a good job of articulating the benefits of RDF and Linked Data principles in a way that potential customers can understand clearly and consistently.

Engineers tend to have a different sort of knowledge gap. I have a very good friend who designs the advanced fuel injectors. He is able to do this because he has specialized so that he knows everything there is to know about fuel injectors. He doesn't need to know anything about radial tires or airbag inflators or headlamps. But to make his business work, he needs to be able to articulate to potential customers the benefits of his injectors in the context of the entire engine and engine application. Whether the technology Linked Data or fuel injectors, that can be really difficult.

My first guess was that it would be most useful for librarians to understand how indexing and searching are almost the same thing, and that indexing done quite differently in RDF tuple-stores and in relational databases. But on second thought, that's more like telling the trucking company that diesel engines don't need spark plugs. It's good to know, but the higher-level fact that diesels burn less fuel is a lot more relevant? Isn't it more important to know that an RDF tuple-store trades off performance for flexibility? How do you ask the right questions to ask, when you don't know where to start? We find ourselves working across many disciplines each of which are more and more specialized, and we need more communications magic to make everything work together.

I'll try to do some gap-filling next week.
Article any source