Showing posts with label Network Effect. Show all posts
Showing posts with label Network Effect. Show all posts

Monday, January 25, 2010

8 One-Way Business Models for Linked Data

Real Wheels - Travel Adventures (There Goes a Train/Plane/Bus)One of the videos that I was forced to watch many times when my boys were younger was There Goes a Train. It's a pretty good video. I learned many things, including the fact that locomotives don't have steering wheels. Yeah. Pretty obvious if you think about it for even a moment. It's so unfair. Without the switches, the rail network would be pretty useless, but no one will ever make a video entitled There Sits a Switch.

In electronics though, the switch is the star. With a switch, you can modify and route information; with a wire you can only send it from one point to another. You need good switches to make a computer or a network; even though photons are faster and easier to move from one place to another, computers are still based on electrons because electronic switches are so much better than optical switches.

Linked Data is a label for a set of technologies that are trying to make information move around the internet more easily and with more meaning. The Linked Data vision is one where many entities acting cooperatively and globally create a web of data much more powerful and meaningful than any single entity could bring about.

For the Linked Data vision to become a reality, each entity must have a strong motivation to cooperate; each entity must have a viable business model. If the business models were easy, the Linked Data vision would already be a vibrant reality.

Scott Brinker recently launched a round of discussion about seven business models that can make Linked Data viable. Leigh Dodds contributed some important insights in his followup, prompting Brinker to add an eighth model.

Here are Brinker's eight business models for Linked Data (somewhat relabeled based on who's writing checks):
  1. Subsidy. Entities such as governments with a mandate to make information available will pay to have it linked into a global web of information.
  2. Subscription. People will pay for valuable data, and will pay more for data that has been linked to a global web of information.
  3. Advertising. Advertisers will pay to information in raw data feeds.
  4. Authority. People will pay for the validation and certification of data.
  5. Affiliate marketing. Merchants will pay sales commissions on sales resulting from affiliate links in embedded in the global web of data
  6. Service Enhancement. People will pay for services which have been enhanced by data from a global web.
  7. Search Engine Optimization. Search engines will send you more traffic if you give them more meaningful data.
  8. Brand Enhancement. Your reputation will be burnished if you emit lots of good information.
(I should note that Brinker describes each model a bit differently so that he can add a dimension that characterizes whether data is delivered raw or as an application.  I find that this dimension is not at all orthogonal. A data driven subscription service is a service that makes use of data, but the core business model is not to sell a data subscription.)

There are difficulties with all of these business models, but it strikes me that each of them will only work in one direction, like a train track without switches. Either they work for emitting data, or they work for consuming data, but none of the models work in both directions at the same time. If you're providing a service that's either based on Linked Data or enhanced by it, you can pay for the data, but if you send that data back out, your competitors get the data for free. Conversely, if you're emitting data, it's hard for you to pay for it.

Imagine you're in the book metadata business. You can use several of these models to support creation of book metadata, or you can consume book-related metadata to provide book-related services. But what if you want to support an activity of aggregating book data or fixing errors in book metadata? None of these models will work for you because you'll either be competing with the entities you get data from, or you'll be competing with entities you send data to.

What's missing from this list is a business model for the Linked Data switch. Entities that take in Linked Data, improve it or otherwise add value and reemit it as Linked Data have no solid business model to run on. Everyone active so far in the Linked Data business is either a data sink or a data source. To realize the full potential of Linked Data, there need to be viable switches, both collecting and emitting Linked Data.
Reblog this post [with Zemanta]

Article any source

Wednesday, December 2, 2009

Databases are Services, NOT Content

I'm very grateful for advice my fellow entrepreneurs have given me; when you meet someone else who has started a company you have an instant rapport from having shared a common experience. I remember each bit of advice with the same stark clarity that characterizes the moment I realized that Santa Claus was the neighbor dressed up in a white beard and a red suit..

A business owner who I've known since sixth grade gave me this gem: "My secret is providing the best service possible, and charging a lot for it." In executing my own business, I did pretty well at the first part and could have done better at the second part.

As I wrote about the effect of database rights on the postcode economy, I kept wondering if I would have done anything differently in my business if that database protection had been available to me. Would I have charged more for the database that my company developed?

In the comments to that post, I was alerted to a book by James Boyle, called The Public Domain. Chapter 9 in particular parallels many of the arguments I made. One thing I found there was something I had wanted to look for- information about how the database industries in general have done since the "sweat of the brow" theory for copyright was disallowed by the Supreme Court. It turns out that the US database industry has actually outpaced its counterpart in the UK since then by a substantial margin. Why would that be?

I think the answer is that building databases is fundamentally a service business. If your brow is really sweating, and someone is paying you to do it, then it's hard to think of that as a "content" business. Databases always have more content than anyone could ever want; the only reason people pay for them is that they help to solve some sort of problem. If your business thinks it's selling content rather than services, chances are it will focus on the wrong part of the business, and do poorly. In the US, since database companies understand that their competition can legally copy much of their data, they focus on providing high quality added value services, and guess what? THEY MAKE MORE MONEY!

Then there's Linked Data. Given that database provision is fundamentally a service business, is it even possible to make money by providing data as Linked Data? The typical means for prodecting a database service business is to execute license agreements with customers. You make an agreement with your customer about the service you'll provide, how much you'll get paid, and how your customer may use your service. But once your data has been released into a Linked Data Cloud, it can be difficult to assert license conditions on the data you've released.

It's been argued that 'Linked Data' is just the Semantic Web, Rebranded, but it's also been noted both Linked Data is sorely in need of some proper product management. Product management focuses on a customer's problems and how the product can address them. You can believe me because I've not only managed products, I've had 2 whole days of real product management training!

One thing I was taught to do in my Product Management class was to come up with a 1 sentence pitch that captures the essence of the product. When I was an entrepreneur, this was called the elevator pitch. After thinking about it for about 9 months I've come up with a pitch for Linked Data:
Linked Data is the idea that the merger of a database produced by one provider and another database produced by a second provider has value much larger than that of the two separate databases.
or, in a more concise form, V(DB1+DB2)>>V(DB1)+V(DB2).

Based on the products that have been successful this year in the application of semantic web technologies, it looks to me that the most successful have been focused on what I saw Tim Gollins tweet that Ian Davis called "Linked Enterprise Data" (attributing the term to Eric Miller). If the merged databases are contained within the enterprise, the enterprise clearly reaps all the added value. Outside the enterprise, however, the only Linked Open Data winners so far have been the ones who have built services on databases merged from others.

Proper product management would have made it a goal for Linked Open Data to have data contributors share somehow in the surplus value created by the merged services. In the next couple of weeks, I hope to describe some ideas as to how this could happen.
Article any source

Thursday, October 15, 2009

Normal and Inverse Network Effects for Linked Data

The human brain has an amazing capacity to recognize familiar patterns in unfamiliar environments. One manifestation of this is pareidolia, the phenomenon of seeing an image of the Virgin Mary in a grilled cheese sandwich or a mesa on Mars that looks like a face. (picture) Another manifestation is our tendency to apply newly popularized or trendy concepts to to totally inappropriate circumstances. For example, once Clayton Christensen popularized "disruptive innovation", any situation where technology brought about change was all of a sudden being labeled as "disruptive".

My latest peeve is what I perceive to be pareidolic use of the term "network effect" to describe almost any example of positive feedback in markets. For example, here's an example that Tim O'Reilly thinks is a network effect:
Google is better at spidering that network than their competitors. They thus benefit more powerfully from the network that we are all collectively building via our web publishing and cross-linking.
While there is definitely a network that enables Google's spidering, it's not a "network effect" that makes Google a good spiderer. Economies of scale are what make Google a good spiderer, even if that scale has resulted in part from network effects.

The originator of the term "network effect" was Bob Metcalfe, the co-inventer of Ethernet. He used it to refer to a mathematical description of how the value of a network scaled with the number of nodes it connected. His reasoning was that the value of each networked node is proportional to the number of other nodes it can connect with, so that the total value of the network scales with the square of the number of nodes.

It's pretty silly to expect that a scaling rule that works for small networks would continue to apply for large networks, and Andrew Odlyzko and Benjamin Tilley have pointed out that more modest scaling laws are a much better fit to market valuations of networks. Still, their suggestion that inappropriate application of Metcalfe's law was to blame for the internet bubble and its subsequent collapse bears reflection.

Recently there's been some discussion of how to apply Metcalfe's law for the network effect to Linked Data and the Semantic Web. Linked Data is information published using standards so that machines can understand its meaning and make inferences from the totality of data that has been collected. One argument says that the value of any set of Linked Data increases in value with every new bit of linked data is added to the world wide cloud of Linked Data. So how does the value of this Linked Data "Network" really scale with the links it contains?

Since I don't know of any way to value an arbitrary bit of Linked Data, I'll pick a simple system where I can compute utility. I'll focus on the direct effects and benefits of linking data together, and ignore for now indirect benefits such as those which result from the use of standards.

Suppose we have two sets of Linked Data entities, Movies and Actors. Let's also assume that both of these sets are essentially complete. We'll then consider the effect on the system utility of adding random "actedIn" links between Actors and Movies. In our utility computation, we'll assume that answering questions about which actors acted in which movies is the primary utility of our set of links, and the number of these questions the set can answer will be the utility measure.

For the questions "What movies did X act in?" and "who acted in the movie Y?" the value of the link collection scales linearly with the number of "actedIn" links. There's no network effect at all for these questions, because the fact that Marlon Brando acted In On the Waterfront adds no utility to the fact that Humphrey Bogart acted in Casablanca.

For the question "Who else acted in movies that X acted in?" the result is different. For this question, the ability of our collection of links to answer usefully scales as the square of the number of links, just as in the classic network case. For this question, there clearly is a network effect.

For the "Kevin Bacon" question, ("how many acted-with degreees of separation are there between X and Kevin Bacon?") the Network effect is even stronger, with the network value scaling as a higher power of the number of actedIn links. Notice that for Linked Data, the network effect is not inherent in the data, but rather is implicit in the types of queries that are made on the data.

What we really wanted to know was the total value of the set of links. We might guess that the value is proportional to the total number of questions that the set of links can answer. That number grows exponentially with the number of links. We can see that an exponentially increasing value doesn't make sense, however, by considering the total value as a power series. We've already discussed the first two terms of that series, but we've not considered the relative coefficient. It's really hard to argue that the value of a system that can answer only the acted-with questions is hugely more valuable than the one that answers only the acted-in questions, despite the fact that it answers N2 questions compared to only N for the acted-in answering system. It seems to me that the Kevin Bacon answering system is less valuable than the other two systems, despite an even larger number of questions (about N12) that it would be able to answer; they're just really stupid questions.

Even if network effects are not inherent in Linked Data, threshold effects can result in the existence of a "critical mass", above which positive reinforcement kicks in to drive the entire system. In our toy system, we can easily imagine that a collection of links might be worthless until there were a sufficient number of links to exceed critical mass. A system that can tell me 90% of the movies that someone has acted in is a lot more than nine times as valuable as a system that can tell me only 10%. That's because an acted-in answering system is worthless unless it's better than the random guy sitting next to me at the bar! So this is kinda-sorta a network effect, but really it's a threshold effect.

It's rather easy to confuse threshold effects for network effects. I'll put it this way: it's not a network effect that causes me to avoid doing my laundry until I have a full load, it's a threshold effect! Never mind that it's really three loads.

Rod Beckstrom, currently CEO of ICANN has described the "inverse" network effect, which occurs in situations where the addition of nodes reduces the value of a network to each participant. Golf clubs are cited as examples- they have an optimum size of about 500 members because additional members make it more difficult for existing members to get playing time. I see two types of inverse network effects at play in the Linked Data world. The first is the cost of expanding a database; the second is the law of diminishing returns.

In an optimally designed database, the cost of accessing any given record is proportional to the logarithm of the number of records. This is a slowly increasing function- if you have this scaling, you can increase from a million records to 10 million records and only increase your cost by 18%. Alas, optimal design is rarely achieved, and to get to that optimum, you have costs that scale much less gently. The result has been that most practical applications of Linked Data use only the most relevant subsets of available data. If data network effects were pervasive and stronger than inverse network effects, this would generally not happen.

In my movie and actor example, I made the assumption that the value of any particular query was equal in value to any other query. In practice this is not true. Most people would agree that there's more utility in knowing that Humphrey Bogart acted in Casablanca, than in knowing that Michael Ripper acted in The Reptile. In many data sets, the 80/20 rule applies (also known as the Pareto Principle) - 80% of the real-world queries would exercise only 20% of the links. The least useful data is typically the most expensive data to acquire, so if we start by adding the most valuable actedIn links, then every additional link reduces the average value of the links in the collection. This mimics an inverse network effect, as the the total value of the collection grows more slowly with every added link, rather than growing more quickly.

The main take-away from this is that you can't look at Linked Data objectively and conclude that it exhibits strong network effects without taking into account the application it's being used for. Some applications will exhibit strong, even exponential network effects, others may exhibit inverse network effects. And sometimes a grilled cheese sandwich is just a sandwich.

Reblog this post [with Zemanta]

Article any source