Showing posts with label knowledgebases. Show all posts
Showing posts with label knowledgebases. Show all posts

Saturday, January 8, 2011

Inside the Dataculture Industry

wild blueberries
I don't really know how all the food gets to my table. Sure, I've gathered berries, baled hay, picked peas, baked bread and smoked fish, but I've never slaughtered a pig, (successfully) milked a cow or roasted coffee beans. In my grandparents generation, I would have seemed rather ignorant and useless. Agriculture has become an industry as specialized as any other modern industry; increasingly inaccessible to the layperson or small business.

I do know a bit about how data gets to my browser. It gets harvested by data farmers and data miners, it gets spun into databases, and then gets woven into your everyday information diet. Although you've probably heard of the "web of data", you're probably not even aware of being surrounded by data cloth.

The dataculture industry is very diverse, reflecting the diversity of human curiosity and knowledge. Common to all corners of the industry is the structural alchemy that transmutes formless bits into precious nuggets of information.

In many cases, this structuring of information is layered on top of conventional publishing. My favorite example of this is that the publishers of "Entertainment Week" extract facts out of their stories and structure them with an extensive ontology. Their ontologists (yes, EW has ontologists!) have defined an attribute "wasInRehabWith" so that they can generate a starlet's biography and report to you that she attended a drug rehabilitation clinic at the same time as the co-star of her current movie. Inquiring minds want to know!

If you look at location based services such as Facebook's "places", Foursquare, Yelp, Google Maps, etc, they will often present you with information pulled from other services. Often, a description comes from Wikipedia and reviews come from Yelp or Tripadvisor and photos come from Panoramio or Flickr. These services connect users to data using a common metadata backbone of Geotags. Data sets are pulled from source sites in various ways.

Some datasets are produced in data factories. I had a chance to see one of these "factories" on my trip to India last month. Rooms full of data technicians (women do the morning shift, men the evening) sit at internet connected computers and supervise the structuring of data from the internet. Most of the work is semi-automated, software does most of the data extraction. The technicians act as supervisors who step in when the software is too stupid to know when it's mangling things and when human input is really needed.

There's been a lot of discussion lately about how spammers are using data scraped from other websites and ruining the usefulness of Google's search results. There are plenty of companies that offer data scraping services to fuel this trend. Data scraping is the use of software that mimics human web browsing to visit thousands of web pages and capture the data that's on them. This works because large websites are generated dynamically out of databases; when machines assemble web pages, machines can disassemble them.

A look at the variety of data scraping companies reveals a broad spectrum. Scraping is an essential technology for dataculture; as with any technology, it can be used to many ends. One company boasts of their "massive network of stealth scrapers capable of downloading massive amounts of data without ever getting blocked. Some companies, such as Mozenda, offer software to license. Others, such as Xtractly and Addtoit are strictly service offerings.

I spoke to Addtoit's President, Bill Brown, about his industry. Addtoit got its start doing projects for Reuters and other firms in the financial industry; their client base has since become more "balanced". Companies such as Bloomberg, Reuters and D&B get paid premiums to provide environments rich in structured data by customers wanting a leg up on competitors. Brown's view is that the industry will move away from labor intensive operations to being completely automated, and Addtoit has developed accordingly.

A small number of companies, notably Best Buy, have realized that making their data easily available can benefit them by promoting commerce and competition. They have begun to use technologies such as RDFa to make it easy for machines to read data on their web sites; scraping becomes superfluous. RDFa is a method of embedding RDF metadata in HTML web pages; RDF is the general data model standardized by the W3C for use on the semantic web, which has been discussed much on this blog.

This doesn't work for many types of data. Brown sees very slow adoption of RDFa and similar technologies but thinks website data will gradually become easier to get at. Most websites are very simple, and their owners see little need or benefit in investing in newer website technologies. If people who really want the data can hire firms like Addtoit to obtain the data, most of the potential benefits to website owners of making their data available accrue without needing technology shifts.

The library industry is slowly freeing itself from the strictures of "library data" and is broadening its data horizons. For example, many libraries have found that genealogical databases are very popular with patrons. But there is a huge world of data out there waiting to be structured and made useful. One of the most interesting dataculture companies to emerge over the last year is ShipIndex. As you'd expect from the name, ShipIndex is a vast directory of information relating to ships. Just as place information is tied together with geoposition data, ShipIndex ties together the world of information by identifying ships and their occurrence in the world's literature. The URIs in ShipIndex are very suitable for linking from other resources.

The Götheborg
ShipIndex is proof that a "family farm" can still deliver value in the dataculture industry. The process used to build ShipIndex. Nonetheless, in coming years you should expect that technologies developed for the financial industry will see broader application and will lead to the creation of data products that you can scarcely imagine.

The business model for ShipIndex includes free access plus a fee-for-premium-access model. One question I have is how effectively libraries will be able leverage the premium data provided with this model. Imagine for example the value you might get from a connection between ShipIndex and a geneological database bound by passenger manifests. I would be able to discover the famous people who rode the same ship that my parents took to and from the US and Sweden (my mom rode the Stockholm on the crossing before it collided with the Andrea Doria). For now though, libraries struggle to leverage the data they have; better data licensing models are way down on the list of priorities for most libraries.

Peter McCracken
ShipIndex was started by Peter and Mike McCracken, who I've known since 2000. Their previous company (SerialsSolutions) and my previous company (Openly Informatics) both had exhibit tables in the "Small Press" section of the American Library Association exhibit hall, where you'll often find the next generation of innovative companies serving the library industry. They'll be back in the Small Press Section at this weekend's ALA Midwinter meeting. Peter has promised to sing a "shanty" (or was that a scupper?) for anyone who signs up for a free trial. You could probably get Mike to do a break dance if you prefer.

I'll be floating around the meeting too. If you find me and say hello, I promise not to sing anything.
Enhanced by Zemanta

Article any source

Tuesday, November 9, 2010

Infochimps and the scaling of dataset value

Image representing Infochimps as depicted in C...Image via CrunchBaseSure, a picture is worth a thousand words, but what is a thousand words worth? How about a million? If I had a dataset of the most recent trillion words spoken by humanity, (anonymized and randomized of course!) would that be worth any more than the set of words in this blog post?

These are real questions. A Texas company called Infochimps has datasets quite similar to these, ready for you to use. Some of the datasets are free, others you have to pay for. More interesting is that if you have a dataset you think other people might be interested in, or even pay for, InfoChimps will host it for you and help you find customers. (Infochimps just announced they had raised $1.2 million in its first round of institutional funding.)

One of the datasets you can get from Infochimps for free is the set of smileys used on twitter in tweets sent between March 2006 and November 2009. It's free. It tells you that the smiley ":)" was used 13,458,831 times, while ";-}" was only used 1,822 times.

If you're willing to fork over $300, you can get a 160MB file conatining a month-by-month summary of all the hashtags, URLs and smiley's used on twitter during the same period. That dataset wil tell you that during September of 2009, the hashtag #kanyeisagayfish was used 11 times while #takekanyeinstead was used 141 times.

If you're a scrabble player, you can spend $4 for a list of the 113,809 official words, with definitions. Or you can get them free, without the definitions.

courtesy of Infochimps, Inc. CC-BY-A
I had a great talk with Infochimps President and Co-Founder Flip Kromer a few weeks ago before his presentation to the New York Data Visualization Meetup. I fell in love with one of the visualizations he showed in his presentation, and he's given me permission to reproduce it here. (Creative Commons Attribution License) It's derived from the same Twitter data set you can get from Infochimps, and shows networks of characters that are found in the same tweet. So if ♠ and ♣ appear in the same tweet over and over again, the two characters will have a strong connection in the network of characters.

The character connection data was fed to a program called Cytoscape, which is an open source visualization program used in bioinformatics; Mike Bergmann has a nice article about its use for large RDF graphs. The networks are laid out using a force-directed algorithm (which is pretty much the simplest thing you can do). Coloring is applied arbitrarily.

As you might expect, the main character networks that show up are associated with languages, but there are some anomalies. For example, the katakana character ツ (tu) sticks out. Katakana is a set of phonetic characters used in Japanese for non-Japanese words. The reason "tu" is set apart from all the other katakana is that people use it on Twitter as a smiley.

The other anomalous character subnet is labeled "???" in the graph. A closer look reveals this to be the set of characters that look like upside down roman text.

Kromer has noticed that the price (or perhaps cost) of a partial data set follows a non-monotonic curve (see graphic). Small amounts of data are essentially free, but a peak value is reached when portions of the data set are extracted from the full data set. If we were discussing book metadata, for example, peak value might accrue for a set of the 100,000 top selling books.

There's much less value, according to Kromer, in having a large incomplete chunk of a data set. Data for 10,000,000 books, for example, would have less value than the 100,000 book data set, because it's not complete. Complete data sets become extremely expensive because of the logistics involved, and because of the value of having the complete set.

This pattern seems plausible to me, but I'd like to see some clearer examples. I've previously written about having too much data, but that article looked at the effect of error rates on data collection; Kromer's curve is about utility.

For me, the most interesting thing about Infochimps is the idea that the best way to make data flow in large volumes and create new types of knowledge is to provide the right incentives for data producers through the establishment of a market. This makes a lot of sense to me; however I'm not sure that the Infochimps market has also established incentives needed for data set maintenance; the world's most valuable and expensive data sets are one that change rapidly.

Kromer contrasted the Infochimps approach to that of Wolfram, whose Alpha service is produced by "putting 100 PhDs and data in a lab". He also feels that much of the work being put into the semantic web is a "crock" because its technology stack solves problems that we don't have. Humans are pretty good at extracting meaning from data, given a good visualization.

We can even recognize upside-down text.
Enhanced by Zemanta

Article any source

Thursday, February 25, 2010

Named Graphs, Argleton and the Truth Economy

Depending on the map provider you're using, there may be a street running through my kitchen. After driving through my kitchen, perhaps you'd like to visit Argleton, town in Lancashire, UK, that only exists on Google Maps. I expect the street through my kitchen is a real mistake, but map companies are known to intentionally insert "trap streets" into their maps to help expose competitors who are just copying their maps.

Errors in information sources can be inadvertant or intentional, but either way, on the internet the errors get copied, propagated and multiplied, resulting in what I call the Information Freedom Corollary:
Information wants to be free, but the truth'll costya.
If you accept the idea that technologies such as Linked Data, web APIs and data spidering are making it much easier to distribute and aggregate data and facts on the internet, you come to the unmistakeable conclusion that it will become harder and harder to make money by selling access to databases. Data of all types will become more plentiful and easy to obtain, and by the laws of supply and demand, the price for data access will drop to near zero. In fact, there are many reasons that making data free increases its value, because of the many benefits of combining data from different sources.


The Attention Economy: Understanding the New Currency of Business
If you want a successful business, it's best to be selling a scarce commodity. Chris Anderson and others have been promoting "free" as a business model for media with the idea that attention is a increasingly scarce commodity (an observation attributed to Nobel prize winning economist Herbert Simon). John Hagel has a good review of discussions about "the Economics of Attention" Whether or not this is true, business models that sell attention are very hard to execute when the product is factual information. Data is more of a fuel than a destination.

The Economics of Attention: Style and Substance in the Age of InformationThere is something that becomes scarce as the volume and velocity of information flow increases, and that's the ability to tell fact from fiction. As data becomes plentiful, verifiable truth becomes scarce.

Let's suppose we want to collect a large quantity of information, and think about the ways that we might practically reconcile conflicting assertions. (We're also assuming that it actually matters to someone that the information is correct!)

One way to resolve conflicting assertions is to evaluate the reputation of the sources. The New York Times has has a pretty good reputation for accuracy, so an assertion by to the Times might be accepted over a conflicting assertion by the Drudge Report. An assertion about the date of an ancestor's death might be accepted if it's in the LDS database, and might be trusted even more strongly if it cites a particular gravestone in a particular cemetary (has provenance information). But reputation is imperfect. I am absolutely, positvely sure that there's no street through my kitchen, but if I try to say that to one of the mapping data companies, why should they believe me in preference to a planning map filed in my town's planning office? What evidence are they likely to accept? Try sending a correction to Google Maps, and see what happens.

Another method to resolve conficts is voting. If two or more independent entities make the same assertion, you can assign higher confidence to that assertion. But as it becomes easier to copy and aggregate data, it becomes harder and harder to tell whether assertions from different sources are really independent, or whether they're just copied from the same source. The more that data gets copied and reaggregated, the more that the truth is obscured.

The semantic web offers another method of resolving conficting assertions, consistency checking. Genealogy offers many excellent examples of how data consistency can be checked against models of reality. A death date needs to be after the birth date of a person, and if someone's mother is younger than 12 or older than 60 at their birth, some data is inconsistent with our model of human fertility. Whatever the topic area, a good ontological model will allow consistency checks of data expressed using the model. But even the best knowledge model will be able to reconcile only a small fraction of conflicts- a birth date listed as 03-02 could be either February or March.

Since none of these methods is a very good solution, I'd like to suggest that many information providers should stop trying to sell access to data, and start thinking of themselves as truth providers.

How does an information provider become a truth provider? A truth provider is a verifier of information. A truth provider will try to give not only the details of Barack Obama's birth, but also a link to the image of his certificate of live birth. Unfortunately, the infrastructure for information verification is poorly developed compared to the infrastructure for data distribution, as exemplified by standards developed for the Semantic Web. Although the existing Semantic Web technology stack is incomplete, it comes closer than any other deployed technology to making "truth provision" a reality.

Although there have been an number of efforts to develop vocabularies for provenance of Linked Data (mostly in the context of scientific data), I view "named graphs" as an essential infrastructure for the provision of truth. Named graphs are beginning to emerge as vital infrastructure for the semantic web, but they have not been standardized (except obliquely by the SPARQL query specification). This means that they might not be preserved when information is transferred between one system and another. Nonetheless, we can start to think about how they might be used to build what we might call the "true" or "verified" semantic web.

On the Semantic Web, named graphs can be used to collect closely related triples. The core architecture of the Semantic Web uses URIs to identify the nouns, verbs, and adjectives; named graphs allow URIs to  identify the sentences and paragraphs of the semantic web. Once we have named graphs, we can build machinery to verify the sentences and paragraphs.

The simplest way to verify named graphs using their URIs is to use the mechanism of the web to return authoritative graph data in response to an http request at the graph URI. Organizations that are serious about being "truth providers" may want to do much more. Some data consumers may need much more extensive verification (and probably updates) of a graph- they may need to know the original source, the provenance, the change history, the context, licensing information, etc. This information might be provided on a subscription basis, allowing the truth provider to invest in data quality, while at the same time allowing the data consumer to reuse, remix, and redistribute the information without restriction, even adding new verification layers.

Consumers of very large quantities of information may need to verify and update information without polling each and every named graph. This might be done using RSS feeds or other publish/subscribe mechanisms. Another possible solution is to embed digital signatures for the graph in the graph URI itself, allowing consumers posessing the appropriate keys to cryptographically distinguish authentic data from counterfeit or "trap street" data.

Named graphs and data verification. I think this is the beginning of a beautiful friendship.
Reblog this post [with Zemanta]

Article any source

Monday, January 25, 2010

8 One-Way Business Models for Linked Data

Real Wheels - Travel Adventures (There Goes a Train/Plane/Bus)One of the videos that I was forced to watch many times when my boys were younger was There Goes a Train. It's a pretty good video. I learned many things, including the fact that locomotives don't have steering wheels. Yeah. Pretty obvious if you think about it for even a moment. It's so unfair. Without the switches, the rail network would be pretty useless, but no one will ever make a video entitled There Sits a Switch.

In electronics though, the switch is the star. With a switch, you can modify and route information; with a wire you can only send it from one point to another. You need good switches to make a computer or a network; even though photons are faster and easier to move from one place to another, computers are still based on electrons because electronic switches are so much better than optical switches.

Linked Data is a label for a set of technologies that are trying to make information move around the internet more easily and with more meaning. The Linked Data vision is one where many entities acting cooperatively and globally create a web of data much more powerful and meaningful than any single entity could bring about.

For the Linked Data vision to become a reality, each entity must have a strong motivation to cooperate; each entity must have a viable business model. If the business models were easy, the Linked Data vision would already be a vibrant reality.

Scott Brinker recently launched a round of discussion about seven business models that can make Linked Data viable. Leigh Dodds contributed some important insights in his followup, prompting Brinker to add an eighth model.

Here are Brinker's eight business models for Linked Data (somewhat relabeled based on who's writing checks):
  1. Subsidy. Entities such as governments with a mandate to make information available will pay to have it linked into a global web of information.
  2. Subscription. People will pay for valuable data, and will pay more for data that has been linked to a global web of information.
  3. Advertising. Advertisers will pay to information in raw data feeds.
  4. Authority. People will pay for the validation and certification of data.
  5. Affiliate marketing. Merchants will pay sales commissions on sales resulting from affiliate links in embedded in the global web of data
  6. Service Enhancement. People will pay for services which have been enhanced by data from a global web.
  7. Search Engine Optimization. Search engines will send you more traffic if you give them more meaningful data.
  8. Brand Enhancement. Your reputation will be burnished if you emit lots of good information.
(I should note that Brinker describes each model a bit differently so that he can add a dimension that characterizes whether data is delivered raw or as an application.  I find that this dimension is not at all orthogonal. A data driven subscription service is a service that makes use of data, but the core business model is not to sell a data subscription.)

There are difficulties with all of these business models, but it strikes me that each of them will only work in one direction, like a train track without switches. Either they work for emitting data, or they work for consuming data, but none of the models work in both directions at the same time. If you're providing a service that's either based on Linked Data or enhanced by it, you can pay for the data, but if you send that data back out, your competitors get the data for free. Conversely, if you're emitting data, it's hard for you to pay for it.

Imagine you're in the book metadata business. You can use several of these models to support creation of book metadata, or you can consume book-related metadata to provide book-related services. But what if you want to support an activity of aggregating book data or fixing errors in book metadata? None of these models will work for you because you'll either be competing with the entities you get data from, or you'll be competing with entities you send data to.

What's missing from this list is a business model for the Linked Data switch. Entities that take in Linked Data, improve it or otherwise add value and reemit it as Linked Data have no solid business model to run on. Everyone active so far in the Linked Data business is either a data sink or a data source. To realize the full potential of Linked Data, there need to be viable switches, both collecting and emitting Linked Data.
Reblog this post [with Zemanta]

Article any source

Saturday, January 2, 2010

Ten Predictions for the Next Ten Years


I didn't do so well in 2000 when I made predictions for the coming year; a year later, I determined that only one of my seven predictions came true.

I'm ten years older and wiser, and I guarantee, triple your money back, that at least 3 of this years predictions will come true. In 2000 I didn't have Twitter to try my first draft on.
  1. The number of public libraries in 2020 will be less than half today's number. Addendum: the number of public library locations will be 50% more in 2020 than today.

    I will write a full post about this, but I believe the driving force for this will be e-books and book digitization, and the result will be consolidation, outsourcing and shuttering of public libraries. Update: I've written a full post.

  2. By the end of 2014, the world's largest aggregation of bibliographic metadata will not be WorldCat. By 2020, no one will care which aggregation is largest.

    Currently, the growth curve for LibraryThing makes it look like it will pass WorldCat in a few years. SerialsSolutions' Summon is definitely in the running. Google can't be discounted. But by the middle of the decade, the size question will seem silly, sort of like "What's the largest computer chip in the word?" or "Who has the most powerful nuclear bomb?" In 2010, we don't care about these questions. In 2020, data quality and currency will be much more important than data completeness. Also, see my article on "When are you collecting too much data?".

    Thanks, @DataG for the comments!

  3. In 2020, general purpose quantum computers will not be useful for any purpose.

    If there's one thing I learned from doing physics, it's there ain't no such thing as a free lunch. If you spend a billion dollars on quantum computing, you might be able to factor an unfactorable integer or two by 2020.

  4. Open Linked Data will hockey-stick in 2012 on standardization of of quad (named graphs?) transport.

    I've been meaning to write more about quad transport, but if you read my article on Pat Hayes' Surfaces, Leigh Dodds' article on Named Graphs, and the DERI proposal on N-quads, you'll know more than I do.

  5. In 2020, the search engine era will be ending. Search engines will give way to less centralized "knowledge fabrics".

    Search engines have a specific topology: spiders pull in data from millions of distributed sites and add it to one big pile that can be searched on. This topology works great if what you want to do is search, but have you ever noticed that Google can't count? Understanding the connections in rapidly changing data will require new topologies and new business models. In 2020, we'll know what they are.

  6. In 2020, China will be seen as having a more modern, sensible, and practical copyright regime than the US.

    In 2010, China has a poor reputation enforcement of Copyright. China will certainly mature in this respect, but to expect it to adopt the regime currently prevailing internationally is to ignore the best interests of China. I think that China will look to the original intent of the US Constitution and invent a copyright regime optimized "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

  7. In 2020, more than half of the book industry's revenue will be facilitated by a Book Rights Registry.

    The Book Rights Registry that would be created by the Google Books Settlement Agreement is too good of an idea to be tied to the settlement agreement. It will happen whether the settlement is approved or not. People will complain about it... all the way to the bank. Note that my prediction uses the indefinite article. There may be more than one book rights registry!

  8. In 2020, the New York Times will be profitable, and will not have gone bankrupt.

    It's easy to predict that the newspaper industry will contract- it's already happening! But the New York Times is uniquely positioned to take advantage of the market gaps that will open when local newspapers fail. Because they do expensive original reporting, they will have little competition. Because they're family-controlled, like Ford, they won't fall victim to the stupidities of the equity markets.

  9. In 2020, Twitter will be a distant memory; Facebook will still be with us.

    Facebook has demonstrated ability to purposefully evolve and extend. Twitter seems not to understand itself. While my neighbor David Carr thinks that Twitter Will Endure, his argument applies to the idea, not the company. Twitter the company will be squeezed between multipurpose networks like Facebook on the high end and non-proprietary protocols on on the low end.

    Thanks, @CodyBrown for the comments!

  10. On January 1, 2020, when I review this list of predictions, I will use a Mac to do it.

    It's been almost 25 years that I've been using a Mac. Do you really think that the mythical Apple tablet of 2020 will not be a Mac?

Reblog this post [with Zemanta]

Article any source

Wednesday, December 2, 2009

Databases are Services, NOT Content

I'm very grateful for advice my fellow entrepreneurs have given me; when you meet someone else who has started a company you have an instant rapport from having shared a common experience. I remember each bit of advice with the same stark clarity that characterizes the moment I realized that Santa Claus was the neighbor dressed up in a white beard and a red suit..

A business owner who I've known since sixth grade gave me this gem: "My secret is providing the best service possible, and charging a lot for it." In executing my own business, I did pretty well at the first part and could have done better at the second part.

As I wrote about the effect of database rights on the postcode economy, I kept wondering if I would have done anything differently in my business if that database protection had been available to me. Would I have charged more for the database that my company developed?

In the comments to that post, I was alerted to a book by James Boyle, called The Public Domain. Chapter 9 in particular parallels many of the arguments I made. One thing I found there was something I had wanted to look for- information about how the database industries in general have done since the "sweat of the brow" theory for copyright was disallowed by the Supreme Court. It turns out that the US database industry has actually outpaced its counterpart in the UK since then by a substantial margin. Why would that be?

I think the answer is that building databases is fundamentally a service business. If your brow is really sweating, and someone is paying you to do it, then it's hard to think of that as a "content" business. Databases always have more content than anyone could ever want; the only reason people pay for them is that they help to solve some sort of problem. If your business thinks it's selling content rather than services, chances are it will focus on the wrong part of the business, and do poorly. In the US, since database companies understand that their competition can legally copy much of their data, they focus on providing high quality added value services, and guess what? THEY MAKE MORE MONEY!

Then there's Linked Data. Given that database provision is fundamentally a service business, is it even possible to make money by providing data as Linked Data? The typical means for prodecting a database service business is to execute license agreements with customers. You make an agreement with your customer about the service you'll provide, how much you'll get paid, and how your customer may use your service. But once your data has been released into a Linked Data Cloud, it can be difficult to assert license conditions on the data you've released.

It's been argued that 'Linked Data' is just the Semantic Web, Rebranded, but it's also been noted both Linked Data is sorely in need of some proper product management. Product management focuses on a customer's problems and how the product can address them. You can believe me because I've not only managed products, I've had 2 whole days of real product management training!

One thing I was taught to do in my Product Management class was to come up with a 1 sentence pitch that captures the essence of the product. When I was an entrepreneur, this was called the elevator pitch. After thinking about it for about 9 months I've come up with a pitch for Linked Data:
Linked Data is the idea that the merger of a database produced by one provider and another database produced by a second provider has value much larger than that of the two separate databases.
or, in a more concise form, V(DB1+DB2)>>V(DB1)+V(DB2).

Based on the products that have been successful this year in the application of semantic web technologies, it looks to me that the most successful have been focused on what I saw Tim Gollins tweet that Ian Davis called "Linked Enterprise Data" (attributing the term to Eric Miller). If the merged databases are contained within the enterprise, the enterprise clearly reaps all the added value. Outside the enterprise, however, the only Linked Open Data winners so far have been the ones who have built services on databases merged from others.

Proper product management would have made it a goal for Linked Open Data to have data contributors share somehow in the surplus value created by the merged services. In the next couple of weeks, I hope to describe some ideas as to how this could happen.
Article any source

Friday, November 13, 2009

The New York Times Gets It Right; Does Linked Data Need a CrossRef or an InfoChimps?

I've been saying this long enough that I don't remember whether I was quoting someone else: whenever the internet disintermediates a middleman, two new intermediaries pop up somewhere else. It's disintermediation whack-a-mole, if you will. The reasons for this are:
  1. The old middlemen became fat on mark-ups an order of magnitude larger than needed by internet-enabled middlemen.
  2. Internet-enabled middlemen add value in ways that the old ones didn't.
My last business functioned as an intermediary that aggregated linking data. We'd get data from publishers, clean it up and add it to our collection, then provide feeds of that data to our customers (libraries and library systems vendors). Our customers got good data and support if was a problem. The companies who provided the data didn't have to deal with hundreds of libraries or system vendors, and they came to understand that we would help their customers link to their content.

Some companies, especially the large ones, were initially uncomfortable with the knowledge that we were selling feeds of data that they were giving out for free. They felt that somehow there was money left on the table. Other companies were fearful of losing control of the information, even though they didn't really have control of it in the first place. Once we explained to them how their data contained mangled character encodings, fictitious identifiers, stray column separators and Catalan month names, they began to see the value we provided.

While my company focused on the data needs of libraries (and did pretty well), a group of the largest academic publishers put up some money and formed a consortium to pool a different type of linking data in a way that let the publishers have more control of the data distribution. This consortium, known as Crossref, just celebrated its 10th anniversary. Crossref has not only paid back the money that its founders invested in it; it has arguably done more to push academic publishing into the 21st century than any other organization on the planet.

As academic publishing companies began to understand the benefits of distributing linking data through Crossref, my company, and others like it, they became more comfortable opening up their content and reaping the financial benefits. Despite the global recession, and despite predictions of its impending collapse, STM publishing has been financially healthy with companies such as Elsevier reporting increased profits. This is rather unlike the newspaper industry, for example.

Before I get to the newspaper industry, I should note yesterday's news that InfoChimps are publishing a collection of token data harvested from Twitter.
Today we are publishing a few items collected from our large scrape of Twitter’s API. The data was collected, cleaned, and packaged over twelve months and contains almost the entire history of Twitter: 35 million users, one billion relationships, and half a billion Tweets, reaching back to March 2006.
InfoChimps is positioning itself as a marketplace to buy, sell, and share data sets of any size, topic or format. Yet another intermediary has popped up!

Two weeks ago, I wrote a somewhat alarmist article about problems in an exciting set of Linked Data being released by the New York Times. I am pleased to be able to be report that the New York Times is now getting it right! The most important thing that they're doing right is that they're listening to the people who want to consume their data. They've started a Google Group based community for the specific purpose of understanding how best to deliver their data. They've also corrected the problems pointed out by myself and others. It's not perfect, but it's not reasonable to expect perfect. The New York Times has set a very hopeful example for other companies that want to start publishing semantic linking information on the open web.

If, as many of us hope, many publishers decide to follow the lead of the Times and make more data collections available, will more intermediaries such as InfoChimps arise to facilitate data distribution, as happened with linking data in scholarly publishing? Will ad hoc groups such as "the Pedantic Web" become key participants in a less centralized data distribution environment? Or maybe large companies will turn off the spigots as "the suits" grow increasingly worried about their ability to control data once it is let out into the web of data.

Perhaps the time is ripe for a set of forward-looking publishers to emulate the nervous-but-smart journal publishers who started Crossref 10 years ago and start a similar consortium for the distribution of Linked Data.
Reblog this post [with Zemanta]

Article any source

Thursday, June 4, 2009

When are you collecting too much data?

Sometimes it can be useful to be ignorant. When I first started a company, more than 11 years ago, I decided that one thing the world needed was a database, or knowledgebase, of how to link to every e-journal in the world, and I set out to do just that. For a brief time, I had convinced an information industry veteran to join me in the new company. One day, as we were walking to a meeting in Manhattan, he turned to me and asked "Eric, are you sure you understand how difficult it is to build and maintain a big database like that?" I thought to myself, how hard could it be? I figured there were about 10,000 e-journals total, and we were up to about 5000 already. I figured that 10,000 records was a tiny database- I could easily do 100,000 records even on my Powerbook 5400. I thought that a team of two or three software developers could do a much better job sucking up and cleaning up data than the so-called "database specialists" typically used by the information industry giants. So I told him "It shouldn't be too hard." But really, I knew it would be hard, I just didn't know WHAT would be hard.

The widespread enthusiasm for Linked Data has reminded me of those initial forays into database building. Some important things have changed since then. Nowadays, a big database has at least 100 million records. Semantic Web software was in its infancy back then; my attempts to use RDF in my database 11 years ago quickly ran into hard bits in the programming, and I ended up abandoning RDF while stealing some of its most useful ideas. One thing that hasn't changed is something I was ignorant of 11 years ago- maintaining a big database is a big, difficult job. And as a recently exed ex-physicist, I should have known better.

The fundamental problem of maintaining a large knowledgebase is known in physics as the second law of thermodynamics, which states that the entropy of the universe always increases. An equivalent formulation is that perpetual motion machines are impossible. In terms that non-ex-physicist librarians and semantic websperts can understand, the second law of thermodynamics says that errors in databases accumulate unless you put a lot of work into them.

This past week, I decided to brush off my calculus and write down some formulas for knowledgebase error accumulation so that I wouldn't forget the lessons I've learned, and so I could make some neat graphs.

Let's imagine that we're making a knowledgebase to cover something with N possible entities. For example, suppose we're making a knowledgebase of books, and we know there are at most one billion possible books to cover. (Make that two billion, a new prefix is now being deployed!) Let's assume that we're collecting n records at random, so for any entity, each record has a 1/N chance of covering any specific entity. At some point, we'll start to get records that duplicate information we've already but into the database. How many records, n, will we need to collect to get 99% coverage? It turns out this is an easy calculus problem, one that even a brain that has spent 3 years as a middle manager can do. The answer is:
Coverage fraction, f = [1- exp(-n/N)]
So to get 99% of a billion records, you'd need to acquire about 4.6 billion records. Of course there are some simplifications in this analysis, but the formula gives you a reasonable feel for the task.

I haven't gotten to the hard part yet. Suppose there are errors in the data records you pull in. Let's call the the fraction of records with errors in them epsilon, or ε. Then we get a new formula for the errorless coverage fraction, F
Errorless coverage fraction, F = [exp(-εn/N) - exp(-n/N)]
This formula behaves very differently from the previous one. Instead of rising asymptotically to one for large n, it rises to a peak, and then drops exponentially to zero for large n. That's right, in the presence of even a small error rate, the more data you pull in, the worse your data gets. There's also a sort of magnification effect on errors- a 1% error rate limits you to a maximum of 95% errorless coverage at best; a 0.1% error rate limits you to 99.0% coverage at best.

I can think of three strategies to avoid the complete dissipation of knowledgebase value caused by accumulation of errors.
  1. stop collecting data once you're close to reaching the maximum. This is the strategy of choice for collections of information that are static, or don't change with time.
  2. spend enough effort detecting and resolving errors to counteract the addition of errors into the collection.
  3. find ways to eliminate errors in new records.
A strategy I would avoid would be to pretend that perpetual motion machines are possible. There ain't no such thing as a free lunch.
Article any source