Showing posts with label Newspaper industry. Show all posts
Showing posts with label Newspaper industry. Show all posts

Saturday, January 2, 2010

Ten Predictions for the Next Ten Years


I didn't do so well in 2000 when I made predictions for the coming year; a year later, I determined that only one of my seven predictions came true.

I'm ten years older and wiser, and I guarantee, triple your money back, that at least 3 of this years predictions will come true. In 2000 I didn't have Twitter to try my first draft on.
  1. The number of public libraries in 2020 will be less than half today's number. Addendum: the number of public library locations will be 50% more in 2020 than today.

    I will write a full post about this, but I believe the driving force for this will be e-books and book digitization, and the result will be consolidation, outsourcing and shuttering of public libraries. Update: I've written a full post.

  2. By the end of 2014, the world's largest aggregation of bibliographic metadata will not be WorldCat. By 2020, no one will care which aggregation is largest.

    Currently, the growth curve for LibraryThing makes it look like it will pass WorldCat in a few years. SerialsSolutions' Summon is definitely in the running. Google can't be discounted. But by the middle of the decade, the size question will seem silly, sort of like "What's the largest computer chip in the word?" or "Who has the most powerful nuclear bomb?" In 2010, we don't care about these questions. In 2020, data quality and currency will be much more important than data completeness. Also, see my article on "When are you collecting too much data?".

    Thanks, @DataG for the comments!

  3. In 2020, general purpose quantum computers will not be useful for any purpose.

    If there's one thing I learned from doing physics, it's there ain't no such thing as a free lunch. If you spend a billion dollars on quantum computing, you might be able to factor an unfactorable integer or two by 2020.

  4. Open Linked Data will hockey-stick in 2012 on standardization of of quad (named graphs?) transport.

    I've been meaning to write more about quad transport, but if you read my article on Pat Hayes' Surfaces, Leigh Dodds' article on Named Graphs, and the DERI proposal on N-quads, you'll know more than I do.

  5. In 2020, the search engine era will be ending. Search engines will give way to less centralized "knowledge fabrics".

    Search engines have a specific topology: spiders pull in data from millions of distributed sites and add it to one big pile that can be searched on. This topology works great if what you want to do is search, but have you ever noticed that Google can't count? Understanding the connections in rapidly changing data will require new topologies and new business models. In 2020, we'll know what they are.

  6. In 2020, China will be seen as having a more modern, sensible, and practical copyright regime than the US.

    In 2010, China has a poor reputation enforcement of Copyright. China will certainly mature in this respect, but to expect it to adopt the regime currently prevailing internationally is to ignore the best interests of China. I think that China will look to the original intent of the US Constitution and invent a copyright regime optimized "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."

  7. In 2020, more than half of the book industry's revenue will be facilitated by a Book Rights Registry.

    The Book Rights Registry that would be created by the Google Books Settlement Agreement is too good of an idea to be tied to the settlement agreement. It will happen whether the settlement is approved or not. People will complain about it... all the way to the bank. Note that my prediction uses the indefinite article. There may be more than one book rights registry!

  8. In 2020, the New York Times will be profitable, and will not have gone bankrupt.

    It's easy to predict that the newspaper industry will contract- it's already happening! But the New York Times is uniquely positioned to take advantage of the market gaps that will open when local newspapers fail. Because they do expensive original reporting, they will have little competition. Because they're family-controlled, like Ford, they won't fall victim to the stupidities of the equity markets.

  9. In 2020, Twitter will be a distant memory; Facebook will still be with us.

    Facebook has demonstrated ability to purposefully evolve and extend. Twitter seems not to understand itself. While my neighbor David Carr thinks that Twitter Will Endure, his argument applies to the idea, not the company. Twitter the company will be squeezed between multipurpose networks like Facebook on the high end and non-proprietary protocols on on the low end.

    Thanks, @CodyBrown for the comments!

  10. On January 1, 2020, when I review this list of predictions, I will use a Mac to do it.

    It's been almost 25 years that I've been using a Mac. Do you really think that the mythical Apple tablet of 2020 will not be a Mac?

Reblog this post [with Zemanta]

Article any source

Sunday, October 18, 2009

My Optimized Baseball Media Diet and Why Motoko Rich Can't Count

40 years ago I started following the Philadelphia Phillies. I think that it started the month that my family rented a beach house on Long Beach Island. Every morning I would walk to the store to buy a newspaper- the Philadephia Inquirer- because I wanted to read everything about Apollo 11. After the astronauts got home safely I continued my morning newspaper ritual, and that's when I started reading about the baseball.

Between then and now, there were some years when it was hard to follow my team, and I don't mean because they were bad. When I lived in California, the local newspapers barely covered my team, even though it was a National League city. I would study the boxscores and the three sentences in the AP summaries to retain an emotional connection to my team. That's when I first imagined a newspaper of the future that could be customized and printed for me so that I could have the New York Times front page along with an Inquirer sports page with my morning coffee.

Then the internet happened, and all of a sudden I could track the Phillies games on Yahoo Sports and read articles on the Inquirer's web site, Philly.com, even though I was living in Mets and Yankees-land. I barely read the sports section of my local paper any more. With cable television, I could watch games whenever the Phils played Atlanta or the Mets. My baseball media diet had reverted to what it was growing up, except now I got it over wires instead of the airwaves and on paper.

Over the past three years, however, my baseball media diet has changed profoundly, and not just because the Phillies won the World Series. This season, I was able to watch most games on my iPhone or on my laptop via MLB.com. Every day I read the blog of the best sports writer covering the Phillies, Jason Weitzel. I get breaking news via Twitter from Scott Lauber, a writer for some paper in Wilmington, Delaware. I read game summaries from Todd Zolecki and other writers who work for MLB.com. I read news about Phillies prospects at PhuturePhillies.com, and I read stat-head analysis (partly enabled by huge volumes of game transactional data released by major league baseball) at the Hardball Times.

In my optimized Phillies media diet, there's not much role for traditional media, or even for transitional media aggregators like Yahoo or Philly.com. The media providors I've ended up with have all specialized in areas of strength. I don't have to endure sports writers I don't like just because they've managed to gain special access to the flow of information.

The same sort of change is happening all over the landscape of news reporting. Two weeks ago, I had a chance to see first hand how the professional media reported a rather minor event in a story I'd been following quite closely. I went to a federal courtroom in New York and witnessed a meeting of a judge and the parties of a lawsuit involving Google, copyright and ebooks.

At the end of my report, I added links to other reports published about the same event. It's interesting to read these reports, and think about how they fit into an optimized media diet. The most knowledgable report was by James Grimmelmann, a professor at NYU law school. Those of us who have followed the lawsuit closely have come to rely on Prof. Grimmelmann's blog for insight into the relevant law. The best written coverage, in my opinion, was that of Motoko Rich of the New York Times. She condensed the event down to its bare essence, and chose exactly the right story lines. At the event, I watched her in action. After its conclusion, she made a beeline for a publishing executive, sitting two seats away from me, and asked him exactly the right question.

But it seems Motoko Rich made a small mistake. If you compare her well-written story with my notes-dump, you'll note a tiny discrepancy. She reports that there were "fewer than 70 people" in the courtroom. I was amazed to see so many people, and it was my very first time in a Federal courtroom, so I decided to make a careful count. There were four rows of benches, filled with 12 people each. There were 8 members of the press seated in the jury box. There were 8 attorneys for the parties and the Department of Justice at the lawyers tables. Seated along the back wall were 16 people, eight on each side. So not including the Judge and his staff or the courtroom official and security, there were 80 people in the courtroom.

Where did Motoko Rich's "fewer than 70" number come from? Perhaps she meant to write "more than 70". Perhaps an editor or fact-checker could believe that so many people could fit in the courtroom. I don't know. I left a comment on the Times' website, but for whatever reason it was not approved. Perhaps the correction was considered so trivial that it was better to leave the mistake in the story. In fact, version of the story put the number at "approximately 70 people", and the print version omitted any reference to the audience size.

This episode got me thinking about the proper role of professional reporters in my media diet. I don't expect a reporter to have the expertise of a Law Professor, but I really want people like Motoko Rich to be asking the right people piercing questions. Although I can go to the same event that she can, it's just not my job to badger people with questions, even if I do happen to know them. But having been accustomed to the accountability of sports reporting that has to stand up to hundreds of reader comments, I would really like to see similar accountability in the news reporting I read. It should matter more, not less.

I'm also worried about the business models that support my news sources. I hope that Jason Weitzel is making enough from his blog to support himself- he's probably made only a few dollars from me (I bought his book last year). I'm glad Scott Lauber is supported by his newspaper, but it has close to zero revenue from me. Major League Baseball is getting significant revenue from me- I hope they're smart enough to add to the media that they support.

Whith the whole news industry experiencing the wholesale rearrangement of roles that has already happened for me for baseball, what is a reporter to do? Should she focus on developing contacts, asking questions and crafting stories, or should she focus more on building a reader contituency? Should a "newspaper" business focus on aggregating news or nurturing reporters? Should it be building a information access platform, or should it be developing a community news resource? Maybe it should be contributing to the cloud of linked data.

I don't have answers for these questions, but I can tell you why you should trust me to count courtroom spectators more accurately than Motoko Rich. I'm taller than she is. I can see better over people's heads. And somehow we should figure out a way for Motoko Rich's physical stature to not be relevant to her stature as a reporter.
Reblog this post [with Zemanta]

Article any source

Thursday, July 16, 2009

The New York Times is NOT Being Disrupted by Innovation

Clayton Christensen coined the phrase "disruptive innovation" to describe a recurring pattern of incumbent technology companies being unable to maintain their market leadership through a particular type of technology transition. If you have not read the book, or watched one of his lectures, you should take two minutes right now to watch his video, or else stop reading this post NOW.

It really bugs me when people who have not read the book or have not taken the time to understand Christensen's insights steal the phrase "disruptive innovation" or "disruptive technology" and plaster it onto something that doesn't fit Christensen's model. For example, one characteristic of disruptive technology is that incumbent companies fail to adopt a new technology because it doesn't meet the needs of the market, i.e. their existing customers. This characteristic gets twisted by some entrepreneurs and technologists so that a technology's failure to address customers' needs (or to have customers in the first place) is cited as evidence of the technology's disruptive nature!

Another common misunderstanding of "disruptive innovation" is to assume that a technology is disruptive just because it poses a threat to an incumbent technology. Here's an easy way to tell if a "threatening" technology is a good fit to the disruptive innovation model: ask yourself "is the new technology a threat because it delivers higher performance, with a hope that its cost will be driven down to challenge current technology? Or is the technology a threat because it's really cheap, and has a hope to increase performance to be able to challenge current technology?" The high-performance technology is what Christensen labels a "sustaining technology"; the low-cost technology is what Christensen label a "disruptive technology".

In a previous post on whether scientific publishing is about to be disrupted, I argued that the problems of newspaper industry were not germane to the future of the scholarly publishing industry. In this post, I want to examine whether the newspaper industry fits the Christensenian model of incumbents facing disruptive innovation. Michael Nielsen's article argues in favor of disruption, suggesting that a blogs like Techcrunch, by adopting low-cost technical infrastructure, are disruptive innovators. I agree that the low cost infrastructure fits the disruptive model- there are no printing companies that have attempted to develop blogging infrastructure, for example. But that doesn't make Techcrunch a disruptive innovator, or newspapers a disrupted industry. The reason is that both Techcrunch and newspapers are really in the business of selling advertising. The advertising that Techcrunch sells is actually at the high-performance, highly targeted, expensive end of the market compared to the advertising that the New York Times sells.


In Christensen's model, incumbent companies abandon low-margin market segments to the disruptors because they want to focus on the most profitable parts of their business. But this is the opposite of what has happened in newspapers. Real estate listings and other classified ads have huge margins. Internet sites such as Zillow and Craiglist exploited these huge margins to make businesses out of delivery of high-performing ads.

I find it much more useful to think of the newspaper industry not as one being disrupted by innovation, but rather as one being fragmented by innovation. The internet allows information services to be profitable at much smaller sizes than previously possible. The result is that many markets previously served by newspapers became vulnerable to competition from smaller, more focused services.

I can think of a number of industries afflicted by fragmentation, and the outlook for incumbent companies is not nearly so dire as for industries afflicted by disruption. The television broadcasting and semiconductor industries are good current examples. Although many companies fail to adapt to a fragmented market and disappear, many survive and remain vital. There are a number of strategies for survival- the "roll-up", the "smaller but focused company", and of course the "climb up the food chain" and "move down the food chain" strategies. There are also strategies for failure, most prominently, the "pretend nothing's wrong" strategy.

The bottom line here is that I think there's hope for companies in the newspaper industry. Unless the New York Times shrinks its typeface and crossword puzzle so loyal readers like me can't read it anymore, it might not go bust.


Article any source

Friday, July 10, 2009

Spherical Livestock and the Alleged Disruption of Scientific Publishing

Physicists have a joke about "spherical cow approximations" referring to their tendency to simplify a problem to make calculations easier, even though such simplifications bring into question the solution's application to reality. My favorite version of the joke, which I first heard directly from Hans Bethe, has Nikita Khrushchev asking his most elite scientists to help the Soviet Union with its difficulty meeting its five year plan for the dairy industry. The biologists and the chemists are completely stumped by the problems of increasing milk production, but the physicists proudly announce they have solved the milk production problem, but only for the case of spherical cows.

In a post entitled "Is scientific publishing about to be disrupted?", quantum information theorist Michael Nielsen describes what he thinks is a general explanation for why businesses and industries fail, and goes on to draw an analogy between the newspaper industry and the scientific publishing industry. Although the post is well written and highly entertaining, (I find his discussion of "immune systems" particularly delicious) I find part of his analysis to be even worse than a spherical cow approximation- he's trying to study milk production by analyzing the spherical chicken! Let me explain.

Nielsens "spherical chicken" is illustrated in this graph from his blog:

In the graph, he plots some sort of measure of success versus some sort of configuration parameter that presumably could be tuned to turn the New York Times into TechCrunch, or vice versa. He goes on to say that
The problem is that your newspaper has an organizational architecture which is, to use the physicists’ phrase, a local optimum. Relatively small changes to that architecture - like firing your photographers - don’t make your situation better, they make it worse. So you’re stuck gazing over at TechCrunch, who is at an even better local optimum, a local optimum that could not have existed twenty years ago
The problem with this analysis is that TechCrunch is completely immaterial to the difficulties that the newspaper industry is undergoing. The financial health of the New York Times and the newspaper industry is not being undermined by news blogs, it's being undermined by non-news sites such as Craigslist, Zillow, and the internet as a whole. Craigslist has focused on classified ads, and only classified ads, and unburdened by the expense of producing the rest of a newspaper, it is able to provide a much more effective solution for the classified advertiser. Zillow has done the same thing in the real estate advertising category. Another big revenue source for newspapers is display advertising to consumers. But nowadays, when someone wants to buy something or find a service, their first thought is to go directly to the internet. Want to find when a movie is playing? You used to pull out a newspaper, now you go to the internet. A company like BestBuy used to communicate with customers through newspaper ads; while they still do so to some extent, the internet allows them to communicate directly with consumers through their web site. None of the newspapers' real competitors are in the news business at all, and there is no configuration parameter of any sort that could be tuned to transform the New York Times into Craigslist.

The news industry's core problem is not, as Nielsen suggests, their inability to adopt disruptive technologies, but rather the disintegration of the linkage between their main activity and their revenue streams. In the past, good news would attract readership, and readership would attract advertisers. The biggest difficulty for newspapers today is not so much the loss of readership, it's that advertisers now have many more ways to connect to that readership. In applying the lessons of the newspaper industry to the evolution of the scientific publishing industry, it's the stability of activity-revenue linkage that needs to be closely examined.

Even a cursory look at the scholarly publishing industry reveals a very different situation from that of the newspaper industry. First of all, there is much more business-model diversity in scholarly publishing. There are huge companies like Elsevier competing with cottage companies which produce a single journal. There are large non-profit societies such as the American Physical Society that produce extremely cost effective journals and who make much of their content available for free. There are journals that have long survived primarily on advertising and journals that have long survived primarily on society member dues. There is also a lot of experimentation with business models going on, including author-paid open access publishers, and mixed "open choice" business models. This business model diversity gives scientific publishing industry robustness against the prospect of any one business model being severely disrupted. In addition, the transition to digital delivery which is giving the newspaper industry such difficulty is to a significant extent already being accomplished in the journal publishing industry.

The scientific publishing industry does have a similar activity-revenue linkage problem that it needs to pay attention to. The people who write the biggest checks to scientific publishers are institutional libraries. But scientific journals, for the most part, do not cater to libraries, they cater to author communities, because the biggest determinant of a scientific journal's success has been the quality and quantity of articles it is able to attract. As long as libraries continue to be attracted to the authorship attracted by journals, and continue to attract the institutional funding they need to support their subscription, the biggest revenue stream for scientific publishers will be secure. But suppose that institutions start deciding to outsource their libraries or begin to require researchers to directly fund their journal subscriptions? Or suppose that libraries are successful in attracting authors directly into open-access institutional repositories?

A better analogy from physics for the scholarly publishing business might be the polaron. A polaron is the combination of a particle and interactions with the environment that it moves in, and the combination has a mass significantly larger that the "bare" particle moving on its own. In the case of the scientific publishing business, the interactions with its environment include the way tenure committees rely on the prestige of a journal that has published a candidates work, or the way accreditation boards require libraries to subscribe to certain numbers of journals. The polaronic industry thus gains mass and inertia, allowing it continue longer than it might otherwise do. Computer operating systems work in the same way- they induce the creation of third party software that interact with the operating system and thus increase its mass and inertia in the market.

Strongly interacting polarons can distort their environments so much that the become trapped by their cloud of interactions- think of a celebrity trying to walk though a crowd of fans. For a business this can be a fatal situation if objectives change, and there is no possibility to adapt.

How's that for a spherical cow?


Article any source

Friday, June 26, 2009

Why the Times took 8 days to Announce its Linked Data Announcement

It took 8 full days for the New York Times to make the same announcement on its "Open" blog that it made last week at the Semantic Technology Conference. Being that it's Friday afternoon, I present here my purely hypothetical speculations on what took so long, based on reading of tea leaves and semantic hyperparsing of the subtle, almost hidden differences between today's text and a transcription of the announcement of last Wednesday.
  1. A pitched battle between entrenched factions within the New York Times has waged over the past week, pitting a radical cabal of openists versus the incumbent "we've always done it that way" faction. The openists slipped the announcement of the announcement into their blog while the traditionalists were occupied with the battle over the type size of the headline for the Michael Jackson story today.
  2. The TimesOpen team missed last week's deadline for the "Sunday Styles" Announcements section.
  3. Normally, announcements like these take two weeks to process, but the business section was starting to get worried that USAToday was going to scoop them with a front pager on Monday.
  4. The written announcement was held up because a patent lawyer feared that the admission that the Thesaurus was "almost 100 years" old could hurt the Times' efforts to obtain a patent on the semantic web.
  5. The announcement was actually made last Thursday, but the printf() command in the Blog's subtitles crashed some key RSS syndication agents.
  6. The fact checker was on vacation.
If you have ever worked in an organization of even moderate size, you know that the real reason is almost certainly banal and boring.

On a more serious note, I think it's important to understand how organizations (not just the New York Times) adapt their internal processes to enable semantic technology in general. Over the past 15 years, the necessity to produce a web site has required many organizations to overhaul many of their internal processes, resulting in new efficiencies and capabilities that go well beyond the production of a website. At last weeks Semantic Technology Conference, there were a number of presentations that solved problem X using semantic technologies, raising immediate questions about what was so wrong with solving problem X the conventional way. Implicit in the presentations was an assumption that by approaching problems using semantic techniques, one could achieve a level of interoperability and software reuse that is not being achieved with current approaches. That's a sales pitch that's been made for many other technologies. What is certainly true is that many problems that are causing pain these days can only be solved by reengineering of corporate processes; maybe semantic technologies will be a catalyst for this re-engineering, at least in the publishing industry.

A very thoughful review of last weeks conference has been posted by Kurt Cagle. I leave you with this quote from Kurt:
There comes a point in most programmers careers where they make a startling realization. Computer programming has nothing to do with mathematics, and everything to do, ultimately, with language. It’s a sobering thought.
A reassuring thought as well.
Article any source

Monday, June 22, 2009

The New York Times and the Infrastructure of Meaning

The big announcement at last week's Semantic Technology Conference came from the New York Times. Rob Larson and Evan Sandhaus announced that the New York Times would be releasing its entire thesaurus as Linked Data sometime soon (maybe this year). I've been very interested in looking at business models that might support the creation of maintenance of Linked Data, so I've spent some time thinking about what the New York Times is planning to do. Rob's presentation spent a lot of time evoking the history of the New York Times, and tried to make the case that the Adolph Ochs' decision in 1913 to publish an index to the New York Times played a large part in the paper's rise to national prominence as one of the nation's "Newspapers of Record". The decade of that decision was marked by an extremely competitive environment for New York newspapers- the NYT competed with a large number and variety of other newspapers. I don't know enough about the period to know if that's a stretch or not, but I rather suspect that the publication of the index was a consequence of a market strategy that proved to be successful rather than the driver of that strategy. The presentation suggested a correspondence between the decade of the 1910's and our current era of mortal challenges to the newspaper business. The announcement about linked data was thus couched as a potentially pivotal moment in the paper's history- by moving decisively to open its data to the semantic web, the New York Times would be sealing its destiny as a cultural institution integral to our society's infrastructure of meaning.

The actual announcement, on the other hand, was surprisingly vague and quite cautious. It seems that the Times has not decided on the format or the license to be used for the data, and it's not clear exactly what data they are planning to release. Rob Larson talks about the releasing the "thesaurus", and about releasing "tags". These are not the terms that would be used in the semantic web community or in the library community. A look at the "TimesTags API" documentation gives a much clearer picture of what Rob means. Currently, this API gives access to the 27,000 or so tags that power the "Times Topics" pages. Included as "tags" in this set are
  • 3,000 description terms
  • 1,500 geographic name terms
  • 7,500 organization name terms
  • 15,000 person name terms
The Times will release as linked data "hundreds of thousands" of tags dating back to 1980, then in a second stage will release hundreds of thousands more tags that go back to 1851. They want to community to help normalize their tags, and connect them to other taxonomies. According to Larson, "the results of this effort, will in time, take the shape of the Times entering (the linked) data cloud." I presume this to meant that the Times will create identifiers for entities such as persons, places, organizations, and subjects, and make these entities available for others to use. Watch the announcement for yourself:


I've found that it's extremely useful to think of "business models" in terms of the simple question "who is going to write the checks?" The traditional business model for newspapers has been for local advertisers and subscribers to write the checks. Advertisers want to write checks because newspapers deliver localized aggregates of readers attracted by convenient presentations of local and national news together with features such as comics, puzzles, columns and gossip. Subscribers write checks because the the paper is a physical object that provides benefits of access and convenience to the purchaser. Both income streams are driven by a readership that finds reading the newspaper to be an important prerequisite to full participation in society. What Adolph Ochs recognized when he bought control of the Times in 1896 was that there was an educated readership that could be attracted and retained by a newspaper that tried to live up to the motto "All the news that's fit to print". What Ochs didn't try to do was to change the business model.

The trials of the newspaper industry are well known, and the business model of the New York Times is being attacked on all fronts. Newspapers have lost their classified advertising business because Craigslist and the like serve that need better and cheaper. Real estate advertising has been lost to Zillow and the online Multiple Listing Service. The New York Times has done a great job of building up its digital revenue, but the bottom line is that hard news reporting is not as effective an advertising venue as other services such as search engines. Subscribers, on the other side, are justifiably unwilling to pay money for the digital product, because the erection of toll barriers makes the product less convenient rather than more convenient. Nonetheless, the digital version of the New York Times retains the power to inform its readership, a power that advertisers will continue to be willing to pay for. It's also plausible that the New York times will be able to provide digital services that some subscribers will be willing to pay for. So, assuming they don't go bankrupt, the business model for the future New York Times does not look qualitatively different from the current model (at least to me), even if the numbers are shifting perilously in the near future.

So let's examine the stated rationales for the New York Times to join the Linked Data community, and how they might help to get someone to send them some checks. The first and safest stated rationale is that by entering the linked data cloud, traffic to the New York Times website will increase, thus making the New York Times more attractive to advertisers. So here's what puzzles me. What Rob Larson said was that they were going to release the thesaurus. What he didn't say was that they were also going to release the index, e.g. the occurrence of the tags in the articles. Releasing the index together with the thesaurus could have a huge beneficial impact on traffic, but releasing the thesaurus by itself will leave a significant bottleneck on the traffic increase, because developers would still have to use an API to get access to the actual article uri's. More likely, most developers who want to access article links would try to use more generic api such as those you'd get from Google. Why? If you're a developer, not so many people will write you checks for code that only works with one newspaper.

I would think that publication of occurrence coding would be a big win for the NYT. If you have articles that refer to a hundred thousand different people, and you want people interested in any of those people to visit your website, it's a lot more efficient for everyone involved (and a lot less risk of "giving away the store") for you to publish occurrence coding for all of these people than it would be for everyone who might want to make a link to that article to try to do indexing of the articles. The technology behind Linked Data, with its emphasis on dereferencable URI's, is an excellent match to business models that want to drive traffic via publication of occurrence coding.

Let's look at the potential costs of releasing the index. Given that the Times needs to produce all of the occurrence data for its website, the extra cost of releasing the linked data for the index should be insignificant. The main costs of publishing occurrence data as Linked Data are the risks to the Times' business model. By publishing the data for free, the Times would cannibalize revenue or prevent itself from being able to sell services (such as the index) that can be derived from the data, and in this day and age, the Times needs to hold onto every revenue stream that it can. However, I think that trying to shift the Times business model towards data services (i.e. selling access to the index) would a huge risk and unlikely to generate enough revenue to sustain the entire operation. Another serious risk is that a competitor might be able to make use of the occurrence data to provide an alternate presentation of the Times that would prove to be more compelling than what the Times is doing. My feeling is that this is already happening to a great extent- I personally access Times articles most frequently from my My Yahoo page.

The other implied rationale for releasing data is that by having its taxonomy become part of Linked Data infrastructure, the New York Times will become the information "provider of record" in the digital world the way the index helped it become one of the nation's "newspapers of record". The likelihood of this happening seems a bit more mixed to me. Having a Times-blessed set of entities for people, places and organizations seems useful, but in these areas, the Times would be competing with more open, and thus more useful, sets of entities such as those from dbpedia. For the Times to leverage its authority to drive adoption of its entities, it would have to link authoritative facts to its entities. However, deficiencies in the technology underlying linked data make it difficult for asserted facts to retain the authority of the entities that assert them. Consider a news article that reports the death of a figure of note. The Times could include in the coding for that article an assertion of a death date property for the entity corresponding to that person. It's complicated (i.e. it requires reification) to ensure that a link back to the article stays attached to the assertion of death date. More likely, the asserted death date will evaporate into the Linked Data cloud, forgetting where it came from.

It will be interesting to see how skillful the Times will be in exploiting participation in linked data to bolster its business model. I'll certainly be reading the Times' "Open" blog, and I hope, for the Times' sake that the go ahead and release occurrence data along with the thesaurus. The caution of Rob Larson's announcement suggests to me that the Times is a bit fearful of what may happen. Still, it's one small step for a gray lady. One giant leap for grayladykind?
Reblog this post [with Zemanta]

Article any source