Showing posts with label Truth. Show all posts
Showing posts with label Truth. Show all posts

Sunday, September 30, 2012

CC BY and the Truth-Printing Business

Why are dollars worth anything? Why are digits on a bank statement worth anything? When my server tells our payments provider to move bits from your credit card, why does it matter to you?

In practical terms, dollars are valuable because other people will give you stuff or do things for you in exchange. Or at least they will if you can convince their bank to change the digits in their bank account. Their bank has to trust your bank which has to trust you. It all works because we all trust it will work. And why do we trust that it will work?

There are governments and laws to back them up. Why do we trust the government and laws? In practical terms we trust the government and laws because... well... they have ballot boxes. And judges and police forces. But mostly we trust the government and legal system because it sort of works and is often not abusive. At the bottom, it's because there's this web of trust which collectively holds everything together. Until of course, it doesn't. Because there isn't a bottom, it's turtles all the way down.

If you haven't heard of Bitcoin, let me give you this non-technical summary. Bitcoin is a recent implementation of the idea that money based on a web of cryptographically secured assertions is sounder than money based on a web of governmentally secured assertions. If as many people believed in cryptography as believe in astrology, we'd be using Bitcoin today.

The magic result is that an entity that gets society to trust its currency can then print money.

When the currency is truth rather than coin, judges and guns don't work so well. Traditional hierarchical authority systems are breaking down. What's replacing them is open authority systems. Systems such as wikipedia which allow everyone to participate in the construction of truth, not by being correct, but by being fixable. And to the frustration of many, Wikipedia delegates all its authority to things that are "citeable".

So how do you get to be an authority that Wikipedia believes? The two criteria that seem to matter most are
  1. Openness. If wikipedians can't read you, you don't exist. 
  2. Authority. People need to believe you. 
If you notice the circularity here, you'll see that printing truth and printing money are not so different.

As usual, I take a long time getting around to my point. Which is this: If you want to be in the business of printing truth, the best license to choose for your business is the Creative Commons Attribution License (CC BY). For now. And if you're printing science, medicine, technology or even philosophy, I really hope you want to print truth.

The Creative Commons part speaks to the need to be open. In the age of the internet, you can't print truth and keep it secret. No one will believe you.

The Attribution part builds your most valuable asset, your reputation. No one believes anonymous assertions.

You might ask about other options, for example, Non-Commercial (NC), No Derivatives(ND), Share-Alike (SA).

I've written about reasons to use NC and ND. Those reasons don't apply to the truth-printing business.

Can you imagine if your dollar bill said "This note is legal tender for all non-commercial debts public or private". That would be silly. The whole point of money is that it doesn't change depending on its use. And its the same with truth. There ain't no such thing as non-commercial truth. You can't control the uses of the truth you print. You can't even demand that people who consume your truth share that truth the same as you do..

A lot of people get confused about using no-derivative licenses. They think that if you print that the sky is blue, your credibility will be hurt if someone reprints a derivative of your truth and says the sky is black. But that's exactly what the attribution requirements prevent. But more than that, if you print your truth as chiseled in stone, then no one will believe it in a few years or so, because we all know that the truth hasn't been chiseled in stone for at least two thousand years. Nowadays we can make cryptographically strong proofs that assertions aren't being fiddled with and were made by the entities they're attributed. We can track the trail of assertions through history. And the provider of that chain of provenance is you, the truth printing proprietor. The longer the trail of conflicting assertions, the more crucial your authority as a truth printer becomes.

The problem of turning the currency of truth into harder currency is left as an exercise for the reader.
Enhanced by Zemanta

Article any source

Saturday, October 15, 2011

How Can We Change the Future? The Tomorrow Project

It turns out that Intel, the giant chip maker, employs a full time futurist. His name is Brian David Johnson, and he actually gets paid to go around asking people what the future might be like. Intel says they're the "Sponsors of Tomorrow", so I guess they want to have a clue about what they're sponsoring. When I worked at Intel in the early 80's, we could have sponsored a thousand futurist studies, and not one of them would have predicted that Intel would someday employ a "Chief Futurist" leading a "Tomorrow Project".

None of those futurists would have predicted that over a hundred thousand people would show up at New York Comic-Con, either. But it's happening. The show is completely sold out. Jacob Javits Convention Center is packed to the gills with zombies, otaku, wood nymphs, transformers and girls with blue, purple or red hair- i don't know the word for them.

Many of them packed a very serious session hosted by Johnson featuring Cory Doctorow, the science fiction writer, blogger, and activist. The session was entitled "Sci-Fi Prototyping: Designing the Future Panel". No on in the audience was disappointed not to hear about the future of the panel, and we also did without Doug Rushkoff, whose appearance was scheduled to make the panel a panel, but who failed to predict his future schedule well enough to participate.

Doctorow, Johnson, Rushkoff, and will.i.am, who was accurately predicted to not be present, have contributed to The Tomorrow Project Anthology which had its launch today. Nostalgically enough, this is a book. Less nostalgic, but perhaps just as dated, it's a 1.8MB PDF file. Made available for free, by Intel. Doctorow's contribution is a novella by Doctorow called The Knights of the Rainbow Table which so far (I'm on p.17), is a fun read. It's about the nano-apocalypse that will occur in the near future when it's easy for a group of grad student low-lifes to crack everybody's website password security.

Johnson framed the session as a discussion about the ways in which science fiction can provide a narrative to steer the future. I'm a bit skeptical. I don't think that the "narrative" of Star Trek communicators caused Motorola engineers to create the flip-phone, even if they were fans of the show while growing up. Doctorow had a really interesting analogy, though. He said a science fiction story was like a Petri dish that lets an microscopic idea grow into a huge colony of micro-organisms visible to the naked eye. That strikes me as a really useful way to think of how fiction influences the world.

The problem with ascribing power to narrative is evident if you look at the world around us. Narratives compete with other narratives, and their relative power derives not from their truth or their skill, but rather from their fit. Narratives warning about the death of privacy, for example, have scant power compared to the offer of a free movie, or even a free PDF download. No one pays attention to a narrative unless it fits with what they want to do today.

Before the panel, Doctorow expressed to me his strong commitment to making his works available with Creative Commons licenses; He'll certainly release The Knights of the Rainbow Table that way. But let's work on Intel to change the future a bit. Why can't they release The Tomorrow Project Anthology with a similar license? (the current license is all rights reserved, you can download it, but you can't redistribute it) You CAN help change the future- file a request to post the whole ebook using this form.
 
In yesterday's tomorrow, androids dream of electric sheep, cars fly around LA, and in Blade Runner, people read newspapers on paper. What will today's tomorrow look like tomorrow? I wonder how much of today's best writing about the future will be available to people ten years from now. Unfortunately, the answer depends on licensing details that most creators don't think much about. Doctorow is an exception.
Enhanced by Zemanta

Article any source

Sunday, September 11, 2011

The Smell of a Book

There's one part of the human brain that seems programmed to never forget things. It's somewhere in the limbic system, and it connects smells to emotions. This past week, deep in the bowels of New York Penn Station, that part or my brain was momentarily triggered by an acrid smell. Perhaps it was smoking train brakes or hot diesel oil, but it evoked a sad memory from ten years ago.

Smells connect us across decades, maybe across millennia. Some smells are hardwired to be pleasant or noxious, other smells are neutral and imprintable. Think of the smell of a new-born baby or the smell of your grandmother. Think of the smell of Starbucks, or of bread baking in the oven at home. Imagine being in a damp cave, or a medieval cathedral.

Scientists have studied this. It's now thought that the primal connection between smell and memory is a result of direct connections between our olfactory lobes and the hippocampus. Some scientists in Israel used functional MRI to see directly the involvement of the hippocampus in memories imprinted with strong smells. (Notes 1 and 2 and the picture.)

It's odd that so many people claim to love the smell of books. It's even stranger that people claim to love the smell of libraries or used bookstores. It's just old glue, ink, dust, mold, and decay. Odd, until you think about the time-travel aspects of smell.

In preparation for the upcoming launch of Unglue.it, I've been talking to a lot of people about the books that they love. "Love" in this context is not the "love" people might use casually to describe their relationship with a product for sale. Instead, people seem to relate to books the way they relate to people. There's the love for a teacher who makes a difference in your life. Love for a friend you helps you feel joy. The thrill of discovering a soul mate. And among authors, there's the blind love for a child that goes beyond all rationality.

The intensity of these emotions must get bound up with smells in the hippocampus to create a lasting impression on book lovers. When we smell a book all of these feelings resonate across time and they comfort us. Even in the future when all our reading is done on ebook readers or other screens, we'll keep real books around us like the clothing of a spouse or a parent lost to a tragedy, left in the bed to warm and comfort. And then we'll find strength to move on, but the spirit of the book will remain.


Notes:
  1. Jonah Lehrer's article on the Israeli fMRI study is very accessible.
  2. That study, "The Privileged Brain Representation of First Olfactory Associations" was written by Yaara Yeshurun, Hadas Lapid, Yadin Dudai and Noam Sobel in Current Biology 19(21), 1869-1874, (9 November 2009) and is available at http://www.cell.com/current-biology/abstract/S0960-9822(09)01857-0
  3. Another human sensation mediated by the hippocampus is laughter. I suffered repeated bouts of this affliction upon reading a website claiming to promote an aerosol spray. I was almost unable to finish this post.

Article any source

Wednesday, June 8, 2011

Our Metadata Overlords and That Microdata Thingy

On June 2, our Metadata Overlords spoke. They told us that they'll only listen when we tell them things using a specialized vocabulary they've now given us at the schema.org website. Although we can still use our stone tablets if that's what we're using now, we're expected to migrate to a new Microdata Thingy, assuming that we really want them to pay attention to our website metadata supplications.

There are among us believers, who, led by druids enraptured by the power of stone tablets to carry truth, will shun the new thingy, but most of us will meekly comply with the edicts of the overlords. We're not able to distinguish the druidic language of the tablets from the new liturgy of of the state church. Many things are difficult to articulate in the new vocabulary, but gosh, those tablets were heavy to carry around. And the new thingy doesn't seem so awful, although it's difficult to tell with the mumbled sermons and hymn singing and all.

I hope the overlords don't try to take our pagan rituals of Friending and Liking away from us, though. The incantations used to invoke and bless the Like ritual also use the druidic language, and the help scrolls tell us we might confuse the overlords if we use more than one language in our prayers.

My soul remains troubled, however, at the thought that the Overlords care not for truth and for justice. Sometimes it seems as though the overlords want only for our offerings of attention and seek only to feed our lust for food, drink, entertainment, debauchery and money. Yes, there are new words for our books and learning, but we can say so little about these in schema.org language that our wizards and mages will be mute if they ever choose to enter that realm.

I myself was present at a conclave of such mages and wizards dedicated to the entwinement of data from libraries, museums and archives in full openness. When tweet of the new order came, we endeavored to learn more of schema.org and its thingy. We questioned whether the thingy was an abomination against openness, or whether we might exploit its Overlord endorsement to make our own spells more powerful. We agreed to teach each other our new thingy spells, even as our colleagues elsewhere figured out how to chisel the new vocabulary into stone. Word came from other lands that the new vessel would founder trying to cross the seas.

We then visited the temple of the archive and found the servers cool to the touch. We heard words from a past oracle, ate as they never ate in Rome, drank cool drafts, and returned home emboldened with an enlarged appreciation of intermingled bits.

So it was said, so shall we do.

Notes:
  1. Google's blog post on adopting microdata was signed by R. V. Guha who had a bit to do with the creation of RDF.
  2. It's not really a surprise that Google doesn't care about RDFa. In my article on RDFa from 2009, I pointed to mistakes that Google made in their RDFa documentation. They never fixed it.
  3. Schema.org can't even list all of its schemata- the web page, chock full of non-breaking spaces, is truncated!.
  4. The current microdata spec is in an odd state where it's confused about how to define an itemtype. In fact, the mechanism for defining new itemtypes is gone! Here's what it says:
    The item type must be a type defined in an applicable specification.

    Except if otherwise specified by that specification, the URL given as the item type should not be automatically dereferenced.

    A specification could define that its item type can be derefenced to provide the user with help information, for example. In fact, vocabulary authors are encouraged to provide useful information at the given URL.
    Apparently, stuff was removed for some sort of political reason- it's there in the WHAT-WG version; note that Google links to the W3C version, which is not fully baked.
  5. the Schema.org terms of service are creepy when you get to the part about patents.
  6. The big selling point for RDFa was that Google, Yahoo and Bing supported it for Rich Snippets and the like. But Microdata's inability to easily support complex markup turned out to be an key feature for the search engines. The moral of the story for standards developers: your best customers are always righter than the others.
  7. In the video, Brewster Kahle reads from the last page of A Manual on Methods of Reproducing Research Material by Robert C. Binkley (1936). OCLC Number 14753642. Peter Binkley, a meeting participant, donated a copy of his grandfather's book to the Internet Archive, along with permission to make it free to the public.
  8. Henri Sivonen has written a very readable and informed discussion about Microdata, RDFa, Schema.org and the process of making standards that you should read if you are interested in why things are the way they are in HTML5.

Article any source

Sunday, August 15, 2010

Charlie Chan Actor Warner Oland Not Mongolian, Say Wikipedia

When my mom was pregnant with her third child, my dad loved it when people asked if they were expecting a boy or a girl. "Well" he'd answer with a twinkle in his eye. "They say one of every 3 children born in the world are Chinese, so for our third child, that's what we're expecting!"

My parents were Swedish. My father was born in Gary, Indiana, but grew up in northern Sweden; my mother was born in Sweden and her mother was a Lapp, or Saami. After his retirement, my father became very interested in genealogy, and he traced his ancestors and relatives, almost 10,000 of them. Since about the year 1400 Sweden has done a very good job of recording births and deaths in church records, and since people didn't move around much, it's not hard for us to trace people. In the farming villages where my parents came from, everybody is related to everybody else.

I've inherited my dad's database and I've put it online. Doing so has has put me in touch with a fascinating variety of distant cousins. Among my distant relatives was the actor Warner Oland, who became famous for portraying Charlie Chan in Hollywood movies. Warner Oland, whose real name was Johan Verner Ölund, was a third cousin to my father's mother. My father noted in his database that he remembered when Warner Oland came to their village by car and met my grandparents. It must have been the same year Warner Oland died, 1938.

Naturally, I pay attention whenever Oland in mentioned in the media. Over the last week, I've read articles in the New Yorker and in the New York Times about a new book by UCSB English Professor Yunte Huang. The book is entitled Charlie Chan: The Untold Story of the Honorable Detective and His Rendezvous with American History; it tells the story of the "real" Charlie Chan, a detective in Honolulu, Hollywood's portrayal of Charlie Chan, and Huang's own story as a chinese immigrant in America. A significant part of the book recounts the odd story of how a Swedish actor came to portray the quintessential Chinese detective.

When I read the New Yorker article, I immediately put the book on my "must read" list. (Unfortunately, it's not available as an ebook, and is sold out of my local bookstores!) But one sentence of the New Yorker review, written by Harvard history professor Jill Lapore, stuck out for me:
Oland, born in Sweden in 1880, had, beginning in 1917, specialized in playing Oriental villains, including Dr. Fu Manchu. (Oland's mother was Russian, and he had slavic features.)
Oland's mother was NOT Russian. Oland's mother was my grandmother's third cousin. His father was a 5th cousin to my grandmother. The Swedish genealogist Sven-Erik Johansson has specialized in the digitization of the church records in the region of northern Sweden where Oland and my grandparents came from and has published an ancestor chart for Warner Oland going back 5 generations. None of those ancestors come from Russia. To top it off, Warner Oland was born in 1879, not 1880 as reported in the New Yorker.

So where did the idea that Oland had a Russian mother come from? Doesn't the New Yorker have fact checkers? I went to Wikipedia to find out. The Wikipedia article said that "His mother was Russian of Mongolian descent.", referencing a "page not found" Internet Movie Database (IMDB) article. I refound that article, which says:
He didn't need make-up when he played Charlie Chan; all he would do is curl down his moustache and curl up his eyebrows. In fact, the Chinese often mistook him for one of their own countrymen. He attributed this to the fact that his Russian grandmother was of Mongolian descent.
So IMDB says it's his grandmother who's Russian and of "Mongolian descent"; the key thing to note is the attribution. I immediately edited the Wikipedia article to omit to spurious information. A day later, a wikipedian had put back the Mongolian bit, but more accurately worded as being something Oland said. A proper reference, to a book by Ken Hanke, Charlie Chan at the Movies: History, Filmography, and Criticism (Google Books, Amazon) had been added. That book says:
"Even before the role of Charlie Chan came his way, Oland was a frequent onscreen Oriental, despite the fact that he was born in Sweden to a mixture of Swedish and Russian Parents. Physically, he had an exotic look to begin with, and the addition of an Oriental-style mustache and beard made the transformation complete. "I owe my Chinese appearance to the Mongol invasion," he once told Keye Luke. "That's true," Luke agrees, "because the Mongols did get up there around Sweden and Finland and naturally sired some children, and so, he said, 'I come by it naturally.' And, his whole family looked like that." There was never any need for elaborate make-up. "All he did," explains Luke. "was put that little goatee on his chin. Otherwise, he had his own mustache. Everything was just like that. No make-up. It's just amazing."
At this point, I must make an observation. Please look at the photo and decide for yourself. As far as I can judge, Warner Oland didn't look the least bit Oriental. He looked like most everyone else living in that area of northern Sweden would look if they put on a smudge of eyebrow makeup. But the resemblance to that Chinese detective in the movies is uncanny!

There is, however, a story I remember my dad telling about a deserter from the Russian army. (The Russians burned down the closest city, Umeå, in a war in 1720.) It was said that this deserter hid in the woods or disguised himself as one of the locals. The way my dad told it, it was quite a scandal, even 200 years later. So maybe Warner Oland was joking when he said his mother was "Russian". In any case, the mysterious Russian in my family does not appear in the church records!

If Oland really had exotic features, it's much more likely he got them from a source other than a stray Mongolian. The closest the Mongols got to northern Sweden was Lithuania. In the area where Oland's family originated, the ethnic mix was dominated by Finns, Swedes, and Saami.

Take a look at a photo of my mother's cousin (unrelated to Oland), a pure Saami. With a bit of make-up (and some acting talent), she would have easily been able to play a Chinese woman. The Saami look quite different from the Finns and the Swedes. They are an indigenous people of Scandinavia, and no one really knows where they came from. Though their language is related to Finnish, they are not genetically related to the Finns. A recent DNA study (PDF, 399KB) published in the American Journal of Human Genetics suggests that they are related to the Berbers of northern Africa. It may well be that Oland thought they might be related to the Mongols.

There's another interpretation of Oland's references to his "Mongolian" blood. In his time, children with Down's syndrome were referred to as "Mongoloid". In the 19th century, Down's syndrome was regarded as an expression of genetic "degeneration" toward the inferior "Mongoloid" races. It could well be that jokes about Mongolian ancestry reflected a belief that cases of Down's Syndrome were a result of racial contamination. My father's database shows many examples of women with large families bearing children into their 40's; his own familiy of 11 included one Down's child.

So it seems likely that Warner Oland's statements about his ancestry were either inventions or jests. What's interesting to me is how this truth is constructed. It's not hard for people to look at the evidence now available and decide that a genealogist working with church records is probably more reliable than a co-star's recollection of an actor's constructed persona with regard to Oland's ancestry. Yunte Huang, the author of the new book, emailed me to say he agreed that it was a jest of Oland, who was known to be "quite a wisecracker". Now THAT sounds like my Dad's family!

At first glance you might say Wikipedia is totally unreliable, because anyone can change it. But compared to the New Yorker, IMDB, and a book published in 2004, Wikipedia is more reliable because it CAN be changed, and because it supports a version history and a culture of citation and transparency for any information that might be disputed. While I'm optimistic about Wikipedia's ability to construct truth, I'm worried about systems that extract facts from Wikipedia articles and feed them in to the semantic web. While editing the article on Warner Oland, I deleted the assertion that he was a "Swedish Person Of Russian Descent". I wonder about the lifespan of this assertion as it has been copied and distributed throughout the world. There are really no good mechanisms to de-sert this sort of assertion. It's only with context that assertions can build truth.

As it happens, I married into a family that really IS Chinese. I remember showing my mother-in-law old pictures of Saami ancestors in their traditional dress. "Those look like Manchu people!" she exclaimed. It's true. If you put aside the lens of race, we all look more or less alike, and we all look a bit exotic.

Update: The author of the New Yorker article, Jill Lapore, got back to me to report that her article relied on the entry for Warner Oland in American National Biography (Oxford:  Oxford University Press, 2000) for the assertion about Oland's mother's ancestry.

Update, August 23: Some additional research shows that Oland is also a third cousin of my grandmother.
Enhanced by Zemanta

Article any source

Thursday, February 25, 2010

Named Graphs, Argleton and the Truth Economy

Depending on the map provider you're using, there may be a street running through my kitchen. After driving through my kitchen, perhaps you'd like to visit Argleton, town in Lancashire, UK, that only exists on Google Maps. I expect the street through my kitchen is a real mistake, but map companies are known to intentionally insert "trap streets" into their maps to help expose competitors who are just copying their maps.

Errors in information sources can be inadvertant or intentional, but either way, on the internet the errors get copied, propagated and multiplied, resulting in what I call the Information Freedom Corollary:
Information wants to be free, but the truth'll costya.
If you accept the idea that technologies such as Linked Data, web APIs and data spidering are making it much easier to distribute and aggregate data and facts on the internet, you come to the unmistakeable conclusion that it will become harder and harder to make money by selling access to databases. Data of all types will become more plentiful and easy to obtain, and by the laws of supply and demand, the price for data access will drop to near zero. In fact, there are many reasons that making data free increases its value, because of the many benefits of combining data from different sources.


The Attention Economy: Understanding the New Currency of Business
If you want a successful business, it's best to be selling a scarce commodity. Chris Anderson and others have been promoting "free" as a business model for media with the idea that attention is a increasingly scarce commodity (an observation attributed to Nobel prize winning economist Herbert Simon). John Hagel has a good review of discussions about "the Economics of Attention" Whether or not this is true, business models that sell attention are very hard to execute when the product is factual information. Data is more of a fuel than a destination.

The Economics of Attention: Style and Substance in the Age of InformationThere is something that becomes scarce as the volume and velocity of information flow increases, and that's the ability to tell fact from fiction. As data becomes plentiful, verifiable truth becomes scarce.

Let's suppose we want to collect a large quantity of information, and think about the ways that we might practically reconcile conflicting assertions. (We're also assuming that it actually matters to someone that the information is correct!)

One way to resolve conflicting assertions is to evaluate the reputation of the sources. The New York Times has has a pretty good reputation for accuracy, so an assertion by to the Times might be accepted over a conflicting assertion by the Drudge Report. An assertion about the date of an ancestor's death might be accepted if it's in the LDS database, and might be trusted even more strongly if it cites a particular gravestone in a particular cemetary (has provenance information). But reputation is imperfect. I am absolutely, positvely sure that there's no street through my kitchen, but if I try to say that to one of the mapping data companies, why should they believe me in preference to a planning map filed in my town's planning office? What evidence are they likely to accept? Try sending a correction to Google Maps, and see what happens.

Another method to resolve conficts is voting. If two or more independent entities make the same assertion, you can assign higher confidence to that assertion. But as it becomes easier to copy and aggregate data, it becomes harder and harder to tell whether assertions from different sources are really independent, or whether they're just copied from the same source. The more that data gets copied and reaggregated, the more that the truth is obscured.

The semantic web offers another method of resolving conficting assertions, consistency checking. Genealogy offers many excellent examples of how data consistency can be checked against models of reality. A death date needs to be after the birth date of a person, and if someone's mother is younger than 12 or older than 60 at their birth, some data is inconsistent with our model of human fertility. Whatever the topic area, a good ontological model will allow consistency checks of data expressed using the model. But even the best knowledge model will be able to reconcile only a small fraction of conflicts- a birth date listed as 03-02 could be either February or March.

Since none of these methods is a very good solution, I'd like to suggest that many information providers should stop trying to sell access to data, and start thinking of themselves as truth providers.

How does an information provider become a truth provider? A truth provider is a verifier of information. A truth provider will try to give not only the details of Barack Obama's birth, but also a link to the image of his certificate of live birth. Unfortunately, the infrastructure for information verification is poorly developed compared to the infrastructure for data distribution, as exemplified by standards developed for the Semantic Web. Although the existing Semantic Web technology stack is incomplete, it comes closer than any other deployed technology to making "truth provision" a reality.

Although there have been an number of efforts to develop vocabularies for provenance of Linked Data (mostly in the context of scientific data), I view "named graphs" as an essential infrastructure for the provision of truth. Named graphs are beginning to emerge as vital infrastructure for the semantic web, but they have not been standardized (except obliquely by the SPARQL query specification). This means that they might not be preserved when information is transferred between one system and another. Nonetheless, we can start to think about how they might be used to build what we might call the "true" or "verified" semantic web.

On the Semantic Web, named graphs can be used to collect closely related triples. The core architecture of the Semantic Web uses URIs to identify the nouns, verbs, and adjectives; named graphs allow URIs to  identify the sentences and paragraphs of the semantic web. Once we have named graphs, we can build machinery to verify the sentences and paragraphs.

The simplest way to verify named graphs using their URIs is to use the mechanism of the web to return authoritative graph data in response to an http request at the graph URI. Organizations that are serious about being "truth providers" may want to do much more. Some data consumers may need much more extensive verification (and probably updates) of a graph- they may need to know the original source, the provenance, the change history, the context, licensing information, etc. This information might be provided on a subscription basis, allowing the truth provider to invest in data quality, while at the same time allowing the data consumer to reuse, remix, and redistribute the information without restriction, even adding new verification layers.

Consumers of very large quantities of information may need to verify and update information without polling each and every named graph. This might be done using RSS feeds or other publish/subscribe mechanisms. Another possible solution is to embed digital signatures for the graph in the graph URI itself, allowing consumers posessing the appropriate keys to cryptographically distinguish authentic data from counterfeit or "trap street" data.

Named graphs and data verification. I think this is the beginning of a beautiful friendship.
Reblog this post [with Zemanta]

Article any source

Thursday, November 5, 2009

The Blank Node Bother and the RDF Copymess

There were many comments on my post about the problems in the Linked Data released by the New York Times, including some back and forth by Kingsley Idehen, Glenn MacDonald, Cory Casanave and Tim Berners-Lee that many readers of this blog may have found to be somewhat inexplicable. On the surface, the comments appeared to be about how to deal with the potentially toxic scope of "owl:sameAs". At a deeper level, the comments surround the issue of how to deal with a limitation of RDF. A better understanding of this issue will also help you understand difficulties faced by the New York Times and other enterprises trying to benefit from the publication of Linked Data.

Let's suppose that you have a dataset that you want to publish for the world to use. You've put a lot of work into it, and you want the world to know who made the data. This can benefit you by enhancing your reputation, but you might also benefit from others who can enhance the data, either by adding to it or by making corrections. You also may want people to be able to verify the status of facts that you've published. You need a way to attach information about the data's source to the data. Almost any legitimate business model that might support the production and maintenance of datasets depends on having some way to connect data with its source.

One way to publish a dataset is to do as the New York Times did, publish it as Linked Data. Unfortunately, RDF, the data model underlying Linked Data and the Semantic Web, has no built-in mechanism to attach data to its source. To some extent, this is a deliberate choice in the design of the model, and also a deep one. True facts can't really have sources, so a knowledge representation system that includes connections of facts to their sources is, in a way, polluted. Instead, RDF takes the point of view that statements are asserted, and if you want to deal with assertions and how they are asserted in a clean logic system, the assertions should be reified.

I have previously ranted about the problems with reification, but it's important to understand that the technological systems that have grown up around the Semantic Web don't actually do reification. Instead, these systems group triples into graphs and keep track of data sets using graph identifiers. Because these identified graphs are not part of the RDF model they tend to be implemented differently from system to system and thus the portability of statements made about the graph as a whole, such as those that connect data to their source, is limited.

At last week's International Semantic Web Conference Pat Hayes gave an invited talk about how to deal with this problem. I've discussed Pat's work previously, and in my opinion, he is able to communicate a deeper understanding of RDF and its implications than anyone else in the world. In his talk (I wasn't there, but his presentation is available.) he argues that when an RDF graph is moved about on the Web, it loses its self-consistency.

To see the problem, ask yourself this: "If I start with one fact, and copy it, how many facts do I have?" The answer is one fact. "one plus one equals two" is a single fact no matter how many times you copy it! You can think of this as a consequence of the universality of the concepts labeled by the english words "one" and "two".

I haven't gotten to the problem yet. As Pat Hayes points out, the problem is most clearly exposed by blank nodes. Blank nodes are parts of a knowledge representation that don't have global identity; they're put in as a kind of glue that connects parts of a fact. For example, lets suppose that we're representing a fact that's a part of the day's semantic web numerical puzzle: "number x plus number y equals two". "number x" and "number y" are labels we're assigning to a number that semantic web puzzle solvers around the world might attempt to map to a univeral concept. Now suppose I copy this fact into another puzzle. How many facts do I have? This time, the answer is two, because "number x" might turn out to be a different number in the second puzzle. So what happens if I copy a graph with a blank node a hundred times? Do the blank nodes multiply while the universally identified node don't? Nobody knows!

I hope you can see that making copies of knowledge elements and moving them to different contexts is much trickier than you would have imagined. To be able to manage it properly you need more than just the RDF model. In his talk, Pat Hayes proposes something he calls "Blogic" which adds the concept of "surfaces" to provide the context for a knowledge representation graph. If we had RDF surfaces, or something like that, then the connections between data and its source would be much easier to express and maintain across the web. Similarly, it would be possible to limit the scope of potentially toxic but useful assertions such as "owl:sameAs".

There are of course other ways to go about "fixing up" RDF, but I'm guessing the main problem is a lack of enthusiasm from W3C for the project. The view of Kingsley Idehen and Tim Berners-Lee appears to be that existing machinery, perhaps bolstered by graph IDs or document IDs is good enough and that we should just get on with putting data onto the web. I'm not sure, but there may be a bit of "information just wants to be free" ideology behind that viewpoint. There may be a feeling that information should be disconnected from its source to avoid entanglements, particularly of the legal variety. My belief is a bit different- it's that knowledge just wants to be worth something. And that providing solid context for data is ultimately what gives it the most value.

P.S. Ironically, in the very first comment on my last post, Ed Summers hints at a very elegant way that the Times could have avoided a big part of the problem- they could have used entailed attribution. It's probably worth another post just to explain it.

Reblog this post [with Zemanta]

Article any source

Sunday, October 18, 2009

My Optimized Baseball Media Diet and Why Motoko Rich Can't Count

40 years ago I started following the Philadelphia Phillies. I think that it started the month that my family rented a beach house on Long Beach Island. Every morning I would walk to the store to buy a newspaper- the Philadephia Inquirer- because I wanted to read everything about Apollo 11. After the astronauts got home safely I continued my morning newspaper ritual, and that's when I started reading about the baseball.

Between then and now, there were some years when it was hard to follow my team, and I don't mean because they were bad. When I lived in California, the local newspapers barely covered my team, even though it was a National League city. I would study the boxscores and the three sentences in the AP summaries to retain an emotional connection to my team. That's when I first imagined a newspaper of the future that could be customized and printed for me so that I could have the New York Times front page along with an Inquirer sports page with my morning coffee.

Then the internet happened, and all of a sudden I could track the Phillies games on Yahoo Sports and read articles on the Inquirer's web site, Philly.com, even though I was living in Mets and Yankees-land. I barely read the sports section of my local paper any more. With cable television, I could watch games whenever the Phils played Atlanta or the Mets. My baseball media diet had reverted to what it was growing up, except now I got it over wires instead of the airwaves and on paper.

Over the past three years, however, my baseball media diet has changed profoundly, and not just because the Phillies won the World Series. This season, I was able to watch most games on my iPhone or on my laptop via MLB.com. Every day I read the blog of the best sports writer covering the Phillies, Jason Weitzel. I get breaking news via Twitter from Scott Lauber, a writer for some paper in Wilmington, Delaware. I read game summaries from Todd Zolecki and other writers who work for MLB.com. I read news about Phillies prospects at PhuturePhillies.com, and I read stat-head analysis (partly enabled by huge volumes of game transactional data released by major league baseball) at the Hardball Times.

In my optimized Phillies media diet, there's not much role for traditional media, or even for transitional media aggregators like Yahoo or Philly.com. The media providors I've ended up with have all specialized in areas of strength. I don't have to endure sports writers I don't like just because they've managed to gain special access to the flow of information.

The same sort of change is happening all over the landscape of news reporting. Two weeks ago, I had a chance to see first hand how the professional media reported a rather minor event in a story I'd been following quite closely. I went to a federal courtroom in New York and witnessed a meeting of a judge and the parties of a lawsuit involving Google, copyright and ebooks.

At the end of my report, I added links to other reports published about the same event. It's interesting to read these reports, and think about how they fit into an optimized media diet. The most knowledgable report was by James Grimmelmann, a professor at NYU law school. Those of us who have followed the lawsuit closely have come to rely on Prof. Grimmelmann's blog for insight into the relevant law. The best written coverage, in my opinion, was that of Motoko Rich of the New York Times. She condensed the event down to its bare essence, and chose exactly the right story lines. At the event, I watched her in action. After its conclusion, she made a beeline for a publishing executive, sitting two seats away from me, and asked him exactly the right question.

But it seems Motoko Rich made a small mistake. If you compare her well-written story with my notes-dump, you'll note a tiny discrepancy. She reports that there were "fewer than 70 people" in the courtroom. I was amazed to see so many people, and it was my very first time in a Federal courtroom, so I decided to make a careful count. There were four rows of benches, filled with 12 people each. There were 8 members of the press seated in the jury box. There were 8 attorneys for the parties and the Department of Justice at the lawyers tables. Seated along the back wall were 16 people, eight on each side. So not including the Judge and his staff or the courtroom official and security, there were 80 people in the courtroom.

Where did Motoko Rich's "fewer than 70" number come from? Perhaps she meant to write "more than 70". Perhaps an editor or fact-checker could believe that so many people could fit in the courtroom. I don't know. I left a comment on the Times' website, but for whatever reason it was not approved. Perhaps the correction was considered so trivial that it was better to leave the mistake in the story. In fact, version of the story put the number at "approximately 70 people", and the print version omitted any reference to the audience size.

This episode got me thinking about the proper role of professional reporters in my media diet. I don't expect a reporter to have the expertise of a Law Professor, but I really want people like Motoko Rich to be asking the right people piercing questions. Although I can go to the same event that she can, it's just not my job to badger people with questions, even if I do happen to know them. But having been accustomed to the accountability of sports reporting that has to stand up to hundreds of reader comments, I would really like to see similar accountability in the news reporting I read. It should matter more, not less.

I'm also worried about the business models that support my news sources. I hope that Jason Weitzel is making enough from his blog to support himself- he's probably made only a few dollars from me (I bought his book last year). I'm glad Scott Lauber is supported by his newspaper, but it has close to zero revenue from me. Major League Baseball is getting significant revenue from me- I hope they're smart enough to add to the media that they support.

Whith the whole news industry experiencing the wholesale rearrangement of roles that has already happened for me for baseball, what is a reporter to do? Should she focus on developing contacts, asking questions and crafting stories, or should she focus more on building a reader contituency? Should a "newspaper" business focus on aggregating news or nurturing reporters? Should it be building a information access platform, or should it be developing a community news resource? Maybe it should be contributing to the cloud of linked data.

I don't have answers for these questions, but I can tell you why you should trust me to count courtroom spectators more accurately than Motoko Rich. I'm taller than she is. I can see better over people's heads. And somehow we should figure out a way for Motoko Rich's physical stature to not be relevant to her stature as a reporter.
Reblog this post [with Zemanta]

Article any source