Showing posts with label Piracy. Show all posts
Showing posts with label Piracy. Show all posts

Tuesday, January 3, 2012

Foreign Libraries Will Be Infringing Sites Under SOPA


On January 1, "Public Domain Day", another year was added to the copyright gap between the US and the rest of the world. In most of the world, New Year's Day marked the end of copyright for works by authors who died in 1941. But not in the USA. Copying and distribution of gap works may be legal and unrestricted in most countries, but these activities are criminal acts of copyright infringement in the US, punishable by up to 10 years of prison.

One effect of SOPA, the "Stop Online Piracy Act" (which almost means "garbage" in Swedish) is that the US Attorney General will be able to extend the effect of US copyright law to foreign web sites. For example, Project Gutenberg Australia (PGA) distributes electronic versions of F. Scott Fitzgerald's The Great Gatsby, which is still under copyright in the US. SOPA would allow the US Attorney General to make a determination that Project Gutenberg Australia, by allowing access from the US, is a "U.S. directed site" that would be subject to forfeiture by the Attorney General for acts prohibited under section 2319 of the U.S. Criminal Code (ongoing copyright infringement), if it were based in the US.

If Project Gutenberg Australia persisted in its criminal activity, SOPA would allow the Attorney General to force internet service providers in the US to block access to the PGA domain. It would also allow the Justice Department to force Google and other US-based search engines to remove PGA links from its search results for the US. It could force Wikipedia to remove links to PGA. Most damaging to PG Australia, it would allow Justice to cut off PGA's revenue from Google's advertising services.

As libraries around the world move aggressively into the digital environment, they will de-emphasize the cataloguing of printed objects in favor of delivery of electronic content, especially public-domain and public-commons content that can be delivered without per-copy fees. They will run up against the same copyright extraterritorial issues exhibited today by Project Gutenberg Australia. If SOPA is enacted as currently written and our system of perpetual copyright persists, we can assume that most of the world libraries will sooner or later be blacklisted from the American internet. And the real bad-guy sites will easily circumvent the blacklist.

The Swedish word "sopa" is not really used as a singular noun. As a verb, it means "to sweep". The plural "sopor" is the stuff you sweep up. "Sopa bort" is what should be done with SOPA.
Enhanced by Zemanta

Article any source

Saturday, December 31, 2011

2011: The Year the eBook Wars Broke Out

Open war is upon us, whether we would have it or not. These incidents in 2011  seemed like twitter-inflamed kerfuffles as we lived through them, but with the perspective of time, we can see they were preludes to a fight to the death.

1. Harper-Collins and Overdrive Stop Pretending

In a year or two, libraries may consider the Harper Collins limit of 26 circulations of a list price ebook through Overdrive to be a relative bargain, as all of the other large publishers will withdraw from "pretend-its-print" ebook licensing.

2. Amazon occupies Overdrive

Libraries mostly welcomed the possibility to lend their Overdrive ebooks to patrons with Kindles. Libraries are fundamentally service-oriented institutions and ebooks on Kindle is what the users wanted. But at what cost? Do the traditional library values of privacy go right out the door? Do libraries realize that patrons gone to Amazon might not come back?

3. The Penguin Strikes Back

The big publishers have watched Amazon's market power grow and see a future of slavery to an internet commerce master. Only Penguin allowed hostilities to break out, however, as the Amazon occupation of Overdrive broke the penguin's back. The target of opportunity was library lending. Evidently Penguin decided that a frontal assault on Amazon would be suicidal.

4. Prime Pretends to be a Library

Amazon added ebook borrowing features to their Amazon Prime service, revealing it as Amazon's answer to Netflix, and without even thinking about it, as a service that could eventually compete directly with public libraries. Now we see why Amazon wanted to get in on that library thing.

5. Publishers Decide Google is a Lesser Evil

Publishers looked back on the halcyon days when Google Books seemed poised to establish a new world order for ebooks with nostalgia. A separate, anticlimactic settlement between Google and the Association of American Publishers appears to be in the offing. It's Amazon that they're afraid of now.

6. Authors Lob Legal Grenades at Hathitrust

Spurned by the publishers in their joint crusade against the Google heathens, the Authors Guild decided that Hathitrust might be a less formidable opponent. And indeed it was, the lawsuit exposed a number of copyright blunders by the library cooperative. But the Guild's suit seemed hasty and ill-contrived. This sort of thing happens in wartime.

7. Amazon Obliterates Borders.

Although Borders was tactically weak in many ways, it was Amazon and the rise of ebooks that killed it strategically. Barnes and Noble, if it survives, won't look anything like the book marketing machine that it is today.

8. Libraries Muster the Resistance

The emergence of the Digital Public Library of America (DPLA) as a rallying point for libraries' continuing presence in the cultural life of America was a surprise, as it went against the prevailing tea-party currents for smaller government and increased reliance on the private sector. It's not clear how the symbolic presence of a library in Zuccotti Park could point the way to a digital future, but many things that have not yet come to pass are shrouded in darkness.

9. Anti-Piracy Hysteria Threatens Freedom Loving Citizens

The powerful publishing and media industries, in a paroxysm of inept do-something-ism, seem to have convinced Congress that it would be a good thing if the intenet could be censored for copyright infringement. Sadly, the solution they've fixed on, SOPA, will be ineffective against unlicensed content and will put the Justice Department smack in the middle of our nation's information infrastructure. Carpet bombing never ends well.

There's hope.

I have learned that whenever it seems that you're falling into the abyss, you must reach for a rope. There is always a rope.
Enhanced by Zemanta

Article any source

Friday, December 16, 2011

How to Dig for Book Data Treasure


To me, surest indicator of an impending doom for book publishing is hearing a publisher cite the advertising of Attributor, an anti-piracy solutions company, as if it were science. It's not the attitude towards piracy that bothers me, that's entirely sensible. It's the implied devaluation of honest data that depresses me.

There's hope though. I've gotten to know quite a number of people throughout the reading ecosystem with whom I can use the word "data" as high praise, roughly equivalent to the word "gold". If you're reading this, chances are you're a member of this secret society, and what follows is a sketch of a treasure map.

In a recent post, I promised to suggest ways that we might measure the effects of library ebook lending on book sales. If you think about it, there are many parallels between attempting such a measurement and previous studies that have tried to measure the effect of ebook piracy on book sales. Unfortunately, the only objective study I know of was a small study done by Brian O'Leary, and the effects observed in that study were small and in a direction counter to popular narratives (and thus rarely noted in the sort of presentations that cite Attributor advertising).

In that study, O'Leary looked for time-domain correlations between sales figures for books from two publishers and the appearance of the same books on BitTorrent. A similar study focused on library Lending could be much more compelling, because library circulation data is a much more direct measure of distribution than any sort of torrent tracking, and librarians are much better than pirates at sharing data.

With the cooperation of booksellers, library circulation and holdings could be compared and correlated to store-by-store sales. For example, you could look at a book that's held in a significant fraction of libraries and look for correlations (positive AND negative) between areas where a library is circulating the book and stores where the book is selling. You've have to remove regional and demographic variance, of course, but with enough data, almost anything is possible.

With the cooperation of a large publisher, rigorous experiments could be done. Scientific experiments derive rigor from the use of controls. To prove that lending influences sales, it's not enough to do lending and look for sales. A rigorous experiment would have both a trial where books are lent and an identical trial where the same books are not lent.

One way to control a lending experiment would be to make a random selection of a publisher's catalog available for lending. Imagine if Penguin had worked with the library community on an experimental withholding of a random part of its catalog from Overdrive. The sales could be analyzed for patterns and trends.

It's important that data analysis of this sort be done objectively by researchers with integrity. In any large collection of data, it's possible to focus on data which supports one narrative over another. If lending-sales studies were done, my guess is that some types of books would show correlations very different from others.

I've used the word "cooperation" several times already. I'm not so naïve as to think that data sharing will materialize out of thin air. Perhaps the sort of eco-system wide organization envisaged by the same Brian O'Leary could be the vehicle to make data treasure digging possible. Opportunity in Abundance for the win!

Enhanced by Zemanta

Article any source

Monday, December 12, 2011

SOPA Could Put Common Library Software in the Soup


The "Stop Online Piracy Act", or SOPA, is promoted as something that will... stop online piracy. So I was a bit surprised when I learned how it's supposed to work. A key provision of SOPA will shut down "notorious" websites by setting up a national web filter based on domain names. I'm sure the pirates had a great laugh about that one. They'll be the ones benefiting while the rest of us figure out how to avoid collateral damage. Members of Congress should consult the nearest available 14-year-old on the ease of web filter evasion: school teachers in my town routinely access their filter-blocked Facebook accounts by asking students to show them how it's done.

Rerouting domain names to alternate IP addresses is pretty easy to do, and can be very useful as well. One type of software used to accomplish this is called a "proxy server". It's called that because it acts as your web browser's proxy in requesting files from a web site. For example, after connecting to a proxy server in Stockholm, my requests for web pages would appear to issue from a computer in Sweden instead of from my computer in New Jersey.

Libraries often use proxy servers to simplify IP authentication of their networks to digital information providers. When an academic library buys access to a database, for example, they'll give the IP address of their proxy-server to the database provider, which then puts the IP address on an "allow" list. Then everyone at the school accesses the database through the address of the proxy server. In effect, those proxy-authenticated users circumvent the IP address-based filter that blocks unauthorized users.

Passage of SOPA would inevitably spawn the creation of a network of proxy servers hosted in countries that reject filtering of the internet. Users in the US could then connect transparently to  blocked sites by connecting through a constantly shifting network of proxy servers. The key to that connection would be a Proxy Auto-config, or PAC file- essentially a mini DNS file installed in the user's web browser software.

SOPA contains provisions that allow the US Attorney General to
bring an action for injunctive relief against any entity that knowingly and willfully provides or offers to provide a product or service designed or marketed for the circumvention or bypassing of [domain name blocking] and taken in response to a court order issued pursuant to this subsection, to enjoin such entity from interfering with the order by continuing to provide or offer to provide such product or service.
Proxy servers meet the condition of being designed to route around filters and therefore fall into the category of services that could be subject to injunctive action under SOPA. The proxy servers most frequently used in libraries are OCLC's EZProxy and the open-source software known as SQUID, but there are many others in use.

In particular, SQUID makes use of PAC files, and thus could be vulnerable if the Justice Department decides that PAC files make it too easy to evade SOPA blockages. Conceivably, the Justice department could force browser developers to omit support for PAC files, or perhaps to restrict their transmission.

Similar concerns about important software have been raised by Jim Fruchterman on behalf of Benetech, a non-profit that among other things, provides ebooks to the reading disabled. Benetech is also one of the largest developers of software for human rights activists around the world. They operate TOR servers designed to foster anonymous communications. On Beneblog, Fruchterman worries that Benetech services could be impacted by SOPA. In response, a commenter signing in as "Copyright Alliance" argues that such action would be unlikely because "The State Department is strongly committed to advancing both Internet freedom and the protection and enforcement of intellectual property rights on the Internet." Too bad it's the Justice Department that gets to decide which services constitute circumvention.

I don't think that libraries will have their proxy servers taken away anytime soon, even if SOPA is enacted. But it's likely that the widespread development of SOPA-circumventing infrastructure would degrade the ability of rights holders to find and prosecute copyright violators. Knowledge of the actual locations of unauthorized files would by hidden offshore in distributed proxy servers, completely out of the reach of US law enforcement. The "file lockers" of today would dissolve into ungraspable bit vapors, and the online piracy problem would just get worse and worse.

There are many ways to address the online piracy problem- too many to list in this post. My own company is working on a piracy-neutering business model for ebooks. I don't know enough to evaluate the possible effectiveness of the payment and advertising network components of SOPA. But it appears to me that from the technical point of view, the internet filter component of SOPA will be a charm of powerful trouble, like a hell-broth, boil and bubble.

Notes:
  1. @amac has a good post on SOPA's scope issues, as well as links to other articles.
  2. I focus here on SOPA, but there are similar issues with PROTECT IP, as described by Steve Crocker and 4 other prominent internet engineers.
  3. The Crocker paper describes a number of other ways that domain name filtering might be circumvented. These include using replacing .hosts files on the user's computer (similar to PAC file installation) and switching the user to using a non-filtered DNS server. Apparently this is done transparently by some types of computer malware. This can only end badly.


Enhanced by Zemanta

Article any source

Wednesday, November 17, 2010

Real Research Gets Reproduced

It's not often that I'm identified as a physicist, as Richard Curtis did in his commentary on my followups of Attributor's piracy "demand" report. But it's true, I worked in materials physics research at Bell Labs from 1988-1998.

Crystal structure of YBCO
Those were great years to be in materials physics. In 1986, two guys at IBM Zurich discovered some amazing new superconducting materials. By the end of that year, a team in Japan had reproduced their results; a group I was a part of at Stanford did the same after talking with the Japan group in December. By March, so many groups around the world had made exciting discoveries that the American Physical Society meeting in New York became known as "the Woodstock of Physics".

A blue semiconductor laser.
A few years later, a guy in Japan reported that he had made a semiconductor light emitting diode (LED) glow blue. His work was a lot harder to reproduce; it took years for anyone to come close to what his team reported; although he published many details, it was hard work. I even sawed one of his LEDs in half to try to understand how it worked. Today, my kitchen (and the screen of my MacBook) is lit by white LED's made from that same semiconductor.

Around that time, some chemists in Utah announced a truly amazing discovery: they saw fusion reactions occurring in palladium electrochemical cells. Since they were respected electrochemists, their results were taken seriously, and lots of people tried to reproduce the incredible results. The promise of a seemingly magical, unlimited power source seemed almost too good to be true. This time, however, nobody could reproduce the results. Some scientists saw odd things happen, but they were different in every lab. At Bell Labs, the scientists trying to reproduce so-called "cold fusion" became convinced that the guys in Utah were being led astray by their excitement.

In science, it's usual that a surprising result will only be accepted once it has been reproduced by someone else. My scientific training has sometimes gotten me in trouble in the world of libraries and publishing. When presented with something that seems surprising to me, I ask for the evidence. In cultures that are more comfortable assigning and recognizing authority, my questions have sometimes been seen as irritants.

It's been that way with my questions about the Attributor report. I was surprised at some of the findings, and I tried to reproduce them. My results can't reproduce some of the key findings reported by Attributor. It would be nice to better understand the factor of a hundred difference between my results and those of Attributor; much might be learned from such an analysis. Attributor is a company that sells anti-piracy services; one would hope that the data they report is somehow rooted in fact, even though they benefit from overestimates of privacy.

In Richard Curtis' article, Jim Pitkow, Attributor's CEO, is quoted:
Our study’s rigorous methodology ensured highly accurate results that align with actual consumer behavior. We analyzed 89 titles, using multiple keyword permutations per title, across different days of the week, with very high bids to ensure placement – each of which is fundamental in guaranteeing accuracy and legitimacy. Each of these variables impact the findings, and analyzing all variables together produce highly accurate results. We stand by our research, and we’re confident that the study addresses an accurate portrayal of the consumer demand for pirated e-books.
If Attributor really stands by its research, it will make it easier for people like me to reproduce their results. In particular, they should publish the complete list of the "869 effective keyword terms" used as keywords for their Google AdWords experiment. There are mistakes they might have made in permuting and combining search terms; they might also have thought of a class of effective search terms that my study totally overlooked. As it stands, it's impossible to know.

I can understand why Attributor might not want to release their search term list. First of all, they should expect people to try to tear it to shreds. The marketing department isn't going to like that. That's what happened to the superconductor guys, the blue LED guy, and cold fusion guys. They stood behind their work, and let the scientific community look for weaknesses and make their own judgments.

Cold fusion didn't pan out, and Pons and Fleischmann, the Utah guys, tried for years to figure out what it was they measured. Bednorz and Müller, the guys in Zurich, won the Nobel Prize. Shuji Nakamura, the LED guy, won a Millenium Prize and a lawsuit.

It may be easier to do a followup study without the worry of spurious searches for widely known terms. But at this point, Attributor customers and the book industry as a whole stand to learn a lot from understanding where the irreproducibility of Attributor's study is coming from. Publishers need that information to plan out a response to the threat of ebook piracy, and their needs should come first- no matter what the marketing department says.
Enhanced by Zemanta

Article any source

Friday, November 5, 2010

Consumer Interest in Pirated eBooks is Even Lower Than I Thought

My recent posts following up on Attributor's most recent study on demand for pirated ebooks have been republished on TeleRead, probably the longest running blog covering ebooks and related topics. Paul Biba, the current editor, has been doing a great job bringing together interesting articles from many different perspectives.

The Teleread discussion on my article Attributor ebook piracy numbers don’t add up had some interesting contributions. Jim Pitkow, CEO of Attributor, suggested that the difference between the data I got from Google AdWords and Attributor's data might be a result of a difference in methodology. While I used Google's predicted traffic numbers, Attributor used numbers from an actual AdWords campaign. For example, Attributor bought the keyword “lost symbol free ebook” along with 868 others and counted how many impressions were generated by Google.

AdWords predicts a total query volume of 62 per day for the keyword “lost symbol free ebook”. If each of the other keywords did the same volume of queries, Attributor should have seen 53,900 impressions per day for its ad campaign. It's not at all clear how 53,900 impressions turns into 1.5-3 millions searches for pirated content; that would be a factor of 30-60 larger than the traffic predicted by the AdWords estimator tool. Pitkow's comment seemed to suggest that the result of an actual advertising campaign would be different from the estimate, possibly accounting for the discrepancy.

So I did an actual advertising campaign of my own to see if this was true.

I've used AdWords in the past with amazingly cost-effective results. I would spend about five dollars a month at five cents per click to advertise a service that sold for thousands of dollars per year. My experience, albeit somewhat dated, was that the estimator tool overestimated, not underestimated, the actual traffic. I was interested to see if this was still true.

I constructed a "free ebook survey" landing page for my ad campaign, and added StatCounter analytics so I could see who clicked on my ad. I bought the keyword suggested by Pitkow: “lost symbol free ebook”.

AdWords gives you a number of settings to fine-tune your ad campaign. For example, I checked all the boxes for geographical coverage so my ads would be seen in as many countries and languages as possible.

AdWords offers three important options that determine the distribution of the ads. You can choose to advertise on Google only, Google and its "search partners" or on Google, it's search partners, and Google's "display network". I spoke with Pitkow last week, and he indicated that Attributor's study used both Google and its search partners.

The initial results of my ad campaign were alarming. In just four hours, I accumulated 7,000 impressions and 19 clicks; my campaign halted because my bids were too low. This traffic level would easily support Attributor's estimates. But when I looked into it, I found that I had included the "display network" without meaning to. What's more, the referring sites were really junky. I couldn't imagine who would use sites like "Lost World TV", NDParking and Sebaidu. The 61 cents I spent in those four hours may well have gone to a bunch of sites engaging in click fraud.

A full week of advertising on just Google, by contrast, resulted in a grand total of 15 impressions, much less than Googles estimate of 62 per day. I next added Google's search partners to my campaign, and got an impression rate about three times higher.

Even with the search partners, the reported search volume is about a tenth of the predicted volume. I must therefore revise the estimate I made about "consumer demand for pirated ebooks". Instead of 100,000-300,000 searches per day, 10,000-30,000 per day throughout the world seems to be a better estimate.

As a result of this experiment, the Attributor numbers are even more inexplicable than before. It's worth noting however, the one area where there's no disagreement. Both my investigations and Attributor's show that consumer interest in piracy is mostly located outside the US, UK, and Canada. Jane Litte's recent post on the geographical restrictions conundrum for ebooks (and comments thereto) does an excellent job of describing why that may be so.
Enhanced by Zemanta

Article any source

Tuesday, October 19, 2010

Attributor eBook Piracy Numbers Don't Add Up

In my article on "Consumer Demand for Pirated eBooks", I showed that Google Trends data tells a very different story from the one that anti-piracy services vendor Attributor derived from the very same data. I did not comment, however, on the headline that Attributor gave for its press release. The key finding of the report heralded by that release was that "Daily demand for pirated e-books can be estimated at 1.5-3 million people worldwide." This result has garnered some significant attention, because the number is quite large.

Extracting numbers using the tools used by Attributor is rather involved, and it's taken a while for me to carefully examine the available data. After doing this work, I've decided that when Attributor wrote "can be estimated at 1.5-3 million", they left out the word "blindly". As far as I can tell, Attributor is recklessly inflating the magnitude of ebook piracy; using the very same traffic measurement tools, I estimate the truth to be about 10% of the number they claim.

The Attributor numbers come from data generated by Google's AdWords service. AdWords is designed to help advertisers select advertising keywords and to manage budgets. For example, AdWords will tell you that the keyword "PDF" is used in approximately 101 million searches per month, worldwide, or 3.32 million searches per day. "PDF" is a keyword that a searcher might use in the course of a search for a pirated ebook, so you could reasonably assume that some percentage of these searches involve a consumer looking for a book they can avoid paying for. The trouble with this assumption is that most searches that include "PDF" have nothing to do with ebooks.

Another AdWords tool designed to assist Google advertisers is the keyword suggestion tool. In practice, you use this tool to refine keywords. Here is a table of the top ten refined searches for "PDF":
Keywordpercent of "pdf"
filetype pdf 36.69%
doc to pdf 6.03%
pdf download 3.30%
pdf to swf 3.30%
pdf to xls 2.70%
free pdf 2.70%
pdf free 2.70%
pdf to word 2.21%
pdf to rtf 2.21%
php pdf 1.81%
Of these, it's reasonable to assume that some percentage of the "pdf free" and a smaller fraction of the "pdf download" searches are related to consumers trying to avoid paying for books. The other searches are clearly unrelated to books. We can further use the keyword suggestion tool to refine these estimates. My review of over 700 refined keywords indicates that at most 4% of PDF searches, or 132,000 per day, are looking for ebooks of any kind.

A review of AdWords' suggested refinements for the term "rapidshare" reveals that searcher interest in ebooks is negligible compared to that for movies, TV, music and games. For example, Rapidshare is a "file-locker" site, and might be expected to appear in search terms for illegally distributed files. Of 743 suggested keywords, only one, accounting for 0.24% of "rapidshare" queries, or about 4,000 per day, is clearly related to ebooks:
Keywordpercent of "rapidshare"
files rapidshare 13.45%
rapidshare download 6.03%
download rapidshare 6.03%
download from rapidshare 6.03%
rapidshare megaupload 4.93%
free rapidshare 3.29%
rapidshare free download 2.70%
free rapidshare downloader 2.70%
free rapidshare download 2.70%
rapidshare download free 2.70%
free download rapidshare 2.70%
rapidshare free downloader 2.70%
download rapidshare free 2.70%
free rapidshare downloads 2.70%
download free rapidshare 2.70%
rapidshare searcher 2.19%
rapidshare search 1.80%
search on rapidshare 1.80%
dvdrip rapidshare 1.21%
rapidshare file 1.21%
rapidshare windows 7 1.21%
rapidshare mp3 1.21%
rapidshare dvd 0.99%
windows 7 rapidshare 0.81%
movie rapidshare 0.54%
rapidshare movie 0.54%
rapidshare upload 0.54%
upload rapidshare 0.54%
rapidshare downloader 0.44%
rapidshare file download 0.44%
rapidshare music 0.44%
music rapidshare 0.44%
download rapidshare files 0.36%
movies rapidshare 0.36%
rapidshare files download 0.36%
rapidshare windows xp 0.36%
720p rapidshare 0.36%
rapidshare premium accounts 0.30%
rapidshare password 0.30%
xbox 360 rapidshare 0.30%
game rapidshare 0.30%
password rapidshare 0.30%
rapidshare game 0.30%
rapidshare premium account 0.24%
premium account rapidshare 0.24%
rapidshare account premium 0.24%
premium rapidshare account 0.24%
rapidshare generator 0.24%
rapidshare engine 0.24%
rapidshare engine search 0.24%
up rapidshare 0.24%
rapidshare software 0.24%
software rapidshare 0.24%
rapidshare ebook 0.24%
Harry Potter and the Twilight Saga make appearances farther down the list, but only the titles that exist as movies.

Although direct interest in ebook torrents is so small that AdWords can barely measure it (~1500 searches per day), torrent search sites can give us another way to estimate the magnitude of interest in pirated ebooks. According to "KickassTorrents", the torrents active recently had this composition:
movies 30.04%
music 27.62%
tv 16.22%
apps 13.76%
games 5.52%
anime 5.43%
ebooks 1.42%
About 1.4 million searches using the keyword "torrent" are made on Google daily, according to AdWords. If the distribution of searches mirrors the distribution of files, this would indicate that searches for ebook torrents comprise about 46,200 per day.

All in all, I estimate that about 210,000 searches made on Google per day represent possible interest in pirated ebooks. About 30,000 of these come from the US. The "real" number for all countries could be as high as 300,000 or as low as 100,000. The 1.5-3 million numbers reported by Attributor are not within the range of plausibility.

One difficulty with using Google AdWords to gain insight into piracy is that it measures only a "shadow cast by piracy", as expressed by a commenter on my previous post. Nonetheless, AdWords sheds considerable light on patterns of demand. For example, the tools show clearly that it's common for people to search for movies and TV shows and acquire them extralegally. Also, they indicate that most of the demand, about 82%, for pirated ebooks comes from outside of the US, UK and Canada. Publishers should plan antipiracy strategies accordingly, based on data that can be confirmed independently.

Update: I have a followup post.
Enhanced by Zemanta

Article any source

Thursday, October 7, 2010

Consumer Demand for Pirated eBooks Stopped Growing in 2010

Online piracy of ebooks has been a persistent worry for book publishers who look at the successes and failures of other media that have moved to digital forms. A surprising number and variety of ebooks are easily availabile on file sharing websites and peer-to-peer networks that use bitTorrent and similar protocols. The possibility that this availability will cut into sales of licensed ebooks and even print books is a scary one for an industry that has had many decades of relative stability. At Digital Book World in January, Brian Napack, President of Macmillan, "delivered a passionate call to arms for publishers to fight piracy in the ebook space or risk permanent damage to the underpinnings of publishing as a commercial enterprise".

Adding to the ebook piracy hysteria have been studies of the prevalence of ebook piracy produced by Attributor, a company that sells anti-piracy services. I've previously written critically about Attributor's report that purported to find evidence that "Online Book Piracy Costs U.S. Publishers Nearly $3 Billion".

In their most recent report, Attributor has taken a rather clever approach to the measurement of ebook piracy. Instead of trying to track downloads, Attributor has begun to use Google Trends to gain an understanding of consumer demand for ebooks. Although there are many potential difficulties in using Google for this purpose, Google Trends is a powerful and useful tool for gaining insight into the things that web users around the world are looking for.

Attributor presents their data along with an alarming narrative of growing and pervasive ebook piracy, and points to the iPad as a contributing factor to an increase in demand for pirated ebooks. After playing around with Google Trends for a while, I've come to the conclusion that Attributor has narrowly selected data to fit their narrative; taken as a whole, Google Trends data broadly supports a rather different narrative: that the growth of consumer interest in pirated ebooks slowed significantly in 2009 and stopped in early 2010.

To understand how Google Trends informs the debate about the prevalence of ebook piracy, it helps to understand what activity is being measured. Google Trends measures the frequency that search terms are used. A consumer looking for a free copy of a particular work will typically search on the book title, adding  terms like "free" or "download" or "pdf" to locate downloadable files. A more sophisticated strategy, one that is quickly learned, is to add the name of a preferred download site. If the user prefers peer-to-peer networks, the word "torrent" can be added to locate "seed" files for the item. The file sharing sites most commonly used for this purpose are currently RapidShare, Megaupload, 4shared, and Hotfile. To use Google trends to measure the demand for a pirated ebook, you give it keywords that reproduce these searches. For example, demand for Stephanie Meyer's book Breaking Dawn can be assessed with a query such as this one.

To assess the overall state of ebook piracy, I used data from this query. Note that since the search is for ebooks generically, there's no telling for sure that the ebooks being searched for are really pirated; for the purposes of this study, I assumed that none of the ebooks being searched for are legally available on these sites. Calling them "pirated books" may be inaccurate, but I'll use that term anyway.

Some features of the data are immediately apparent. First of all, searches for pirated ebooks have increased a great deal over the past 5 years. It's worth noting however, that the most intense interest measured by Google occurs in India, the Philippines, Indonesia, Vietnam, Malaysia, Singapore, and eastern Europe. Less than half the search volume comes from the US. It's also easy to see seasonal peaks that obscure the shorter term trends. The peak periods for pirate ebook seeking are the December holidays and the beginning of September, presumably because of the start of school.
To eliminate seasonal variations, I computed the year over prior year growth of pirate ebook search activity. The resulting plot is quite smooth. After a few years of 100% per year growth, 2008 showed a clear slowing of growth. This slowing of growth continued up to the beginning of 2010, and then  flat-lined. Since February of 2010, the growth of interest in pirated ebooks has stopped completely.

It should be noted that this stabilization has occurred during a period of strong sales of ebook reader devices, including Kindle, Nook, and the iPad. Indeed, the unveiling of the iPad was coincident with the stabilization of demand for pirate ebooks.

It's hard to know for sure what's happening, but one interpretation of these patterns is that a broad increase in consumer-friendly availability of properly licensed ebooks over the last 2 years has squelched the growth of demand for ebooks from illicit sources. In that light, the remaining demand can be interpreted as a sign of poor availability for appropriately priced ebooks on college campuses and in developing countries.

While this data has to be seen as an encouraging sign for the book publishing industry, it's too soon to know if it will last. It's entirely possible that too-high prices, cumbersome DRM, or new technologies could reinvigorate the demand for illicitly shared ebook files. For the moment at least, the book publishing industry can exhale.
Enhanced by Zemanta

Article any source

Monday, April 12, 2010

Beware, Comment Spammers!

I had this great idea about how to fight comment spam. If you're not familiar with comment spam, you probably don't have your own blog and you think that "Kathryn" and "Patrick" who try to comment on this blog are just brain dead people. You might be right about the brain dead part, but I'm not sure they're really people.

Do you ever wonder why commenting on blogs can be such a hassle, or why so many blogs require moderation, or why many blogs don't accept comments on older posts, or forbid links in comments? It's because of comment spam. Spammers will submit comments such as "Your post is helpful and informative" or "We need to pay attention to the eco friend environment" that don't address the topic of the post in question. I'm not talking about targeted self-promotion here. It's not comment spam to link to an article you wrote on a similar topic, but it's definitely comment spam if you use a robot to do so. Or if you hire people in Asian boiler rooms to get around the CAPTCHA's that stop your robots.

It used to be that comment spam was done to improve the search engine ranking of websites. That motivation has largely gone away with the development of the "nofollow" tag. Blogs such as "Go To Hellman" attach add rel="nofollow" to any links in the comment threads. This tells spidering robots not to follow the specified links and tells search engines to ignore the links for purposes of site ranking.

I guess the people who have been leaving spam comments on my blog didn't get that memo. It's annoying to have to delete the comments, especially the ones in Chinese where links get hidden around the periods in "...". I went to the Blogger help pages to see if there's any way to report the abusive commenters (this blog restricts anonymous comments, so there's at least a user profile for every comment). There isn't. What's worse, Google tells you that if you don't remove those spam comments, your site's ranking will be hurt. Then I had my bright idea. I clicked on one of the links left in the spam comment. Then I picked some keywords from the page and plugged them into Google to find the site. There, at the bottom of the search result, was an option: Dissatisfied? Help us improve. Google is asking for feedback. I pasted in the URL for my comment spammer's site, and checked the radio button labeled "The results included spam." I clicked send, and my spammer's site was bound for Google oblivion!

Beware, comment spammers, I'm going to report you!

Though I felt good about it, I started to have doubts. A lot of these comment spammers seemed to be Asian; could it be that Asian search engines didn't get the nofollow memo either? Some quick googling confirmed my suspicion, China's leading search engine, Baidu, doesn't pay attention to the nofollow attribute! These comment spammers must be using my blog to juice their Baidu ranking!

Well maybe not. I did a few searches in Baidu. Baidu is probably the worst internet search engine I've ever tried! Baidu gives really stupid results for my vanity search. Baidu doesn't index my blog, my website, or anything I've ever posted. Perhaps China has blacked out the entire Google network, including Blogger, and Baidu doesn't see it any more. Or perhaps "Go To Hellman" has been banned for its post on Qin Shi Huangdi. Baidu has spidered a page from WorldCat that mentions some other Eric Hellman, and has picked up blog mentions of my by John Blyberg and in Dear Author but not much else. It's safe to assume that Baidu's strength is not English-language indexing.

So if Baidu doesn't index my blog, then spammers shouldn't be able to improve their Baidu rankings with comment spam in my blog. There must be some other motivation for the comments.

Another thing I noticed is that Baidu seems to be big on searching for MP3's and PDF's. It ranks sites like Rapidshare rather highly. Maybe Baidu and similar search engines spider websites like my blog to discover the mp3 files, the PDFs, and the video files that Baidu users are really looking for, and the intended audience of the spam comments is these content spiders. My blog has discussed ebooks, piracy and related topics, so maybe the spammers think its a good source for links to content. Who knows?

Another possibility is that the spammers are trying to get bloggers themselves to visit the their sites. "Patrick" from Madras is trying to sell "web templates". It turns out that his site has copied content from another site marketing web templates, which appear to me to be copies of other websites with much of the content stripped out. It's ironic: Patrick seems to be using a template for a web-template selling website to sell web templates.

After a few days, I checked back to see if the website I had complained about had been removed from Google or not. As it turns out, the site actually improved its Google ranking from #5 to #1 in my test search. So much for my career in comment spam scourgedom!
Reblog this post [with Zemanta]

Article any source

Tuesday, January 26, 2010

Deconstructing the Attributor Book Piracy Study

It was a lot of fun to have a post slashdotted and read by 100 times more people than any of my previous posts, however, the fact that it made fun of the way a study on piracy was "spun" overshadowed some serious points.
  1. The study in question, from Attributor, in fact has some substance beneath the spin.
  2. The number of books people get from libraries is comparable to the number they get from bookstores.
First, the Attributor study. Looking past the silly projections of book industry lost sales, there is some useful information there. In the study, illicit copies of 913 books were found on four one-click download sites which display download counts. The reported download counts for these 913 books totaled 3.2 million over 90 days. On average, each book was downloaded 3,500 times total, or 39 times per day.

These numbers are in rough agreement with a much smaller sample I collected for myself. I looked at only 10 books in a single category on a single site, and observed that download totals ranged for 0 to 5000, with an average of 578.

I followed up with some questions for Rich Pearson, the General Manager at Attributor. Most interesting to me, and probably of likely concern to publishers, was this: of 913 titles, chosen from Amazon lists to get a broad distribution rather than to focus on popularity, Attributor was able to find illicit copies of 90% of the books they looked for. I had expected that the most popular titles would be easy to find, but this breadth surprised me.

The extrapolations made to translate these download numbers into economic impact, while understandable from an marketing point of view, are multiply lacking in rigor.

The first assumption made by the study is that the "downloads" reported by  download sites represent potential readers of books. Attributor does not have any special relationship with the download sites to gain access to download statistics, they simply rely on the download counters published by the sites.  Assuming that the counts are not simply manufactured (and I've seen download sites that do just this) it's unlikely that the download sites bother to distinguish between robot activity and human activity. Robot activity is important because many of the download sites do not offer search themselves. This is a tactic that allows them to be unaware of, and thus not liable for,  the content that they host. The sites depend on other sites linking to them to drive traffic; they monetize the traffic by throttling downloads and offering "premium" subscriptions to people want to remove the throttles.

The second assumption made in the Attributor study is the way that they extrapolate numbers reported by 4 sites to the 25 sites that were monitored. What they did was to weight the sites based on the distribution of the 52,000+ takedown notices that Attributor has sent out since launching their monitoring service in July of '09. One problem with this is that while the scope of takedowns was limited to Attributor customers, the 913 monitored books were limited to non-customers. If Attributor's customers were focused in textbooks, for example, this would result in a bias towards sites used to host illicit textbooks. The other problem is "lamp post" bias- the extrapolation is based on what Attributor can find; sites hidden from Attributor would result in undercounting.

Freakonomics: A Rogue Economist Explores the Hidden Side of Everything (P.S.)Third, the sampling extrapolation to cover the entire industry has significant uncertainty. I worry most about price bias. The 1,132 downloads of the book Freakonomics were reported, which seems insignificant against sales of 2.5 million copies. Freakonomics, ranked #77 at Amazon,  can be bought new for $9.35. By contrast, Architect’s Drawings, which was downloaded over 10,000 times and ranks #547,491 at Amazon,  is a $60 hardcover. (The high download count is likely due to use as a textbook). Thus the "lost sales" for Architect’s Drawings will be hugely overweighted in the extrapolation.

Architect's Drawings: A selection of sketches by world famous architects through historyFinally, while the Attributor study does note that the actual effect of downloads on sales is speculative, there is a significant question as to whether the people downloading the book copies could have purchased the books even if they wanted to. It is evident from the chatter around the book downloads that many of the downloaders speak Spanish, French, Portuguese, Arabic, and Indonesian. According to Pearson, Attributor scans sites in many languages linking to illicit books, including Chinese, Japanese, Korean, French, Spanish, German, Italian, Czech, Polish, Russian, Portuguese and Austrian. The impact of piracy is likely to be very different in export markets. Not that this is shocking news to anyone.

Attributor is currently targetting its service (which starts at around $10,000) at large publishers, though it is working with a reseller called Author Guard to service smaller publishers. It points to its superior ability to find and take down illicit content as its comparative advantage. (I've previously surveyed other companies in this space; somehow I managed to overlook Attributor.)

Although many publishers will look at the numbers and conclude that services like Attributor are not worth the expense, I think the real danger from piracy is collective rather than individual, and thus the response should be collective rather than individual. The book publishing industry will only suffer if piracy becomes so widespread that piracy gains cultural acceptance. To prevent this from happening, book publishers need to lead the culture with both carrot and stick. Legal, allmost-free access to books must be provided such as occurs today in libraries, while illicit content should be taken down in as efficient manor as possible. This means that services such as Attributor's should be deployed on behalf of the entire industry, perhaps through a consortium, and not just by individual publishers. Finally, the book industry needs to figure out how to use ebooks to effectively address the needs of the developing world, or else huge markets will be forever out of reach.

The book publishing industry is entering some scary times and needs to decide who its friends are. A don't know whether technology companies with scary marketing will prove to be reliable friends or not. Amazon and Apple might end up being saviours, but publishers shouldn't expect them to be buddies. I'm pretty sure of one thing, though. Libraries are definitely not the enemy.
Enhanced by Zemanta

Article any source

Friday, January 15, 2010

Offline Book "Lending" Costs U.S. Publishers Nearly $1 Trillion

Hot on the heels of the story in Publisher's Weekly that "publishers could be losing out on as much $3 billion to online book piracy" comes a sudden realization of a much larger threat to the viability of the book industry. Apparently, over 2 billion books were "loaned" last year by a cabal of organizations found in nearly every American city and town. Using the same advanced projective mathematics used in the study cited by Publishers Weekly, Go To Hellman has computed that publishers could be losing sales opportunities totaling over $100 Billion per year, losses which extend back to at least the year 2000. These lost sales dwarf the online piracy reported yesterday, and indeed, even the global book publishing business itself.

From what we've been able to piece together, the book "lending" takes place in "libraries". On entering one of these dens, patrons may view a dazzling array of books, periodicals, even CDs and DVDs, all available to anyone willing to disclose valuable personal information in exchange for a "card". But there is an ominous silence pervading these ersatz sanctuaries, enforced by the stern demeanor of staff and the glares of other patrons. Although there's no admission charge and it doesn't cost anything to borrow a book, there's always the threat of an onerous overdue bill for the hapless borrower who forgets to continue the cycle of not paying for copyrighted material.

To get to the bottom of this story, Go To Hellman has dispatched its Senior Piracy Analyst (me) to Boston, where a mass meeting of alleged book traffickers is to take place. Over 10,000 are expected at the "ALA Midwinter" event. Even at the Amtrak station in New York City this morning, at the very the heart of the US publishing industry, book trafficking culture was evident, with many travelers brazenly displaying the totebags used to transport printed contraband.

As soon as I got off the train, I was surrounded by even more of this crowd. Calling themselves "Librarians", they talk about promoting literacy, education, culture and economic development, which are, of course, code words for the use and dispersal of intellectual property. They readily admit to their activities, and rationalize them because they're perfectly legal in the US, at least for now.

Typical was Susanne from DC, who told me that she's been involved in lending operations for over 15 years. This confirms our estimate that "lending" has been going on for over ten years, beyond even Google's memory. Our trillion dollar estimate may thus be on the conservative side. Of course, it's impossible to tell how many of these lent books would have been purchased legally if "libraries" were not an option, but we're not even considering the huge potential losses to publishers when "used" books are resold for pennies on the black markets.

The communications backbone for this vast enterprise appears to be Twitter. Already, there is constant chatter on the #alamw10 hashtag. Most messages are clearly coded references to illicit transactions. For example a trafficker with the alias "@libacat" tweets "Have to be on the bus to the airport at 6:41 tomorrow morning to make it to the airport to get on my plane to #alamw10". At first glance, it seems like a mundane tweet about travel plans, but the breathtaking ordinariness and triple redundancy is more likely a secret code. How else to understand @scolford's (correction: retweet of @SonjaandLibrary replying to @BPLBoston) tweet; "curling my toes in joy at the thought of visiting your library"?

I've attended this meeting before. When I register for the book lending confab, I'll be presented with an encrypted document labeled the "program", which once decoded, will tell me where I can meet other book traffickers, discuss arcane trafficker lore, and drink trafficker beer. It's thick with secret code words like YALSA, LITA and NMRT, and no apparent rhyme or reason in its layout, evidently to frustrate outside investigators. I'll be lucky if I can find a bathroom.

Two places I'll be sure to find this weekend will be the OCLC Blog Salon on Sunday evening and the Chinatown Storefront Library on Saturday afternoon. Say hello if you see me.

A more serious post on Attributor is forthcoming.

Update: here's my post on "Deconstructing the Attributor Study".
Enhanced by Zemanta

Article any source

Thursday, December 31, 2009

Do Libraries Have a Role in the Coming e-Book Economy?

You've probably heard it said that in Chinese, the word for "crisis" is composed from the words for "danger" and "opportunity". In the same presentation, you probably heard that there's no "I" in "TEAM". If you were skeptical of these attempts to extract wisdom from way language is written, you had good reason. The story about the Chinese word for crisis is not true. And even if it was true, it would be about as meaningful as the fact that the English word "SLAUGHTER" contains the word "LAUGHTER".

During my brief time working in "middle management", I was required to do "SWOT Analysis". SWOT stands for "Strengths, Weaknesses, Opportunities, Threats". As a planning exercise, it was quite useful, but it became comical when used as a management tool. Everyone understood the fake Chinese crisis wisdom, and we all made sure that our threats were the same as our opportunities, and our weaknesses were also our strengths.

On this last day of the "0"s, I've been reading a lot of prognostication about the next ten years. It's very relevant to this blog, as I've been using it to help me think about what to do next. Some things are not too hard to imagine: the current newspaper industry will shrink to maybe 10% its current size; the book publishing will reshuffle during the transition to e-books; Google will become middle-aged. The SWOT analysis for these will be easy.

The SWOT analysis that I have trouble with is the one for libraries. What threats to libraries will arise? Will Libraries as we know them even exist in 10 years?

I've heard publishers say they believe that there will be no role at all for libraries in the developing e-book ecosystem. If that's not a threat, I don't know what is! On the other hand, there's the example of the Barnes and Noble e-book reader, the Nook, that has the intriguing feature of being able to read books without buying them while you're in the bookstore! If there's a role for brick and mortar bookstores in the e-book ecosystem, then surely there's a role for libraries.

In thinking about what roles libraries will play when all books are e-books, I keep coming back to a conclusion that sounds odd at first: the prospective role of libraries will be entwined with that of piracy in the e-book ecosystem.

While there are fundamental differences between e-book libraries and e-book pirates, there are important similarities. As I noted in my article on copyright enforcement for e-books, libraries have traditionally played an important role in providing free access to print books; e-book pirates have as their mission the provision of free access to e-books. For this reason, libraries and pirates would occupy the same "market space" in an e-book ecosystem. This is not to say that libraries and pirates would be direct competitors; it's hard to imagine pirate sites appealing to many of the people who patronize libraries.

So where is the "threat" to libraries? Think about how book publishers will need to respond to the threat of e-book piracy. I've argued that publishers should do everything they can to reward e-book purchases, but that addresses only the high price segment of the market. Public libraries address the low-price segment of the market, providing books to people with a low willingness or ability to pay for access, while still providing a revenue stream for the publishers. To keep pirates from capturing this market in the e-book economy, publishers will need to facilitate the creation of services targeted at this market.

An analogy from the video business is appropriate here. DVDs can only satisfy part of the digital video market. Though it's taken a while for the studios to realize it, in order to effectively compete with video pirates, the movie studios need to have digital offerings like hulu.com that offer movies for free.

What will the free e-book services look like? Perhaps they'll be advertising sponsored services like Google Books. Perhaps they'll be publisher- or genre-specific subscription services that provide people a "free book" experience at a fixed monthly price. Unfortunately, it seems a bit unnatural that publishers would turn to libraries to create the sort of services that could replicate the role of the library in the e-book ecosystem- libraries just aren't entrepreneurial in that way.

Somehow I don't think that book publishers will warm to a "Napster for e-Books", even if it was labeled "e-Book Inter-Library Loan".

Still, I'm optimistic. Some horrific mashup of Open Library, Google Books, LibraryThing, WorldCat, BookShare, Facebook, Freebase, RapidShare and the Mechanical Turk is going to just the thing to save both libraries and publishers. You heard it here first. And if you find it scary- don't forget that you can't spell e-Book without BOO!
Reblog this post [with Zemanta]

Article any source

Monday, December 28, 2009

The Case Against Using Spoofed e-Books to Battle Piracy

I've known since I was four years old the difference between the Swedish Santa Claus and the American Santa Claus. The Swedish Santa Claus (the one who comes to our house) uses goats instead of reindeer and enters by the front door instead of the chimney. And instead of milk and cookies, the Swedish Santa Claus (a.k.a. Jul Tomten) always insists on a glass of glögg.

The glögg in our house was particularly good this year (used Cooks Illustrated recipe), so Jul Tomten stayed a bit longer than usual. I had a chance to ask him some questions.

"You're looking pretty relaxed this year, what's up?" I asked.

"It's this internet, you know. What with all the downloaded games, and music and e-books, my sleigh route takes only half the time it used to!"

"Really, that's amazing! I've read about the popularity of Kindle e-books, but I never imagined it might affect you! Are you worried that the sleigh and goat distribution channel will survive?"

"Oh not at all, Eric, remember, Christmas isn't about the goats, it's about the spirit! And even if all the presents could be distributed digitally, someone's got to go and drink the glögg, don't you think?"

"One thing I've been wondering, that list of yours, you know, the naughty and nice list... It must be very different now- do you look at people's Facebook profiles?"

"Ho ho ho ho. At the North Pole, your privacy is important to us, as the saying goes. Well, I'm going to let you in on a little secret. 'Naughty and Nice' is a bit of a misnomer. We never put coal in anyone's stocking. The way we look at it, there's goodness in each and every person."

"I guess I never thought of it that way."

"Just imagine how a child would feel if they woke up Christmas morning to find a lump of coal in their stocking! Even if the child was very naughty, do you a holiday disappointment would suddenly turn the child nice?"

"Besides, if we really wanted to put something useless in a stocking these days, it would be a VCR tape or an encyclopedia volume, not coal."

I've had a chance to reflect a bit on my chat with Santa, particularly about putting coal in naughty people's stockings. I've recently been studying how piracy might effect the emerging e-book market, and I've made suggestions about how to reinforce the practice of paying for e-books. But one respected book industry consultant and visionary, Mike Shatzkin, has made a suggestion that the book industry should take the coal-in-the-stocking approach to pirated e-books.

In an article entitled Fighting piracy: our 3-point program, Shatzkin proposes as point #1:
Flood the sources of pirate ebooks with “frustrating” files. Publishers can use all sorts of sophisticated tricks to find pirated ebooks, like searching for particular strings of words in the text. (You’d be shocked at how few words it takes to uniquely identify a file!) But people looking for a file to read will probably search by title and author. So publishers can find the sources of pirated files most likely to be used by searching the same way, the simple way.

But, then, when publishers find those illicit files, instead of take-down notices, which is the antidote du jour, we’d suggest uploading 10 or 20 or 50 files for every one you find, except each of them should be deficient in a way that will be obvious if you try to read them but not if you just take a quick look. Repeat Chapter One four times before you go directly to Chapter Six. Give us a chapter or two with the words in alphabetical order. Just keep the file size the same as the “real” ebook would be.
Points 2 and 3 of Shatzkin's "program" are reasonably good ideas. But this point 1 is a real clunker.

I'll admit, when I first read Shatzkin's proposal for publishers to put "sludge" on file sharing sites, I thought it an idea worth considering. After having studied the issue, however, I think that acting on the idea would be a foolish and shameful.

First of all, the idea is not original. The tactic of spoofing media files was deployed by the music industry in its battle against the file sharing networks that became popular after the demise of Napster. This tactic was promoted by MediaDefender, a company that also used questionable tactics such as denial of service attacks to shut down suspected pirate sites. Although the tactic was at first a somewhat effective nuisance for file sharers, the file sharing networks developed sophisticated defenses against this sort of attack. They adopted peer-review and reputation-rating systems so that deficient files and disreputable sharers could easily be discriminated. They instituted social peering networks so that untrusted file sharers could be excluded from the network of sharers. The culture of "may the downloader beware" has carried over for e-books. On one site I noted quite a bit of discussion of the true "last word" of Harry Potter and the Deathly Hallows along with chatter about file quality and the like. After seeing all this, Shatztkin's suggested point 1 seems quaint, to put it kindly.

The e-book-coal-in-the-stocking idea could also be dangerous if acted on. The tactic of disguising unwanted matter as attractive content has been widely adopted by attackers going back millennia to the builders of the Trojan Horse. The sludge could be as innocuous as a Amazon "buy-me" link with an embedded affiliate code, or it could be as malicious as a virus that lets a botnet take control of your computer if you open the file. When this really happened happened for video files, it was widely asserted, without any substantiation that the viruses were planted by the film industry operatives themselves. Thus, what began as a modest attempt to harass Napster file sharers ended up resulting in a smeared reputation for the film industry.

Obviously, Shatzkin is not advocating spoofing e-book files with harmful content on file sharing sites. But publishers who are tempted to follow his point #1 should consider the possibility that emitting large amounts of e-book sludge could provide ideal cover for scammers, spammers, phishers, and other cybercriminals. Then they should talk to their lawyers about "attractive nuisances" and "joint and several liability".

Go ahead and accuse me of believing in Santa Claus. I firmly believe that no matter what business you're in, not everybody gets corrupted. You have to have a little faith in people.

Reblog this post [with Zemanta]

Article any source