Showing posts with label Google+. Show all posts
Showing posts with label Google+. Show all posts

Friday, October 11, 2013

Google Patch Reward: tu correggi, lei ti premia

Sembra una di quelle promozioni che si trovano dal benzinaio e invece è tutto vero: Google ti premia, con un cash che oscilla dai 500$ ai 3.133,70$ se riesci a patchare uno dei grandi progetti dell'opensource.
“In short, we decided to try something new [...] Quite a few vulnerabilities trace back to preventable coding mistakes, or are made easier to exploit due to the absence of simple mitigation techniques. We are hoping to address this to some extent.”
La lista dei progetti coinvolti è molto ampia: OpenSSH, BIND, libjpeg, Chromium, Blink, zlib e lo stesso Kernel Linux, incluso KVM, ma prossimamente verranno inclusi anche Apache, nginx, SMTP come Sendmail e Postfix, VPN e molti altri.
In order to qualify, your patch must first be submitted directly to the maintainers of the project, and you must work with them to have it accepted into the repository and incorporated into a shipping version of the program.
Le patch in questione devono riguardare bug prettamente di sicurezza (quindi niente fix di singoli bug) e prima di poter ricevere il premio, dovranno essere valutate dal team mainstream ed accettate nel branch principale; Google stessa quindi non verrà coinvolta nel merging, né avvantaggerà un commit piuttosto che un altro: queste decisioni resteranno esclusive del progetto in questione, al Google Security Team andrà solo la valutazione (inviando i dettagli a security-patches@google.com) della qualità e della bontà della patch e della relativa importanza nel progetto, affinché poi possa quantificare l'ammontare del premio. P.S. Se siete Nordcoreani, Siriani o Sudanesi lasciate perdere, che tanto Google non vi paga. Ah, le tasse di transazione del premio non sono incluse.
Don't Be Evil. Almost.
Any source

Saturday, September 28, 2013

29/9/2013: Happy (one of the) Birthday(s), Google...

Cool graphic mapping evolution of Google over time (click to enlarge):


Sometime recently Google celebrated its 15th or 16th anniversary*. Whatever the date or the age is, Happy Birthday!


* Google was incorporated on September 4, 1998, but domain name google.com was registered on September 15, 1997. Officially the company recognises September 27th as its birthday, but apparently this is a new thing, since it is marked as such consistently only since 2006.


Any source

Thursday, September 19, 2013

Le macchine senza pilota di Google usano Linux

Nel corso della Embedded Linux Conference 2013 Andrew Chatham ha svelato alcune delle tecnologie dietro le macchine senza pilota del colosso di Mountain View, Google.

Durante la presentazione, Chatham rende noto che il sistema di navigazione automatico in sviluppo usa il sistema operativo Linux. O meglio, GNU/Linux.

In futuro ci porteranno a spasso
Doverosa la precisazione sull'uso di GNU/Linux , in quanto per il momento non si usa una versione personalizzata di Android ma una versione modificata di Ubuntu per permettere alla nota distribuzione anglosudafricana di girare al meglio in un'automobile.

Google parte da un presupposto: il 93% degli incidicenti fatali sono dovuti ad errori umani. Ecco quindi che Google Self-Driving, questo il nome del sistema di guida automatica, viene in aiuto.
Ovviamente è bene ricordare che questi sistemi per quanto siano automatizzati, possono comunque richiedere l'intervento umano in casi particolari.

Si viene anche a conoscenza del fatto che le macchine saranno guidate con l'ausilio delle mappe ma allo stesso tempo forniranno informazioni sulle mappe stesse. Ecco quindi che probabilmente le nostre macchine saranno le Google car del futuro.

A seguire il video della presentazione.


Any source

Tuesday, September 17, 2013

Il Play Store Devices arriva in Italia! Finalmente potremo acquistare i Nexus direttamente da Google!


La notizia tanto attesa è finalmente giunta, Google ha aperto il Play Store Devices anche in Italia!
Da oggi sarà finalmente possibile acquistare i device di Google direttamente dal Play Store senza dover aspettare i fornitori e i loro capricci!
Attualmente è possibile acquistare solamente il Nexus 7 (2013).
Any source

Tuesday, July 16, 2013

Richard Stallman intervistato su Snowden, Assange, Prism e la privacy


Nuova intervista al nostro caro Richard Stallman pubblicata sul canale Youtube di Russia Today. Stallman, intervistato dalla giornalista Sophie Shevardnadze, parla di Snowden, Assange e in generale del caso Prism.
Nel video Stallman approfondisce parla della privacy, del perché non ha un telefonino.
Secondo Stallman le prove che Snowden ha fornito circa la raccolta dei dati operata dai Governi non fanno che confermare quello di cui lui ci ha da sempre messo in guardia nel corso degli anni.
Il video è in inglese ma non è difficile da capire per i non anglofoni.
Ringrazio +Nicola Liguori  per aver condiviso il video su Google+.

Any source

Thursday, March 14, 2013

Google Reader chiude dal 1 luglio 2013 ecco le alternative

Terremoto e tragedia in casa Google. Con un annuncio a sorpresa Google annuncia la chiusura di Google Reader a partire dal 1 luglio 2013.



Cliccano su Learn more verrete indirizzati alla pagina http://support.google.com/reader/answer/3028851 dove potrete scaricare i vostri dati utente in formato XML da utilizzare con altri servizi

Ma quali servizi usare?
Personalmente su Linux preferisco installare alcuni programmi specifici, se conoscete le mie guide vi ho sempre suggerito Akregator per KDE e Liferea per tutti gli altri desktop environment.

Se però preferite utilizzare un servizio online simile a Google Reader le principali alternative gratuite che in queste ore stanno riscuotendo più successo da parte degli utenti sono tre, eccole:
Se avete suggerimenti e magari volete condividere con gli altri la vostra esperienza commentate questo post :)
Any source

Wednesday, October 17, 2012

Uno sguardo dall'interno ai data center di Google

Sono pochissime le persone che hanno accesso ai data center di Google e per un motivo molto valido: la nostra priorità assoluta è la sicurezza dei vostri dati e, per questo, facciamo di tutto per proteggerli mantenendo le nostre strutture sotto stretta sorveglianza. Sebbene nel corso degli anni abbiamo condiviso molti elementi della loro progettazione e molte delle migliori pratiche e sebbene pubblichiamo i dati sull'efficienza dei nostri data center dal 2008, soltanto un ristretto gruppo di dipendenti ha accesso diretto ai server.

A partire da questo momento, potrete entrare nei nostri datacenter con una visita virtuale. Grazie a Dove batte il cuore di Internet, il nostro nuovo sito arricchito con le bellissime fotografie di Connie Zhou, potrete avere una panoramica senza precedenti sulla tecnologia, le persone e i luoghi che mantengono Google in funzione. Il sito è disponibile anche in italiano.




Ma c'è di più: il data center di Lenoir, in North Carolina, è disponibile anche su Street View! Entrate dalla porta d'ingresso, salite le scale, girate a destra al tavolo da ping-pong e scendete al piano dei server. Oppure, fate due passi all'esterno dell’edificio per osservare il nostro impianto di raffreddamento ad alta efficienza energetica. Potete anche avventurarvi in un tour virtuale per comprendere meglio cosa state vedendo su Street View e osservare alcune delle nostre apparecchiature in azione.




Infine, abbiamo invitato Steven Levy, scrittore e giornalista di WIRED, a parlare con chi progetta la nostra infrastruttura e guardare per la prima volta come funziona un data center dall'interno. Il suo nuovo reportage, è un’esplorazione della storia e dell'evoluzione della nostra infrastruttura, con il primo reportage in assoluto da dentro un data center Google.

Quattordici anni fa, quando Google era un semplice progetto di ricerca di alcuni studenti, Larry e Sergey alimentavano il loro nuovo motore di ricerca utilizzando un paio di server economici e preconfezionati, combinati in modo piuttosto creativoDa allora siamo un po' cresciuti e ci auguriamo che questo sguardo a ciò che abbiamo costruito sia di vostro gradimento. Nei prossimi giorni condivideremo una serie di post sul Google Green Blog per analizzare alcune delle fotografie in modo più dettagliato, perciò se vi interessa il tema, non perdetevi i prossimi appuntamenti.

Scritto da: da Urs Hölzle, Senior Vice President, Technical Infrastructure
Any source

Tuesday, May 29, 2012

Quando si parla di privacy online tutto questo è "buono a sapersi"

Internet è ormai uno strumento indispensabile nella nostra vita di tutti i giorni, nel lavoro come nella vita privata. Ci sono tuttavia momenti in cui può suscitare delle preoccupazioni e non tutti gli utenti si sentono perfettamente preparati ad affrontarle. Per questo abbiamo sviluppato il progetto Buono a Sapersi, che ha l’obiettivo di aiutare le persone a navigare in piena sicurezza e a gestire con consapevolezza i propri dati online. 

Buono a Sapersi è un’iniziativa sviluppata in Italia in collaborazione con la Polizia Postale e delle Comunicazioni e si articola in quattro sezioni, due generali sul web e due più specifiche sui servizi Google. Il sito, che incorpora anche il Centro per la sicurezza della famiglia, è arricchito inoltre da numerosi contenuti video molto semplici e operativi.

Potrete imparare come navigare sicuri, avere informazioni sul modo in cui i vostri dati sono utilizzati su Google e su tutto il web e ottenere suggerimenti utili per gestire al meglio l'esperienza online della vostra famiglia: come scegliere una password sicura, cos’è un malware e come raddoppiare la sicurezza del vostro account Google con la verifica in due passaggi.

Chi non ha un amico che usa la stessa password per ogni servizio web, lascia il suo computer aperto senza bloccare lo schermo e pensa che una conversazione sui cookie sia un invito a colazione? L'alfabetizzazione digitale diventa ogni giorno sempre più importante, speriamo che anche voi possiate trovare interessanti i contenuti di Buono a sapersi e vogliate aiutarci a diffondere la conoscenza di questi temi sensibilizzando chi vi sta accanto.

Vi aspettiamo su www.google.it/BuonoASapersi per saperne di più.


Any source

Tuesday, May 15, 2012

Google Forms - useful way to get data for your business

This entry is largely based on my own recent discovery.  Google seems to have a service for just about everything, and it seems like every time you turn around, you can discover something new on the site.  This leads me to talk about Google Forms, which is a great (and free) way that you and your business can get data from a small or massive set of people.

Its operation is simple (as, it seems, are most things when it comes to Google): you basically create your own form in whatever format you want.  You ask a question and then specific the type of answer you want to get back - open ended, select one, etc.  When someone is done with the form, they submit it.  You can then track that information in an easy to read spreadsheet.  You can also create a link to the survey that you can then send via E-mail, Social Media, etc.

How can this be useful for your business?  Well, you can create a very easy to respond to survey, RSVP to an event, get opinions, etc.  There is no limit on respondents, the amount of questions that you ask and relatively few formatting constraints.  From a technical perspective, these are easy to design, and the data is easy to read.    This is somewhat similar to other services like SurveyMonkey - but more stripped down.  There are no bells, no whistles (and no advertising on the site!) - its just easy to use, and if you are already addicted to Google products (as I am), this can really save you some time.

What do you think?
Article any source

Thursday, July 28, 2011

A Bay Area Company Challanges Google Ranking

With federal regulators pursuing an antitrust probe over whether Google (GOOG) is abusing its dominance in search to favor its own online products, a company that owns several Bay Area websites promoting local small businesses is taking the rare step of publicly challenging the fairness of the search giant.

ShopCity, the parent company of local sites such as ShopPaloAlto.com, ShopMountainView.com and ShopPleasanton.com, says Google provides it an unfairly low ranking, especially since those sites have the backing of groups such as the city of Menlo Park, the Palo Alto Chamber of Commerce and the Palo Alto Weekly newspaper. A search for "Palo Alto restaurants" on Google this week didn't reveal a ShopPaloAlto.com result until the seventh page of results, while the site ranks at the top for identical searches on Yahoo (YHOO) or Microsoft's Bing.

"The most dangerous man is a man with nothing to lose, and that's the position they've put us in," said Colin Pape, the president of ShopCity, who is considering complaining to federal regulators.

Google defends its rankings as serving users. But ShopCity also says Google is taking its content and displaying it in Google Places, which like ShopCity displays business information such as location, operating hours and customer reviews. The practice is called "scraping," and companies like Yelp and TripAdvisor.com also have complained about the practice.

As Google moves heavily into local services in an effort to attract small advertisers, and faces antitrust scrutiny from the Federal Trade Commission and other regulators, long-standing complaints about how it ranks millions of websites may take on a more problematic ring for the Mountain View Internet giant, as unhappy websites allege anticompetitive behavior. Thursday, the Senate Judiciary Committee said Eric Schmidt, Google's executive chairman, would testify Sept. 21 about antitrust issues. The Texas attorney general also has an active antitrust investigation against Google and other state attorneys general may follow.

A quality issue?

Google says its low ranking of ShopCity sites is fair because the vast majority of its more than 8,100 local sites across the U.S. and Canada do not feature original content. ShopCity acknowledges that all but 44 of its sites do not yet have original content, and the company says it has asked the search giant not to crawl and rank those sites. But Google says it must consider the collective authority of the company's Internet properties, just as someone wouldn't judge a supermarket tabloid as superior to a national daily newspaper based on the accuracy of one story.

"We're committed to returning high-quality sites to our users," said Gabriel Stricker, a Google spokesman. "In the case of ShopCity, this is a network of thousands of sites that appear lower in Google's rankings because nearly 100 percent of the sites violate our quality guidelines. For years, these sites have contained little original content, substantial duplicate content, along with cookie-cutter templates. Our users frequently complain to us about these kinds of sites."

But ShopCity does have 44 sites, including seven Bay Area sites from Gilroy to Menlo Park to Pleasanton, that feature extensive original content, including features like restaurant menus, discount offers from local merchants and community event listings.

The low search ranking has also angered local business leaders who say they are trying to create a quality online presence for independent businesses that can compete against local listings by big companies like Google or Yelp. One business linking extensively to ShopPaloAlto.com is the Palo Alto Weekly, but publisher Bill Johnson says the local listing site is hard to find on Google. He finds the site's low ranking suspicious, because Google search results are based in part on links to other websites.

"Clearly Google is monkeying with the settings manually to prevent that from happening. We're out trying to build a community business website that is providing local businesses with some really great tools they can use to enhance their online presence," Johnson said. "With all this antitrust stuff going on, is Google really trying to make it difficult for those entities who are attempting to compete in the local consumer Web area?"

Search industry expert Danny Sullivan, editor in chief of Search Engine Land, said such suspicions about a site as small as ShopPaloAlto.com are "ludicrous. If that was what (Google) was worried about, you would never find Yelp," a formidable competitor for Google that offers restaurant reviews and business listings, Sullivan said.

But Sullivan said Google should be able to differentiate between higher-quality ShopCity sites such as the Bay Area sites, and placeholder sites waiting until ShopCity makes partnerships with local groups for listings.

Abrupt changes

"Do I think there is something wrong here? Probably," Sullivan said. "Do I think it's because Google has an antitrust agenda? No."

ShopCity is represented by Palo Alto antitrust attorney Gary Reback, who represented a group of companies, including Microsoft, in opposing Google's plan to scan millions of out of print books.

ShopCity's suspicions were triggered when, two days after the June 24 announcement of the FTC inquiry, ShopPaloAlto.com and other Bay Area sites suddenly began ranking better on Google, providing a temporary 400 percent increase in search-engine driven traffic, only to fall back to near zero in mid-July when Google downgraded its sites again. Pape said ShopCity was unable to get a reply from Google about what happened.

Stricker, the Google spokesman, said an earlier automated penalty imposed against ShopCity sites by coincidence had expired at that time, but Google imposed another penalty when it received outside complaints about ShopCity sites. Local partners say they still have high hopes for their network.

"I absolutely think it's a valuable service," said Paula Sandas, president and CEO of the Palo Alto Chamber of Commerce. "This is a very simple and economical way to have the Web presence that the rest of the world seems to have."

Article any source

Wednesday, June 8, 2011

Our Metadata Overlords and That Microdata Thingy

On June 2, our Metadata Overlords spoke. They told us that they'll only listen when we tell them things using a specialized vocabulary they've now given us at the schema.org website. Although we can still use our stone tablets if that's what we're using now, we're expected to migrate to a new Microdata Thingy, assuming that we really want them to pay attention to our website metadata supplications.

There are among us believers, who, led by druids enraptured by the power of stone tablets to carry truth, will shun the new thingy, but most of us will meekly comply with the edicts of the overlords. We're not able to distinguish the druidic language of the tablets from the new liturgy of of the state church. Many things are difficult to articulate in the new vocabulary, but gosh, those tablets were heavy to carry around. And the new thingy doesn't seem so awful, although it's difficult to tell with the mumbled sermons and hymn singing and all.

I hope the overlords don't try to take our pagan rituals of Friending and Liking away from us, though. The incantations used to invoke and bless the Like ritual also use the druidic language, and the help scrolls tell us we might confuse the overlords if we use more than one language in our prayers.

My soul remains troubled, however, at the thought that the Overlords care not for truth and for justice. Sometimes it seems as though the overlords want only for our offerings of attention and seek only to feed our lust for food, drink, entertainment, debauchery and money. Yes, there are new words for our books and learning, but we can say so little about these in schema.org language that our wizards and mages will be mute if they ever choose to enter that realm.

I myself was present at a conclave of such mages and wizards dedicated to the entwinement of data from libraries, museums and archives in full openness. When tweet of the new order came, we endeavored to learn more of schema.org and its thingy. We questioned whether the thingy was an abomination against openness, or whether we might exploit its Overlord endorsement to make our own spells more powerful. We agreed to teach each other our new thingy spells, even as our colleagues elsewhere figured out how to chisel the new vocabulary into stone. Word came from other lands that the new vessel would founder trying to cross the seas.

We then visited the temple of the archive and found the servers cool to the touch. We heard words from a past oracle, ate as they never ate in Rome, drank cool drafts, and returned home emboldened with an enlarged appreciation of intermingled bits.

So it was said, so shall we do.

Notes:
  1. Google's blog post on adopting microdata was signed by R. V. Guha who had a bit to do with the creation of RDF.
  2. It's not really a surprise that Google doesn't care about RDFa. In my article on RDFa from 2009, I pointed to mistakes that Google made in their RDFa documentation. They never fixed it.
  3. Schema.org can't even list all of its schemata- the web page, chock full of non-breaking spaces, is truncated!.
  4. The current microdata spec is in an odd state where it's confused about how to define an itemtype. In fact, the mechanism for defining new itemtypes is gone! Here's what it says:
    The item type must be a type defined in an applicable specification.

    Except if otherwise specified by that specification, the URL given as the item type should not be automatically dereferenced.

    A specification could define that its item type can be derefenced to provide the user with help information, for example. In fact, vocabulary authors are encouraged to provide useful information at the given URL.
    Apparently, stuff was removed for some sort of political reason- it's there in the WHAT-WG version; note that Google links to the W3C version, which is not fully baked.
  5. the Schema.org terms of service are creepy when you get to the part about patents.
  6. The big selling point for RDFa was that Google, Yahoo and Bing supported it for Rich Snippets and the like. But Microdata's inability to easily support complex markup turned out to be an key feature for the search engines. The moral of the story for standards developers: your best customers are always righter than the others.
  7. In the video, Brewster Kahle reads from the last page of A Manual on Methods of Reproducing Research Material by Robert C. Binkley (1936). OCLC Number 14753642. Peter Binkley, a meeting participant, donated a copy of his grandfather's book to the Internet Archive, along with permission to make it free to the public.
  8. Henri Sivonen has written a very readable and informed discussion about Microdata, RDFa, Schema.org and the process of making standards that you should read if you are interested in why things are the way they are in HTML5.

Article any source

Friday, April 1, 2011

Alla ricerca di talenti

Stiamo assumendo in Google. Non è una novità visto che Eric Schmidt, il nostro attuale CEO, ha annunciato a inizio anno che intendiamo incrementare il nostro organico a livello mondiale di altre 6.000 unità in tutto il mondo. E stiamo assumendo anche in Italia; abbiamo molte posizioni aperte e altre si aggiungono a ritmo sostenuto.

Qualsiasi assunzione è fondamentale per Google. Altrimenti non dedicheremmo così tante energie e passione nel riuscire a scovare talenti in tutto il mondo. Per quanto mi riguarda più direttamente, la posizione di Head of Operations and Strategy per l'Italia è particolarmente significativa ed esposta a tutto quanto concerne la direzione strategica di Google nel nostro paese.
Come sempre, qualsiasi candidatura deve essere presentata attraverso il sito web, in modo da essere sottoposta a valutazione da parte dei recruiter coinvolti nel processo di selezione.

Spero d’avervi incuriosito e stimolato al punto giusto per prendere in considerazione questa opzione. Vi garantisco che lavorare per Google è proprio un’esperienza … da vivere!

Any source

Tuesday, October 19, 2010

Attributor eBook Piracy Numbers Don't Add Up

In my article on "Consumer Demand for Pirated eBooks", I showed that Google Trends data tells a very different story from the one that anti-piracy services vendor Attributor derived from the very same data. I did not comment, however, on the headline that Attributor gave for its press release. The key finding of the report heralded by that release was that "Daily demand for pirated e-books can be estimated at 1.5-3 million people worldwide." This result has garnered some significant attention, because the number is quite large.

Extracting numbers using the tools used by Attributor is rather involved, and it's taken a while for me to carefully examine the available data. After doing this work, I've decided that when Attributor wrote "can be estimated at 1.5-3 million", they left out the word "blindly". As far as I can tell, Attributor is recklessly inflating the magnitude of ebook piracy; using the very same traffic measurement tools, I estimate the truth to be about 10% of the number they claim.

The Attributor numbers come from data generated by Google's AdWords service. AdWords is designed to help advertisers select advertising keywords and to manage budgets. For example, AdWords will tell you that the keyword "PDF" is used in approximately 101 million searches per month, worldwide, or 3.32 million searches per day. "PDF" is a keyword that a searcher might use in the course of a search for a pirated ebook, so you could reasonably assume that some percentage of these searches involve a consumer looking for a book they can avoid paying for. The trouble with this assumption is that most searches that include "PDF" have nothing to do with ebooks.

Another AdWords tool designed to assist Google advertisers is the keyword suggestion tool. In practice, you use this tool to refine keywords. Here is a table of the top ten refined searches for "PDF":
Keywordpercent of "pdf"
filetype pdf 36.69%
doc to pdf 6.03%
pdf download 3.30%
pdf to swf 3.30%
pdf to xls 2.70%
free pdf 2.70%
pdf free 2.70%
pdf to word 2.21%
pdf to rtf 2.21%
php pdf 1.81%
Of these, it's reasonable to assume that some percentage of the "pdf free" and a smaller fraction of the "pdf download" searches are related to consumers trying to avoid paying for books. The other searches are clearly unrelated to books. We can further use the keyword suggestion tool to refine these estimates. My review of over 700 refined keywords indicates that at most 4% of PDF searches, or 132,000 per day, are looking for ebooks of any kind.

A review of AdWords' suggested refinements for the term "rapidshare" reveals that searcher interest in ebooks is negligible compared to that for movies, TV, music and games. For example, Rapidshare is a "file-locker" site, and might be expected to appear in search terms for illegally distributed files. Of 743 suggested keywords, only one, accounting for 0.24% of "rapidshare" queries, or about 4,000 per day, is clearly related to ebooks:
Keywordpercent of "rapidshare"
files rapidshare 13.45%
rapidshare download 6.03%
download rapidshare 6.03%
download from rapidshare 6.03%
rapidshare megaupload 4.93%
free rapidshare 3.29%
rapidshare free download 2.70%
free rapidshare downloader 2.70%
free rapidshare download 2.70%
rapidshare download free 2.70%
free download rapidshare 2.70%
rapidshare free downloader 2.70%
download rapidshare free 2.70%
free rapidshare downloads 2.70%
download free rapidshare 2.70%
rapidshare searcher 2.19%
rapidshare search 1.80%
search on rapidshare 1.80%
dvdrip rapidshare 1.21%
rapidshare file 1.21%
rapidshare windows 7 1.21%
rapidshare mp3 1.21%
rapidshare dvd 0.99%
windows 7 rapidshare 0.81%
movie rapidshare 0.54%
rapidshare movie 0.54%
rapidshare upload 0.54%
upload rapidshare 0.54%
rapidshare downloader 0.44%
rapidshare file download 0.44%
rapidshare music 0.44%
music rapidshare 0.44%
download rapidshare files 0.36%
movies rapidshare 0.36%
rapidshare files download 0.36%
rapidshare windows xp 0.36%
720p rapidshare 0.36%
rapidshare premium accounts 0.30%
rapidshare password 0.30%
xbox 360 rapidshare 0.30%
game rapidshare 0.30%
password rapidshare 0.30%
rapidshare game 0.30%
rapidshare premium account 0.24%
premium account rapidshare 0.24%
rapidshare account premium 0.24%
premium rapidshare account 0.24%
rapidshare generator 0.24%
rapidshare engine 0.24%
rapidshare engine search 0.24%
up rapidshare 0.24%
rapidshare software 0.24%
software rapidshare 0.24%
rapidshare ebook 0.24%
Harry Potter and the Twilight Saga make appearances farther down the list, but only the titles that exist as movies.

Although direct interest in ebook torrents is so small that AdWords can barely measure it (~1500 searches per day), torrent search sites can give us another way to estimate the magnitude of interest in pirated ebooks. According to "KickassTorrents", the torrents active recently had this composition:
movies 30.04%
music 27.62%
tv 16.22%
apps 13.76%
games 5.52%
anime 5.43%
ebooks 1.42%
About 1.4 million searches using the keyword "torrent" are made on Google daily, according to AdWords. If the distribution of searches mirrors the distribution of files, this would indicate that searches for ebook torrents comprise about 46,200 per day.

All in all, I estimate that about 210,000 searches made on Google per day represent possible interest in pirated ebooks. About 30,000 of these come from the US. The "real" number for all countries could be as high as 300,000 or as low as 100,000. The 1.5-3 million numbers reported by Attributor are not within the range of plausibility.

One difficulty with using Google AdWords to gain insight into piracy is that it measures only a "shadow cast by piracy", as expressed by a commenter on my previous post. Nonetheless, AdWords sheds considerable light on patterns of demand. For example, the tools show clearly that it's common for people to search for movies and TV shows and acquire them extralegally. Also, they indicate that most of the demand, about 82%, for pirated ebooks comes from outside of the US, UK and Canada. Publishers should plan antipiracy strategies accordingly, based on data that can be confirmed independently.

Update: I have a followup post.
Enhanced by Zemanta

Article any source

Sunday, June 27, 2010

Global Warming of Linked Data in Libraries

Libraries are unusual social institutions in many respects; perhaps the most bizarre is their reverence for metadata and its evangelism. What other institution considers the production, protection and promulgation of metadata to be part of its public purpose?

The W3C's Linked Data activity shares this unusual mission. For the past decade, W3C has been developing a technology stack and methodology designed to support the publication and reuse of metadata; adoption of these technologies has been slow and steady, but the impact of this work has fallen short of its stated ambitions.

I've been at the American Library Association's Annual Meeting this weekend. Given the common purpose of libraries and Linked Data, you would think that Linked Data would be a hot topic of discussion. The weather here has been much hotter than Linked Data, which I would describe as "globally warming". I've attended two sessions covering Linked Data, each attended by between 50 and 100 delegates. These followed a day long, sold-out  preconference. John Phipps, one of the leaders in the effort to make library metadata compatible with the semantic web, remarked to me that these meeting would not have been possible even a year ago. Still, this attendance reflects only a tiny fraction of metadata workers at the conference; Linked Data has quite a ways to come. It's only a few months ago that the W3C formed a Library Linked Data Incubator Group.

On Friday morning, there was an "un-conference" organized by Corey Harper from NYU and Karen Coyle, a well-known consultant. I participated in a subgroup looking at use cases for library Linked Data. It took a while for us to get around to use cases though, as participants described that usage was occurring, but they weren't sure what for. Reports from OCLC (VIAF) and Library of Congress (id.loc.gov) both indicated significant usage but little feedback. The VIVO project was described as one with a solid use case (giving faculty members a public web presence), but no one from VIVO was in attendance.

On Sunday morning, a meeting of the Association for Library Collections and Technical Services (ALCTS), Rebecca Guenther, Library of Congress, discussed id.loc.gov, a service that enables both humans and machines to programatically access authority data at the Library of Congress. Perhaps the most significant thing about id.loc.gov is not what it does but who is doing it. The Library of Congress provides leadership for the world of library cataloguing; what LC does is often slavishly imitated in libraries throughout the US and the rest of the world.  id.loc.gov started out as a research project but is now officually supported.

Sara Russell-Gonzalez of the University of Florida then presented the VIVO which has won a big chunk of funding from the National Center for Research Resources, a branch of NIH. The goal of VIVO is to build an "interdisciplinary national network enabling collaboration and discovery between scientists across all disciplines." VIVO started at Cornell and has garnered strong institutional support there, as evidenced by an impressive web site. If VIVO is able to gain similar support nationally and internationally, it could become an important component of an international research infrastructure. This is a big "if". I asked if VIVO had figured out how to handle cases where researchers change institutional affiliations; the answer was "No". My question was intentionally difficult; Ian Davis has written cogently about the difficulties RDF has in treating time-dependent relationships. It turns out that there are political issues as well. Cornell has had to deal with a case where an academic department wanted to expunge affiliation data for a researcher who left under cloudy circumstances.

At the un-conference, I urged my breakout group to consider linked data as a way to expose library resources outside of the library world as well as a model for use inside libraries. It's striking to me that libraries seem so focused on efforts such as RDA, which aim to move library data models into Semantic Web compatible formats. What they aren't doing is to make library data easily available in models understandable outside the library.

The two most significant applications of Linked Data technologies so far are Google's Rich Snippets and Facebook's Open Graph Protocol (whose user interface, the "Like" button, is perhaps the semantic webs most elegant and intuitive). Why aren't libraries paying more attention to making their OPAC results compatable with these application by embedding RDFa annotations in their web-facing systems? It seems to me that the entire point of metadata in libraries is to make collections accessible. How better to do this than to weave this metadata into peoples lives via Facebook and Google? Doing this will require the dumbing-down of library metadata and some hard swallowing, but it's access, not metadata quality, that's core to the reason that libraries exist.



Enhanced by Zemanta

Article any source

Monday, April 12, 2010

Beware, Comment Spammers!

I had this great idea about how to fight comment spam. If you're not familiar with comment spam, you probably don't have your own blog and you think that "Kathryn" and "Patrick" who try to comment on this blog are just brain dead people. You might be right about the brain dead part, but I'm not sure they're really people.

Do you ever wonder why commenting on blogs can be such a hassle, or why so many blogs require moderation, or why many blogs don't accept comments on older posts, or forbid links in comments? It's because of comment spam. Spammers will submit comments such as "Your post is helpful and informative" or "We need to pay attention to the eco friend environment" that don't address the topic of the post in question. I'm not talking about targeted self-promotion here. It's not comment spam to link to an article you wrote on a similar topic, but it's definitely comment spam if you use a robot to do so. Or if you hire people in Asian boiler rooms to get around the CAPTCHA's that stop your robots.

It used to be that comment spam was done to improve the search engine ranking of websites. That motivation has largely gone away with the development of the "nofollow" tag. Blogs such as "Go To Hellman" attach add rel="nofollow" to any links in the comment threads. This tells spidering robots not to follow the specified links and tells search engines to ignore the links for purposes of site ranking.

I guess the people who have been leaving spam comments on my blog didn't get that memo. It's annoying to have to delete the comments, especially the ones in Chinese where links get hidden around the periods in "...". I went to the Blogger help pages to see if there's any way to report the abusive commenters (this blog restricts anonymous comments, so there's at least a user profile for every comment). There isn't. What's worse, Google tells you that if you don't remove those spam comments, your site's ranking will be hurt. Then I had my bright idea. I clicked on one of the links left in the spam comment. Then I picked some keywords from the page and plugged them into Google to find the site. There, at the bottom of the search result, was an option: Dissatisfied? Help us improve. Google is asking for feedback. I pasted in the URL for my comment spammer's site, and checked the radio button labeled "The results included spam." I clicked send, and my spammer's site was bound for Google oblivion!

Beware, comment spammers, I'm going to report you!

Though I felt good about it, I started to have doubts. A lot of these comment spammers seemed to be Asian; could it be that Asian search engines didn't get the nofollow memo either? Some quick googling confirmed my suspicion, China's leading search engine, Baidu, doesn't pay attention to the nofollow attribute! These comment spammers must be using my blog to juice their Baidu ranking!

Well maybe not. I did a few searches in Baidu. Baidu is probably the worst internet search engine I've ever tried! Baidu gives really stupid results for my vanity search. Baidu doesn't index my blog, my website, or anything I've ever posted. Perhaps China has blacked out the entire Google network, including Blogger, and Baidu doesn't see it any more. Or perhaps "Go To Hellman" has been banned for its post on Qin Shi Huangdi. Baidu has spidered a page from WorldCat that mentions some other Eric Hellman, and has picked up blog mentions of my by John Blyberg and in Dear Author but not much else. It's safe to assume that Baidu's strength is not English-language indexing.

So if Baidu doesn't index my blog, then spammers shouldn't be able to improve their Baidu rankings with comment spam in my blog. There must be some other motivation for the comments.

Another thing I noticed is that Baidu seems to be big on searching for MP3's and PDF's. It ranks sites like Rapidshare rather highly. Maybe Baidu and similar search engines spider websites like my blog to discover the mp3 files, the PDFs, and the video files that Baidu users are really looking for, and the intended audience of the spam comments is these content spiders. My blog has discussed ebooks, piracy and related topics, so maybe the spammers think its a good source for links to content. Who knows?

Another possibility is that the spammers are trying to get bloggers themselves to visit the their sites. "Patrick" from Madras is trying to sell "web templates". It turns out that his site has copied content from another site marketing web templates, which appear to me to be copies of other websites with much of the content stripped out. It's ironic: Patrick seems to be using a template for a web-template selling website to sell web templates.

After a few days, I checked back to see if the website I had complained about had been removed from Google or not. As it turns out, the site actually improved its Google ranking from #5 to #1 in my test search. So much for my career in comment spam scourgedom!
Reblog this post [with Zemanta]

Article any source

Monday, March 22, 2010

Second Sourcing, Application Interfaces, and a 16 Bit Static Ram

While returning my Mac Plus to the attic, I decided to bring out some electronics relics I have from an even earlier era. The photo shows an undiced 1.25" wafer of integrated circuits and two packaged chips from around 1969. (The acorn hat is for scale) There are about 200 transistors on each chip. I think it's a 16 bit static RAM chip- with a magnifying glass, I can see and count the 16 cells. For context, when the Mac 128K Mac came out, it shipped with 64Kb DRAM chips. (about 1000x more dense). Today 4Gb chips are in production, a factor of a billion denser than my relic, and the silicon wafers are 30 cm in diameter.

My father was one of the founders of Solid State Scientific, Inc., (SSSI) a company that made CMOS integrated circuits. SSSI, located in Montgomeryville PA, started out as a second-source supplier for RCA's line of low-power CMOS logic chips. In the electronics industry, it has been a common practice for component manufacturers to license their circuit designs or specifications to other manufacturers so that their customers would be assured of an adequate supply. The second source company could compete on price or performance. For example, engineers could design systems with the 4060 14-bit ripple counter chip with internal oscillator, and know that they could buy a replacement chip from either RCA or SSSI. If RCA's fab was fully booked, SSSI would be able to fill the gap. There was no vendor lock-in.

Second source relationships could be tricky- AMD and Intel famously ended up litigating AMD's second-source status for the 8086 series of microrocessors. Logic family chips were commodities, and profit margins were thin. The second-source gambit was a judgment that a company could make more money by driving prices down and volume up. Companies like SSSI were always chasing after higher profit margins in new applications such as custom circuits for digital watches. The large volume parts would pay for their fabs, and the proprietary circuits would earn the profits, or at least that was the idea. Vendor lock-in, while while it might discourage adoption and reduce volume, is good for profitability.

As chips become more and more complicated, the chip manufacturing industry realligned. Today, apart from giants like Intel, most chips are manufactured by foundry companies that don't do chip design at all. Chip design companies try to maintain high margins with exclusive intellectual property; the foundry companies aggregate volume and drive down cost by manufacturing chips from many different design companies.

I've been thinking about the way that the advance of technology moves application interfaces. In the days of the CMOS logic chips, the application interface was a spec sheet and a logic diagram. That was everything an circuit designer needed to include the component in a design. Today that interface has migrated onto the chip and into software;  chip foundries provide software models for components ranging from transistors to processor blocks for designers to include in their products.

When software engineers talk about application interfaces, they're usually thinking about function calls and data structures that one block of software can use to interact with other blocks of software. These interfaces, once published and relied on, tend to be much more stable over time than the code hidden behind them. To some extent, software application interfaces can hide hardware implementations as easily as they can hide code. One result of this is that new chips may come with software interfaces that persist through different versions of the chip. In something of a paradox, the software interface is fixed while the hardware interface moves around.

Software has become more and more part of our daily work, and interfaces have become important to non-engineers. File formats are a good example of application interfaces that are important to all of us. The files I produced on my Mac Plus 25 years ago are still with me and usable; because of that, but you can read the Ph. D. dissertation I wrote using it. OpenOffice serves as a second-source for Word, and I can use either program with some assurance that I will continue to be able to do so into the future.

There's some backstory there. The "interchange format" for the original Word was "RTF". RTF is a reasonably good format, informed by Donald Knuth's TeX, but it was always a second citizen compared to the native "DOC" format. Microsoft published a spec, but they didn't follow it too closely and they changed it with every new release of Word. One result was that it was difficult to use Word as part of a larger publishing system (which I tried to do back in my days as an e-Journal developer). The last thing Microsoft wanted was for competition to Word develop before it grew to dominate the marketplace.

Cloud based software (software as a service) depends in a interesting way on application interfaces. Consider Google docs. You can send it a ".DOC" file created in Microsoft Word, do something with it, then export it. In a sense, Word is a "second source" for Google Docs, and consumers can use Docs without fear of lock-in. Docs adds its own web API so that developers can use it as a component of a larger web-based system. This is the "platform" strategy.

These new interfaces offer a user lock-in trade-off. While the customer gains the freedom to use a website's functionality with services from other companies, the control of the interface leaves the other companies at the mercy of the  company controlling the API. Developers coding to the interface are in the same situation as a second source chip supplier- always exposed to competition, while the platform provider becomes more and more locked in with every new component that plugs into it.

We now see a very interesting competition in platform strategies emerging. Apple's iPad/iPhone/iTouch software platform tries to lock-in consumers by opening an attractive set of API's for app development. It goes further, though, by attempting to control a marketplace (the app store) and imposing restrictive terms on app vendors. Google's Android platform tries to do the same thing in a much more open environment. Apple seems to have learned an important lesson, though. The biggest difficulty facing a company trying to plug into a platform is profitability, and the iPhone software marketplace appears to be offering viable business models for developers. It remains to be seen whether that condition will last, but it's clear that technology shifts are pushing services (such as phone service) that used to be stand-alone products into large, more complex ecosystems.
Enhanced by Zemanta

Article any source

Monday, March 1, 2010

eBook Pricing Calculus and A/B Testing

You've probably read about how book publisher Macmillan has won a big battle with Amazon over the pricing of ebooks. By shifting to an "agency" model, publishers will gain the ability to control the price that consumers pay for ebooks. A much discussed question has been whether this is really a win or a pyrrhic victory for publishers.

My question is a bit different. How will book publishers determine the correct pricing?

In Econ 101, we learned that markets set pricing by matching supply and demand curves. The publisher's task in the ebook economy is to find a price that will maximize their profits. Too high a price will result is low unit sales, while too low a price will leave money on the table.

The Blind Side (Movie Tie-in Edition) (Movie Tie-in Books)One of the frustrations you encounter trying to apply Econ 101 lessons to the real world is that you quickly find that most supply and demand curves are completely hypothetical. When W. W. Norton & Company set a retail price of $13.95 for The Blind Side (Movie Tie-in Edition) they didn't solve a set of equations that told them their profit would be maximum at this value. Norton doesn't know how many copies they would sell at $99.95, and they don't know how many they would sell at 99¢. It's likely they know how many total copies they're selling, but they probably don't have solid numbers telling them how many of those are selling at $9.81, the current price on Amazon.

But Amazon does.

Booksellers like Amazon can map out a large part of a demand curve using A/B testing. In A/B testing, website visitors are divided into two groups. The A group sees one version of a website and the B group gets another. The behavior of the two groups is then measured and compared. For example, the two groups could be shown different pricing for The Blind Side, and the rate that they purchase the book would be measured. Using repeated measurements of purchase rate vs. price, a dominant retailer such as Amazon is able to measure the consumer demand curve for a book or group of books. Pricing and profit can be optimized accordingly.

Amazon's pricing calculus will be somewhat different from the publisher's calculus, however. If the price they pay publishers is fixed (as it is for books), then the optimum price for Amazon will be higher that the optimum pricing for the publisher. You can do the math.

Amazon is well known for doing A/B Testing- see Bryan Eisenberg's description of the evolution of the Amazon shopping cart for a great example. Google is also notorious for depending on the technique. It even tested 41 shades of blue when it couldn't decide on a color for a design element.

Book publishers, on the other hand, have little experience with running e-commerce websites. A successful web merchant will optimize their site for search engine ranking, and will make it simple for users to find and get what they want.

Try a Google search for "The Blind Side". Since the book has become an Oscar-nominated major motion picture starring Sandra Bullock, it's not surprising that the top hits relate to the movie, not the book, but the complete absence of publisher results is striking. Here are the links my Google search pulls up:
    The Blind Side: Evolution of a Game
  1. Movie times
  2. IMDB (Amazon property, links to Amazon)
  3. the movie web site (Warner Bros., with move commerce links)
  4. Wikipedia (film)
  5. Google News Results (no book links)
  6. Google Image Search results (First one a book cover at AOL shopping)
  7. YouTube (official trailer) (no book links)
  8. Amazon page for the book At last, a place to buy the book!
  9. Rotten Tomatoes (movie reviews, no book links)
  10. Yahoo Movies (no book links)
  11. Apple iTunes Movie Trailers (no book links)
  12. Fandango (no book links)
  13. Moviephone (no book links)
  14. Google video search results. The second result is a link to a YouTube interview with The Blind Side Author Micheal Lewis, labeled "WW Norton: The Blind Side". It seems the publisher ponied up for some promotional video! But are there any links from the video to a book related page? Of course not!
On the second page of google results, the book gets a Wikipedia link and another Amazon link. On page 3, there's a book link to Powell's. On page 5, there's an  excerpt from the book on the NPR website.

Perhaps the publisher web presence for The Blind Side has been swamped by the movie pages.  If we add "book" to the search term we might expect to see a publisher presence for the book. On the third page of that search, there it is: a result from WW Norton. It's their home page, and no mention of The Blind Side at all. A message there tells me that WW Norton has
"recently relaunched our website, and many things have moved around. If you're looking for a book, try the search field above, or browse all books by subject." 
Oh, and when I search Norton for "the blind side", I find this page, which says the book is out of stock! If a competant merchant were running the site, it would tell me that the version without the movie-tie-in cover was in stock, but no such luck. However, there's a tiny link way on the other side of the page that says the book is available on the iPhone/iPod Touch iTunes App Store! Although no one has submitted a review on iTunes, I'm told I can buy it there from Kiwitech for $13.99.

There are so many things wrong with Norton's attempt at e-commerce that pricing is almost the last thing you would want to test with an A/B study.

So the funny thing about the shift to an "agency" model for the selling of ebooks is that the power to call the plays (set prices) now belongs to the one player (Norton, Macmillan, Random House, etc.) that has the poorest view of the ebook playing field; in fact, I'm not sure they all know the rules. The big huge left guard (Amazon) has just been benched even though he blocks like a superstar, because he's urged Norton to run the ball. Norton wants to pass the ball, to his stylish wide receiver, Apple, but the other team's blitzing, and a speedy right defensive end named Google is bearing down on Norton from his blind side.

I'm not sure I want to look.
Enhanced by Zemanta

Article any source

Monday, January 18, 2010

Google Exposes Book Metadata Privates at ALA Forum

At the hospital, nudity is no big deal. Doctors and nurses see bodies all the time, including ones that look like yours, and ones that look a lot worse. You get a gown, but its coverage is more psychological than physical!

Today, Google made an unprecedented display of its book metadata private parts, but the audience was a group of metadata doctors and nurses, and believe me, they've seen MUCH worse. Kurt Groetsch, a Collections Specialist in the Google Books Project presented details of how Google processes book metadata from libraries, publishers, and others to the Association for Library Collections and Technical Services Forum during the American Library Association's Midwinter Meeting.

The Forum, entitled "Mix and Match: Mashups of Bibliographic Data", began with a presentation from OCLC's Renée Register, who described how book metadata gets created and flows though the supply chain. Her blob diagram conveyed the complexity of data flow, and she bemoaned the fact that library data was largely walled off from publisher data by incompatible formats and cataloging practice. OCLC is working to connect these data silos.

Next came friend-of-the-blog Karen Coyle, who's been a consultant (or "bibliographic informant") to the Open Library project. She described the violent collision of library metadata with internet database programmers. Coyle's role in the project is not to provide direction, but to help the programmers decode arcane library-only syntax such as "ill. (some col)". The one instance where she tried to provide direction turned out to be something of a mistake. She insisted that, to allow proper sorting, the incoming data stream should try to keep track of the end of leading articles in title strings. So for example, "The Hobbit" should be stored as "(The )Hobbit". This proved to be very cumbersome. Eventually the team tried to figure out when alphabetical sorting was really required, and the answer turned out to be "never".

Open Library does not use data records at all, instead, every piece of data is typed with a URI. This architecture aligns with W3C web standards for the semantic web, and allows much more flexible searching and data mining than would be possible with a MARC record.

Finally, Groetsch reported on Google's metadata processing. They have over 100 bibliographic data sources, including libraries, publishers, retailers and aggregators of review and jacket covers. The library data includes MARC records, anonymized circulation data and authority files. The publisher and retailer data is mostly ONIX formatted XML data. They have amassed over 800 million bibliographic records containing over a trillion fields of data.

Incoming records are parsed into simple data structures which looked similar to Open Library's, but without the URI-ness. These structures are than transformed in various ways for Googles use. The raw metadata structures are stored in an SQL-like database for easy querying.

Groetsch then talked about the nitty-gritty details of data. For example, the listing of an author on a MARC record can only be used as an "indication" of the authors name, because MARC gives weak indications of the contributor role. ONIX is much better in this respect. Similarly, "identifiers" such as ISBN, OCLC number, LCCN, and library barcode number are used as key strings but are only identity indicators with varying strengths. One ISBN with a chinese publisher prefix was found on records for over 24,000 different books; ISBN reuse is not at all uncommon. One librarian had mentioned to Groetsch that in her country, ISBNs are pasted onto a book to give it a greater appearance of legitimacy.

Echoing comments from Coyle, Groetsch spoke with pride of the progress the Google Books metadata team has made in capturing series and group data. Such information is typically recorded in mushy text fields with inconsistent syntax, even in records from the same library.

The most difficult problem faced by the Google Books team is garbage data. Last year, Google came under harsh criticism for the quality of its metadata, most notably from Geoffrey Nunberg. (I wrote an article about the controversy.) The most hilarious errors came from garbage records. For example, certain Onix records describing Gulliver's Travels carried an author description of the wrong Jonathan Swift. Most of these errors come from garbage records, and when one of these is found, almost always, the same problems can be found in other metadata sources. Google would like to find a way to get corrected records back into the library data ecosystem so that they don't have to fix them again, but that there have been issues with data licensing agreements that still need to be worked out. Article like Nunberg's have been quite helpful to the Google team. Every indication is that Google is in the metadata slog for the long term.

One questioner asked the panel what the library community should be doing to prevent "metadata trainwrecks" from happening in the future. Groetsch said without hesitation "Move away from MARC". There was nodding and murmuring in the audience (the librarian equivalent of an uproar). He elaborated that the worst parts of MARC records were the free text data, and normalization of data would be beneficial whereever possible.

One of the Google engineers working on record parsing, Leonid Taycher, added that the first thing he had had to learn about MARC records was that the "Machine Readable" part of the MARC acronym was a lie. (MARC stands for MAchine Readable Cataloging) The audience was amused.

The last question from the audience was about the future role of libraries in production of metadata. Given the resources being brought to bear on the book metadata by OCLC, Google and others, should libraries be doing cataloguing at all? Karen Coyle's answer was that libraries should concentrate their attention on the rare and unique material in their collections- without their work, these materials would continue to be almost completely invisible.
Reblog this post [with Zemanta]

Article any source