Showing posts with label privacy. Show all posts
Showing posts with label privacy. Show all posts

Sunday, July 21, 2013

NJ Senate Primary Could Become a Referendum on NSA Surveillance

Yesterday, Glen Greenwald, the journalist at the center of the Edward Snowden leaks, essentially endorsed Rush Holt for Senate, and promised to write more this week.

So let me give you some context about our crazy New Jersey Senate race. The open seat in the US Senate was created by the death of 89 year-old Sen. Frank Lautenberg, a Democrat. Our governor, Chris Christie, is a Republican seeking re-election in November (maybe the presidency in 2016). New Jersey is much more "blue in Presidential election years than it is in the off years when turnout is low. Christie didn't want to jeopardize his November chances by having a popular Democrat on the ballot to upset the usual voter turnout pattern, and so we have an odd situation where a Senate primary is being held in August. The voter turnout is expected to be minuscule, and Cory Booker has been considered a shoo-in.

Booker earned widespread admiration and media stardom by ousting a corrupt political machine and becoming the Mayor of Newark, New Jersey's largest city. His name recognition and ability to raise a huge amount of money in a short period of time is a huge advantage. He's already started advertising on cable TV; New Jersey is a very expensive market for politicians because the primary media outlets are New York and Philadelphia. He has 1.4 million followers on Twitter. Despite Booker's big lead, the unusual dynamics of this election mean that another candidate with a very motivated base of support could conceivably pull an upset.

Cory Booker faces three challengers in the Democratic primary. Two of them, Frank Pallone and Rush Holt, are Congressmen from the middle of the state, presumably attracted to a intra-term election that doesn't require them to give up their House seats. The other, Sheila Oliver, is the Speaker of the New Jersey General Assembly, and represents the district where I live. Like Booker, she's from Essex County, and there are a lot of people still upset about Booker's mayoral win.

Holt has put the privacy issue front and center on his website, and has vowed to repeal the PATRIOT Act. He also has the legislative record to do so, as he opposed both the extension of the PATRIOT Act and the FISA Amendments Act. His website features a petition and this statement:
As former chairman of the House Select Intelligence Oversight Panel, I know that these surveillance programs have compromised Americans’ rights while providing only the illusion of security. As a scientist who understands how these massive databases can be used and abused, I am frightened by what this near-universal surveillance suggests for the future of our democracy.
Cory Booker is triangulating a bit on the NSA. Here's what his website says:
I was deeply troubled by recent revelations of the scope of the National Security Agency’s domestic data collection. We failed as a nation to thoroughly debate and create public oversight before this highly questionable data collection began. It is time to bring this program to light and fix that error. 
It is a basic principle of our founding that laws be open to public debate and inspection. We must update the rules that permitted this program to exist and ensure Congress, the courts, and the people have access and oversight. We need to vigorously guard our 4th Amendment privacy protections while still protecting Americans from terrorism. There are serious questions about whether this program successfully does that, and we cannot ask these questions after the fact again. 
As "Liberty Equality Fraternity and Trees" on Daily Kos  points out, Booker "managed to talk about the NSA revelations without mentioning either (1) the PATRIOT Act or (2) the FISA Amendments Act. "

At a "Bloomberg View" lunch in New York, Booker talked about Edward Snowden. Buzzfeed reports:
As for Snowden, who Booker described as “a whistle-blower — however you want to call him,” he noted that “I don’t think we’d be having this conversation now” without him. 
He said Snowden had broken the law, and that fell short of true civil disobedience because he left the country. 
“It’s not heroic to be the person who stands up and you blow a whistle and then you sprint out of Dodge,” he said. “Stay here in your country — stand up.” 
Booker invoked “friends who smoke pot and one of them saying to me, ‘This is my civil disobedience, man.’” 
“Go smoke your joint in front of a policeman and get 100 of your friends to do the same thing,” he said.
There's nothing about the NSA on Pallone's website, but his recent votes have tracked Holts'.

Oliver hasn't campaigned much, but this is what her website says about the NSA surveillance program:
On privacy, Sheila Oliver believes that we must balance the needs of homeland security with protecting the privacy of law abiding citizens. She believes that some of the recent revelations made regarding National Security Administration (NSA) programs reveal that we’ve gone a step too far and that we must work to ensure that the privacy of law abiding citizens is protected.
This is my lawn.
Whichever Democrat wins, they'll face Steve Lonegan in a September election, and Lonegan has been vocal in his outrage at the NSA surveillance. (Lonegan faces only political novice Alieta Eck) At a recent event Lonegan responded to Booker's call for more discussion about the NSA program:
“We had that robust discussion…two-hundred and thirty-seven years ago,” Lonegan told a cheering audience .... “It was called the American Revolution.”

Notes

  1. Lautenberg's Senate seat is being kept warm by Jeffrey Chiasa
  2. Unaffiliated voters can vote in either primary, but the deadline has passed for Republicans to switch affiliation and vote in the Democratic primary, and vice versa.

Update July 26:

In the evening of July 24, the House of Representative voted on an amendment to a defense appropriations bill to defund the NSA's bulk data collection program. The amendment, co-authored by by conservative Republican Justin Amash and liberal Democrat John Conyers was narrowly defeated by a final vote of 205-217 with 12 not voting. The vote was characterized by highly unusual alliances. Rep. Holt voted for the amendment and Rep. Pallone did not vote. You can check how your own representative voted at http://defundthensa.com/ Blogger Jeff Jarvis repeatedly asked @CoryBooker for his position and got no answer. Rep. Holt then introduced his promised "Surveillance State Repeal Act".
Enhanced by Zemanta

Article any source

Friday, July 12, 2013

Datagate: Ecco come Microsoft collaborava con Prism e l'NSA


Nuove scottanti rivelazioni arrivano sulla vicenda Prism. In un articolo pubblicato sull'edizione online del quotidiano britannico Guardian, sono emersi alcuni dettagli su come Microsoft collaborava con l'NSA per il suo programma Prism. Ecco i dettagli emersi:

  • La Microsoft ha aiutato l'NSA ad aggirare la cifratura del nuovo portale Outlook.com al fine di intercettare le chat private scambiate via web
  • L'NSA aveva altresì ottenuto accesso alle email su Outlook.com, Hotmail e Live con un metodo che consentiva di accedere ai dati prima della cifratura.
  • FBI e NSA avevano un accesso facilitato, tramite Prism, a SkyDrive, il servizio di cloud storage di casa Microsoft
  • Nel mese di luglio dello scorso anno, nove mesi dopo che Microsoft ha acquistato Skype, l'NSA si è vantata che una nuova funzionalità introdotta in Skype il 14 Luglio 2012 gli aveva permesso di triplicale il numero di video chiamate intercettate attraverso Prism.
  • Tutto il materiale raccolto in questo modo era condiviso fra FBI, CIA e NSA.
Questo è quello che è emerso relativamente a Microsoft ma va ricordato che l'agenzia aveva stretto rapporti anche con Apple, Google, Facebook e Yahoo.

Mi vien sempre più da pensare che quel sant'uomo di Richard Stallman aveva visto giusto e che il suo estremismo era ben motivato
Voi cosa ne pensate?
Prima di concludere il breve post vi lascio con il video di una canzone il cui ritornello mi vien da canticchiare ogni qual volta penso a Prism....


 


Fonte notizia The Guardian.uk
Immagine via 9GAG
Any source

Tuesday, May 29, 2012

Quando si parla di privacy online tutto questo è "buono a sapersi"

Internet è ormai uno strumento indispensabile nella nostra vita di tutti i giorni, nel lavoro come nella vita privata. Ci sono tuttavia momenti in cui può suscitare delle preoccupazioni e non tutti gli utenti si sentono perfettamente preparati ad affrontarle. Per questo abbiamo sviluppato il progetto Buono a Sapersi, che ha l’obiettivo di aiutare le persone a navigare in piena sicurezza e a gestire con consapevolezza i propri dati online. 

Buono a Sapersi è un’iniziativa sviluppata in Italia in collaborazione con la Polizia Postale e delle Comunicazioni e si articola in quattro sezioni, due generali sul web e due più specifiche sui servizi Google. Il sito, che incorpora anche il Centro per la sicurezza della famiglia, è arricchito inoltre da numerosi contenuti video molto semplici e operativi.

Potrete imparare come navigare sicuri, avere informazioni sul modo in cui i vostri dati sono utilizzati su Google e su tutto il web e ottenere suggerimenti utili per gestire al meglio l'esperienza online della vostra famiglia: come scegliere una password sicura, cos’è un malware e come raddoppiare la sicurezza del vostro account Google con la verifica in due passaggi.

Chi non ha un amico che usa la stessa password per ogni servizio web, lascia il suo computer aperto senza bloccare lo schermo e pensa che una conversazione sui cookie sia un invito a colazione? L'alfabetizzazione digitale diventa ogni giorno sempre più importante, speriamo che anche voi possiate trovare interessanti i contenuti di Buono a sapersi e vogliate aiutarci a diffondere la conoscenza di questi temi sensibilizzando chi vi sta accanto.

Vi aspettiamo su www.google.it/BuonoASapersi per saperne di più.


Any source

Thursday, March 1, 2012

Le norme di Google diventano più semplici ma l'attenzione alla privacy resta immutata

Mi occupo dei progetti Google relativi a privacy e sicurezza dei dati fin da quando sono entrata a far parte dell'azienda, nel 2003. Nel corso degli anni, a mano a mano aumentavano il numero dei servizi offerti e quello degli utenti, ci siamo resi sempre di più conto di quanto sia importante guadagnarsi e conservare la fiducia delle persone. Questa consapevolezza ci ha portato a creare strumenti innovativi che consentono alle persone di mantenere il pieno controllo delle proprie informazioni e a lavorare con impegno affinché gli utenti comprendano facilmente i nostri impegni relativamente alla privacy.

A gennaio abbiamo annunciato l'imminente aggiornamento delle norme sulla privacy di Google e dal 1 marzo queste modifiche diventano effettive. Consci dell'importanza di modifiche di questo tipo, abbiamo intrapreso il più grande sforzo comunicativo nei confronti degli utenti mai attuato prima: gli utenti ricevono una notifica quando accedono al proprio account o quando utilizzano servizi come la Ricerca e abbiamo inviato una email a coloro che dispongono di un account Google. Se utilizzate Gmail, YouTube o qualsiasi altro servizio Google avrete senza dubbio sentito parlare di questo aggiornamento.

Come risultato, le modifiche alle nostre norme sulla privacy sono state al centro di una grande attenzione, che giudichiamo in modo positivo perché la privacy è importante. Tuttavia, ciò ha anche dato luogo ad alcuni fraintendimenti. La cosa più importante da tenere a mente è che facciamo tutto questo per rendere più comprensibile il nostro atteggiamento e i nostri impegni relativamente alla privacy e perché Google sia ancora più efficace per gli utenti.

Prima di tutto, la semplicità. Google nasce nel 1998 come motore di ricerca. Da allora abbiamo aggiunto una vasta gamma di servizi: Gmail, Google Maps, Chrome, Google Documenti, Android e Google+, solo per citarne alcuni. Al lancio di ogni nuovo servizio abbiamo introdotto nuove norme sulla privacy. Lo stesso ogni volta che abbiamo acquisito un servizio: mantenevamo le norme esistenti.

Così facendo, però, si è arrivati al punto in cui leggere tutte queste norme era diventata un’impresa interminabile. Per questo motivo, nel 2010 abbiamo compiuto un primo passo verso la semplificazione, raggruppando una decina di norme di servizi specifici nelle nostre Norme sulla privacy principali. Anche così, però, rimanevano ancora esclusi oltre 70 norme distinte. Il 24 gennaio abbiamo annunciato di aver riscritto con un linguaggio più semplice le Norme sulla privacy principali di Google e di avere consolidato in una le norme di oltre 60 servizi specifici. A partire da oggi è disponibile un unico documento completo che illustra l'impegno in difesa della privacy che adoperiamo nella maggior parte dei servizi Google.

In secondo luogo, il nostro obiettivo è creare un'esperienza migliore per gli utenti. Nella maggior parte dei casi, le nostre norme sulla privacy ci consentivano già di combinare le informazioni raccolte in relazione a un determinato servizio con le informazioni di altri servizi quando gli utenti erano connessi all'account Google. Microsoft, Yahoo e altre aziende del Web fanno altrettanto. Questo ci consente di trattare una persona che usa diversi servizi Google come un singolo utente quando è connessa.

Oggi, ad esempio, è possibile aggiungere immediatamente un appuntamento al Calendario Google quando in Gmail si riceve un messaggio che annuncia una riunione. È possibile condividere indicazioni stradali con una delle proprie cerchie di Google+ senza abbandonare Google Maps. O ancora, è possibile prendere direttamente l’indirizzo di email di qualcuno che si vuole invitare a condividere un documento su Google Documenti da Gmail. Tutto questo è semplicissimo e intuitivo e consente di risparmiare tempo.

Tuttavia, le nostre precedenti norme sulla privacy limitavano la possibilità di combinare le informazioni all'interno di un account per due servizi: la Cronologia web (cronologia delle ricerche per gli utenti che utilizzano il motore di ricerca quando sono loggati nel proprio account) e YouTube, un servizio che abbiamo acquisito nel 2007. Quindi, se un utente loggato cercava ricette di cucina su Google, le nostre vecchie norme sulla privacy non ci consentivano di suggerire video di ricette quando l'utente visitava YouTube, anche se era loggato e utilizzava lo stesso account Google per entrambi i servizi.

Ora la nostre Norme sulla privacy aggiornate rendono chiaro che quando un utente è connesso possiamo combinare le informazioni fornite in un determinato servizio con le informazioni di altri nostri servizi. Riteniamo che questo porterà a rendere più utili informazioni di vario tipo: risultati di ricerca, annunci e qualsiasi altro elemento di interesse per l'utente.

Il nostro approccio alla privacy non cambia. Queste modifiche non comporteranno la raccolta di nuove informazioni, non modificheremo le impostazioni della privacy delle persone né venderemo le informazioni personali dei nostri utenti agli inserzionisti. Desideriamo solo utilizzare le informazioni di cui siamo già in possesso per migliorare la loro esperienza come utenti.

Chi non ritiene che la condivisione di informazioni possa migliorare la sua esperienza come utente non è obbligato a eseguire l'accesso per utilizzare servizi quali Ricerca Google, Google Maps e YouTube. Se si è loggati, è possibile utilizzare i vari strumenti per la privacy di Google, ad esempio per modificare o disattivare la cronologia delle ricerche o la cronologia di YouTube, gestire il modo in cui Google personalizza gli annunci in base ai propri interessi o per navigare sul Web in incognito con Chrome. È addirittura possibile separare le proprie informazioni usando account diversi. In altre parole, è possibile utilizzare la gamma completa di strumenti Google che garantiscono trasparenza e controllo sui propri dati.

La nostra attenzione nei confronti della privacy è rimasta immutata e continueremo a cercare modi per aiutare gli utenti a comprendere e controllare le modalità di utilizzo delle informazioni che ci affidano. Giusto la scorsa settimana ci siamo uniti alla Casa Bianca e ad altre aziende del settore in favore della proposta "Do Not Track", ovvero l'aggiunta di un pulsante che offre agli utenti la possibilità di limitare a piacere le informazioni che possono essere monitorate durante la navigazione e abbiamo spiegato con chiarezza i controlli dei browser. Garantire trasparenza, controllo e sicurezza rimane fondamentale per conservare la fiducia degli utenti; è proprio per gli utenti che Google è stato creato e riteniamo che queste modifiche renderanno i nostri servizi ancora migliori.


Any source

Friday, June 17, 2011

“Io, io e ancora io”: con "Me on the web" la tua identità online è sempre sotto controllo

Strumenti come i social network o i servizi per la condivisione di foto rendono sempre più facile pubblicare informazioni personali online. Per tutelare la tua privacy, puoi ad esempio decidere chi nello specifico può visualizzare tali informazioni, indicando se devono essere visibili solo ad alcuni amici, a familiari o a chiunque sul web. Molto importante è però anche la scelta della modalità con cui si vuole essere identificati quando si postano queste informazioni. Ad esempio, puoi voler essere identificato con il tuo nome quando posti una risposta a una domanda in un forum, ma poter caricare sotto pseudonimo un video su un tema controverso.

Per supportarti in questo, abbiamo integrato all’interno dei prodotti di Google diverse opzioni d’identificazione. Tuttavia, la tua identità online è definita non solo da ciò che pubblichi tu, ma anche da ciò che altri utenti pubblicano su di te - che si tratti di una citazione in un blog, di un tag in una foto o di una risposta a un aggiornamento pubblico del tuo status. Quando qualcuno cerca il tuo nome su un motore di ricerca come Google, i risultati che appaiono sono infatti una combinazione delle informazioni che tu stesso hai pubblicato e di quelle pubblicate da altri su di te.

Oggi introduciamo un nuovo strumento che agevola il monitoraggio del tuo profilo sul web e consente un accesso semplice alle procedure per il controllo di quali informazioni su di te sono presenti sul web. Questo strumento, che si chiama Me on the Web, appare come una sezione della Google Dashboard sotto i dettagli relativi all’Account. Gli utenti più accorti già utilizzeranno gli Alert di Google per ricevere notifiche quando il loro nome o la loro email vengono citati su siti o news online. Me on the Web semplifica ulteriormente il procedimento e suggerisce anche automaticamente termini di ricerca che potresti voler monitorare. Me on the Web fornisce inoltre link a risorse che offrono informazioni su come controllare quali informazioni su di te vengono pubblicate online. Per esempio, linee guida su come entrare in contatto con il webmaster di un sito per chiedere la rimozione di un contenuto o su come pubblicare autonomamente ulteriori informazioni su di te per far sì che siti web contenenti citazioni meno rilevanti appaiano molto più in basso nei risultati di ricerca.
Questo è solo uno dei passi che quotidianamente compiamo per aiutarti a gestire la tua identità online al meglio e in modo sempre più facile.

Any source

Monday, August 23, 2010

Google e la privacy. Tra mito e realtà

Il tema della privacy online è oggi sempre più dibattuto. Per noi, che siamo un'azienda di ingegneri, privacy e tecnologia sono due aspetti che vanno di pari passo. Vorremmo cogliere l'occasione per fare un po' di chiarezza sull'argomento e distinguere tra mito e realtà.

Mito: Google sa chi sono e sa tutto di me

Realtà: Google ha l’obiettivo di creare servizi di valore, non di identificare i propri utenti. Le informazioni che raccogliamo hanno unicamente questo scopo. Ad esempio, nel momento in cui eseguite una ricerca senza avere effettuato login con un account Google, le informazioni che conserviamo nei nostri file di log (come l’indirizzo IP, browser e sistema operativo del vostro computer, ricerca compiuta, data e ora della ricerca, cookie id) non permettono di identificarvi personalmente ma servono per capire se i risultati forniti sono stati utili.
Per tutelare ulteriormente la privacy, cancelliamo una parte degli indirizzi IP dopo 9 mesi e anonimizziamo i cookies dopo 18 mesi. Solo chi ha un account Google é associato ad un nome,
ma é il nome che l’utente ha deciso di attribuirgli.
Ogni utente registrato ai servizi di Google può accedere a tutte le informazioni collegate al proprio account attraverso Google Dashboard (google.com/dashboard), una soluzione tecnologica che è anche la prova concreta della trasparenza verso tutti i nostri utenti.

Mito: Google ci scheda per fini pubblicitari

Realtà: Il modello di business di Google è basato sulla pubblicità, che ci permette di offrire molti dei nostri servizi gratuitamente. Anche in questo caso abbiamo un obiettivo chiaro: dare accesso ad informazioni sempre più rilevanti e utili ed esserne ripagati con un clic. Dal 2009 abbiamo introdotto anche un nuovo servizio: la pubblicità basata sugli interessi. Quest’ultima non effettua alcuna profilazione degli utenti ma si occupa di mostrare pubblicità basate sugli interessi di navigazione espressi tramite browser (che non identifica personalmente alcun individuo, anche perché lo stesso browser può essere usato da più persone e la stessa persona può usare più browser).
La pubblicità basata sugli interessi non può creare degli identikit perché non è associata ad un nome e neppure alle ricerche effettuate dagli utenti sul nostro motore.
Non consente di mostrare annunci pubblicitari associati a categorie di natura sensibile, come preferenze politiche, religiose, sessuali o informazioni di natura sanitaria. Le categorie sono determinate solo dalla navigazione su alcuni siti, quelli che mostrano le nostre pubblicità attraverso il programma AdSense.
Prima di lanciare questo servizio abbiamo voluto progettare delle soluzioni tecnologiche che garantissero la trasparenza e la libertà di scelta dei nostri utenti.
La risposta è data oggi dal pannello di controllo che permette di gestire le preferenze degli annunci associati al proprio browser (google.com/ads/preferences), aggiungere o rimuovere categorie ed effettuare, se lo si desidera, opt-out definitivo dal servizio. Questo pannello è raggiungibile anche mediante il link al nostro Centro Privacy posto sulla home page del motore.

Mito: E’ difficile tornare in possesso delle informazioni date a Google

Realtà: Il valore competitivo per aziende come la nostra é dato dalla fiducia degli utenti, che per Google non significa incatenarli ai propri servizi ma lasciarli in controllo dei propri dati.
Iniziative come Data Liberation Front (www.dataliberation.org) sostengono il diritto degli utenti di controllare le informazioni conservate nei diversi prodotti e servizi Google. Questo consente, ad esempio, di chiudere un account Gmail e trasferire i propri contatti su un altro provider di posta elettronica. Crediamo che la concorrenza stimoli l'innovazione e vogliamo che chi sceglie i nostri servizi lo faccia perché rispondono a dei bisogni.

Mito: Google vende i dati dei propri utenti alle aziende e li comunica ai governi

Realtà. Non cederemo mai a nessuna azienda le informazioni personali che possono identificare i nostri utenti, senza il loro consenso esplicito. Sulla home page del nostro motore è disponibile un link privacy attraverso il quale si può accedere a tutte le informazioni relative alla tutela dei dati personali e leggere che Google collabora con le istituzioni nella repressione del crimine informatico, rispondendo alle richieste di informazioni che sono formulate nel rispetto della legge.
Anche in questo caso la nostra è stata una scelta di trasparenza che si è concretizzata in un sito (google.com/governmentrequest) attraverso il quale é possibile avere informazione relativi alle richieste fomulate a Google dai Governi di tutto il mondo.

Siamo consapevoli dell’importanza assunta oggi dal tema della privacy nel mondo online ed é per questo che manteniamo un dialogo aperto con tutti voi e con le autorità preposte a tutelare questo diritto. Riteniamo che dialogo e tecnologia siano le risposte per proteggere la privacy online e permettere di avere pieno controllo sulle informazioni personali quando utilizzate i nostri servizi.

Any source

Friday, May 21, 2010

Bit.ly Preview Add-on Leaks User Activity; Referer Header Considered Harmful

Inside Intel: Andy Grove and the Rise of the World's Most Powerful Chip CompanyI've been reading a book called "Inside Intel" by Tim Jackson that reports the history of the chip giant up to 1997. At the end of the book, Intel is dealing with the famous flaw in the Pentium's division circuitry. Jackson observes that Intel's big mistake in dealing with the bug was to deal with it as a minor technical issue rather than as the major marketing issue it really was. If Intel's management had promptly addressed consumer concerns by offering to replace chips for any customer that wanted it rather than dismissing the problem as the inconsequential bug it actually was, it could have avoided 90% of the expense it actually incurred. The public doesn't want to deal with arcane technology bugs; they want to know who to trust.

This week FaceBook and MySpace had to deal with the consequences of obscure bugs that leaked personal subscriber information to advertisers. The Wall Street Journal reported that because Facebook and MySpace put user handles in URLs on their sites, these user handles, which can very often be traced back to a user identity, leaked to advertisers via the referer headers sent by browser software.

Reaction on one technology blog reminded me of Intel's missteps. Marshall Kirkpatrick, on ReadWriteWeb, called the the Journal's article "a jaw dropping move of bizarreness", going on to explain that passing referrer information was "just how the Internet works" and accusing the Journal of "anti-technology fear-mongering".

When a web browser requests a file from a website, it sends a bunch of extra information via http headers. One header gives the address of the file, which might be a web page, an image, or a script file. Other headers give the name of the software being used, the language and character sets supported by the browser. The Referer header (yes, that's how it's spelled, blame the RFC for getting the spelling wrong) reports the address of the page that requested or linked to the file. If the request is made to an advertiser's site, the Referer URL identifies the page that the user is looking at. When that page has an address that include private information, the private stuff can leak.

The controversy spurred me to take a look at some library websites to see what sort of data they might leak using referer headers. I used the very handy Firefox add-on called "Live HTTP Headers". I was astounded to see that a well known book database website seemed to be reporting the books I was browsing to Bit.ly, the URL shortening service! In another header, Bit.ly was also getting an identifying cookie. I went to another website, and found the exact same thing. This set off some alarm bells.

I soon realized that a report of EVERY web page I visit is being sent to Bit.ly. The culprit turned out to be Bit.ly's Bit.ly Preview add-on for Firefox. It turns out that for every web page I visit, this line of javascript is executed:
this.loadCss("http://s.bit.ly/preview.s3.v2.css?v=4.2");
This request for a CSS stylesheet has the side effect of causing Firefox to transmit to Bit.ly the address for each and every web page I visit in a referer header.

It's ironic. My last post described how URL shortening services can be abused for evil, but my point was that these abuses were a burden for the services, not that the services were abusive themselves. In fact, Bit.ly has probably done more than any shortening service to combat abuse and the Preview add-on is part of that anti-abuse effort. With Preview installed, users can safely check what's behind any of the short URLs they encounter by hovering over the link in question.

The privacy leak in bit.ly Preview is almost certainly an unintentional product of sloppy coding and deficient testing rather than an effort to spy on the 100,000 users who have installed the add-on. Nonetheless, it's a horrific privacy leak. There are other add-ons that intentionally leak private information, but typically they disclose their activity as a natural part of the add-on's functionality. One example would be GetGlue, which I've written about, and even Bit.ly preview cannot help but leak some info when it's doing what it's supposed to do (expand and preview shortened URLs).

I'm sure that Bit.ly will fix this bug quickly; their support was amazingly fast when I reported another issue. But a larger question remains. How do we make sure that the services we use everyday aren't leaking our info all over the place? The most widely deployed services- Google, Amazon, Facebook, etc. all deserve a  higher level of scrutiny because of the quantity of data at their fingertips. All the privacy policies in the world aren't worth a dime if web sites can't be held accountable for the effects of sloppy coding. It's high time for popular sites to submit to strict third-party privacy auditing, and for web users to demand it. It doesn't matter whether any advertisers actually used the personal information that Facebook sent them; what matters is whether users can trust Facebook.

It's also time for the internet technology community to recognize that referer headers are as dangerous to privacy as they are to spelling. They should be abolished. Browser software should stop sending them. The referer header was originally devised to help dispersed server admins fix and control broken links. Today, the referer header is used for "analytics", which is a polite word for "spying". The collection of referer headers helps web sites to "improve their service", but you could say the same of informants and totalitarian governments.

The pipe is rusty- that's why it leaks. We need to fix it.
Article any source

Thursday, January 21, 2010

Mostrare annunci pubblicitari più rilevanti in Gmail

Sin da quando abbiamo lanciato il servizio di posta elettronica Gmail, abbiamo cercato di mostrare annunci pubblicitari che non fossero intrusivi e che potessero essere rilevanti per l’utente, e lavoriamo costantemente per migliorare i nostri algoritmi e mostrare annunci che siano sempre più utili.

Quando si apre una email in Gmail, spesso si vede un annuncio pubblicitario associato al contenuto di quel messaggio. Diciamo che state leggendo un’email nella quale un albergo di Chicago (o di Roma) conferma la vostra prenotazione. A fianco del messaggio potreste vedere una pubblicità dei voli per Chicago (o per Roma, a seconda del caso!).

E’ importante ricordare che il meccanismo attraverso cui vengono mostrati gli annunci pubblicitari è basato su una scansione completamente automatica: per capirci, non c’è nessuno che legge i messaggi di posta elettronica per decidere quali annunci associare al messaggio. Si tratta dello stesso meccanismo di scansione automatica che la maggior parte dei servizi di email, e non solo Gmail, utilizza per filtrare lo spam o rendere possibile il controllo ortografico. Gli annunci pubblicitari sono selezionati in base a un criterio di pertinenza e vengono pubblicati automaticamente utilizzando la stessa tecnologia di pubblicità contestuale sulla quale si basa il programma AdSense.


A volte, tuttavia, non ci sono annunci pubblicitari sufficientemente pertinenti rispetto ad uno specifico messaggio. Da oggi, talora potreste vedere degli annunci che, invece di essere associati al messaggio che state leggendo, sono associati a un altro messaggio che si trova nella stessa pagina della Inbox del messaggio che state leggendo. Per esempio, state leggendo un messaggio in cui un amico vi augura buon compleanno; se non ci sono annunci pubblicitari pertinenti, potreste vedere visualizzato un annuncio che pubblicizza i voli per Roma e che è associato a quella email di conferma dell’albergo per Roma che si trova nella stessa pagina della Inbox.

Per mostrare questi annunci il nostro sistema non archivia nessuna informazione aggiuntiva, semplicemente seleziona un messaggio diverso con cui fare l’associazione contestuale. Così come non conserviamo alcuna informazione relativamente al testo del messaggio che state leggendo, non conserviamo nessuna informazione nemmeno relativamente al testo dei messaggi che sono stati riscansionati dal sistema per individuare un contenuto a cui associare un annuncio pubblicitario pertinente. Il processo è interamente automatico, non ci sono persone che leggono le email e né le email né informazioni personali vengono condivise con gli investitori pubblicitari.

Abbiamo aggiornato una delle voci del centro assistenza e alcune delle Domande frequenti nelle quali si specificava che la pubblicità mostrata a fianco di una email era associata solo al messaggio che si stava leggendo in quel momento. Non ci sono invece modifiche nelle Informazioni sulla privacy di Gmail. Abbiamo anche realizzato un breve video che illustra i cambiamenti apportati:



Il cambiamento verrà implementato nei prossimi giorni e grazie a questo ci auguriamo che gli annunci pubblicitari in Gmail risultino più interessanti: più annunci su argomenti a cui siete interessati e meno annunci non rilevanti.


Steve Crossan, Gmail Product ManagerAny source

Thursday, November 5, 2009

Una dashboard per gestire i propri dati in modo trasparente

Oggi abbiamo lanciato Google Dashboard, una nuova funzionalità attraverso la quale è possibile gestire tutte le informazioni associate al proprio account Google. Su Google Dashboard potete infatti disporre di un sommario con tutte le informazioni salvate dalle applicazioni che utilizzate, gestire e cambiare le impostazioni dei servizi in modo facile e veloce.

I dati salvati dalle varie applicazioni di Google sono diversi: ad esempio l’account di posta Gmail vi permette di salvare la posta ricevuta ed inviata, le bozze ma anche gli allegati e le conversazioni fatte attraverso la chat. Se decidete di attivare la funzione Cronologia Web, invece, vengono salvate le pagine web che avete visitato in passato, il che consente di ottenere risultati di ricerca ancora più personalizzati. La Dashboard raccoglie tutti questi dati in un unico formato, facile da usare e in grado di offrirvi un livello di accesso e controllo dei dati che ci auguriamo possa esservi d'aiuto.

Google Dashboard è stata sviluppata in Europa dal team di ingegneri di Monaco e Zurigo e oggi è accessibile in 17 lingue al seguente link google.com/dashboard o, in alternativa, tramite la pagina delle impostazioni dell’account di Google.

Any source

Wednesday, October 7, 2009

Street View e la privacy

Abbiamo appena caricato sul nostro canale YouTube un nuovo video che racconta il funzionamento di Street View, la tecnologia per il blurring e gli strumenti che vi permettono di tutelare la vostra privacy online.

Buona visione!





Any source

Thursday, September 24, 2009

Nambu Gets Better and Shortener User Tracking is Undermined


When I last wrote about the tribulations of tr.im and the business of bit.ly, our heroes had just stepped away from their nose-to-nose struggle, with Nambu founder Eric Woodward having announced the shut-down of tr.im, his URL shortener, only to vow its revival a few days later. Bit.ly, with a cozy relationship with Twitter, seemed to have taken a dominant position in the URL shortening business, whatever that turned out to be. I speculated that Bit.ly would use its position in the ocean of usage data to build psychographic profiles of users to help target advertising.

Since then, Woodward decided not to sell the tr.im business and has instead released the tr.im software as free open source, making it that much easier for websites to do their own URL shortening. He's also focused his company's attention on its Nambu Twitter clients (Mac and iPhone), the development of which suffered a major setback when his Chinese developers left for richer opportunities as soon as their contract was up. Nambu for Mac OS X has been my preferred Twitter client; it has a much more Mac-like user interface than others I've tried. When I updated my system to Snow Leopard last week, I was disappointed to find that Nambu had not survived the system update.

After unhappily revisiting Tweetdeck, I decided to try the beta version of Nambu, even though it's described as being not quite done. So far, it looks pretty solid. One change in particular pleased me, and that's the way the new Nambu works with URL shorteners. It seems that by surrendering URL shortening to bit.ly, Nambu is now freer to innovate in the user experience. Nambu now pre-expands all the shortened links so that the user can see the hostname that the links are pointing to. This has a number of consequences:
  1. The user can tell where a link will go. This will help avoid wasted clicks, and will help the user avoid spam and malware sites.
  2. Because all of the links are dereferenced before use, the URL shortening sites will no longer be able to track the user's reading preference. The business model I previously suggested for bit.ly will be defeated, and the user's reading privacy will be protected.
  3. The URL shortener will have to deal with an increased load. Nambu's going to make bit.ly work harder for the privilege of domination the URL shortening space.
Now I understand why bit.ly has been registering a bunch of instantaneous hits whenever I tweeted a link- it was robot agents, not people, that were clicking the links.

I was curious to see if Nambu was querying the URL shorteners directly or whether Nambu was trying to aggregate and cache the expanded links. I installed a nifty program called "Little Snitch" to see the outbound connections being made by programs on my laptop. It turns out that Nambu is doing a direct check for redirection on ALL of the links that it shows me, not just the shortened ones. Although this could break links that are routed as part of a redirect chain, I imagine that sort of link occurrs rarely in a Twitter stream.

The new behavior of Nambu and its effects on usage tracking points up a general problem faced by any system designed to measure and track internet usage. In my post on "bowerbird privacy", I mentioned that I use StatCounter to measure usage on this blog. StatCounter works quite well for now, but I imagine that its methods (based on javascript) might well stop working so well as web client technology evolves. That's one reason I expect that efforts to standardize measurements of usage in the publishing community, such as Projects "COUNTER" and "USAGE FACTOR" are doomed to rapid obsolescence.

Will bit.ly ever get a business model? Will Nambu find peace with the chilly kitty? Find out in next months installment of... As th URL Trns
Article any source

Friday, September 11, 2009

Public Identity and Bowerbird Privacy

My legal name is "Eric Sven Hellman". On Twitter, I'm using "gluejar". On Facebook, I took the username "eshellman", which I also use in a number of other places. For the most part, while I do try to separate my work from my personal life, I don't try to isolate my online identity from my "real world" identity. I used to use the identity "openly" in some circumstances for my work identity, but I sold that name as part of my previous company. My work identities can be easily connected to my personal identity, and as a result, the use of different identities affords me negligible privacy.

The very concept of privacy has changed a great deal over the last 20 years, in large part due to the internet-induced shrinkage of the world and the relentlessly growing power of large databases. Our traditional notions of privacy have had embedded within them an implicit equation of privacy with obscurity. Public documents with records of where we lived, what we owned, who we were married to and how much our house was worth would be available in the sense that anyone could go to a county registrar and get the information. If I walked to town, anyone who saw me and knew me would know where I was. In the not so distant future, it's not hard to imagine that an internet connected camera could see me, recognize my face, and post my whereabouts on the internet so that anyone searching for me on google could discover my whereabouts. It doesn't really matter whether that happens through Linked Data or by discovering my GPS coordinates on Twitter. As any computer security expert will tell you, security-by-obscurity is ultimately doomed to failure, and I'm pretty sure that the same is true of the privacy-by-obscurity.

Yesterday, Wired Magazine writer Evan Ratliff was found. As part of reporting an article about how hard is for someone to "disappear" in the digital age, Wired had offered $5000 to anyone who could track down Ratliff during 30 days starting August 15. Ratliff's downfall was partly that he "followed" a vegan pizza restaurant in New Orleans on Twitter. It should not be surprising that Ratliff was found, given that a Facebook group with 1,000 members formed in an effort to track him down, so the relevance to our everyday privacy is a bit tenuous. Hollywood celebrities are only too aware that privacy retention is much harder for famous people.

Partly inspired by Ratliff's article, I decided to do a bit of investigation of my own. If you've been reading blogs around the topic of e-book technology, you probably have encountered posts by someone that signs posts with the name "bowerbird". Bowerbird's posts are always on-topic, but they are written with oddly short lines, as if bowerbird was typing on a 40 character-wide terminal. Bowerbird's posts are often impolite and sometimes really insulting, and in several forums, the posts have provoked complaints of trolling or that bowerbird is "hiding behind a pseudonym". Bowerbird replies that "bowerbird" is his real identity "in many versions of reality". It's clear that bowerbird is an iconoclast. When bowerbird posted a comment on one of my recent posts, I decided to see what I could find out about him or her.

It turns out that "bowerbird" is really the first name of "bowerbird intelligentleman". He has used this as his professional name since at least the late eighties. The name is written with lowercase letters in the manner of e e cummings, and Mr. intelligentleman, as the New York Times might refer to him, is a performance poet, among other things. In fact, he claims to have started performance poetry as an art form in 1987, and was an early promoter of "poetry jams". In a charming, self-deprecating bio (PDF), he writes "bowerbird is also one of the world's worst poetry producers" and describes how his forays into computer typesetting of poetry magazines led him into the world of electronic publishing and ebooks. He was very active on the Project Gutenberg volunteer discussion list, where his talent for provocation prompted Marcello Parathoner to cathartically excerpt a collection of his postings. Much of his energy in the ebook arena was spent promoting his ideas about "Zen Markup Language" (z.m.l.) whose philosophy can be summed up as "the best mark-up is no mark-up". The short line endings in bowerbird's post appear to be his insistence on using z.m.l. for his posts. Or perhaps they're performance poetry. It's a cute idea, but personally, I find that the formatting makes the posts hard to read in their context.

When bowerbird posted his comment, he left digital footprints. He visited the blog on a link from LanguageLog. He lives in the Los Angeles area (he's posted elsewhere that he can be found in Santa Monica), uses Verizon DSL, and uses version 4.0 of Safari on the Mac as his browser. The blog uses statcounter.com to monitor usage, so a cookie has been placed in his browser so I can tell if he returns for a visit; bowerbird is able to control these cookies using privacy controls in Safari. DSL lines use a pool of IP addresses, so although I know the IP address he used, I can't use that IP address to persistently track him. However, StatCounter can follow him to other sites that use StatCounter. In principle, StatCounter could report his interest in my blog to other sites and perhaps even connect him to other identities he might have, which would bother me a lot and prompt me to stop using StatCounter.

What's interesting to me is that bowerbird has had an online public identity for over 20 years, and although his entire online life, warts and all, is open for examination (how many of us can say the same?) it appears as though he has successfully walled it off from his private life. Even if I go to the register of deeds in Santa Monica, I probably won't be able to discover whether he owns a house. I can't find out from fundrace if he has donated to a political candidate. I can find his cell phone number because he's chosen to post it, but I don't know anything he hasn't chosen to divulge. (He once owed 1-800-GET-POEM!) In the course of leading a poet's life, bowerbird has been living an experiment in public identity and privacy for 20 years!

I've previously written about the evolution and fluidity of personal names. The use of professional names for public identity is quite common in our society. Women who marry and take their husband's family name routinely retain their names professionally. Use of professional names is particularly common among authors, actors, and musicians. For them, the additional privacy afforded by the use of a professional name is particularly valuable. It strikes me that the separation and isolation of identities may become an essential privacy curtain even for people who aren't celebrities.

It's probably too late for me and most people of my generation. But "bowerbird privacy" could be a reasonable solution for the next generation. A significant number of my son's friends use Facebook under not-their-real-names, and I say more power to them. I think that privacy advocacy organizations should be working to put rules in place to prevent Facebook from enforcing its "only your real name" terms of service and prohibit companies such as Twitter and Google and Yahoo (and StatCounter) from working with ISPs to connect online identities with offline identies.

Nature's bowerbird gets its name from the bower, a structure that male bowerbirds construct to attract females. You might think of it as the bird's public identity. It's not sure why the females are attracted to the bower. Maybe it's privacy?

Reblog this post [with Zemanta]

Article any source

Friday, August 28, 2009

Third Wheels on Class Action Coffee Search


Although I'm not a lawyer, (IANAL) I've worked on a fair number of legal agreements. Most typically, I would spend a lot more time working on the agreement than I ever did consulting the terms of the agreement. That's because it's much easier to work out the hard core details of implementation without the lawyers reviewing everything.

In my post on privacy and google book search, I spent a fair amount of time looking at what the settlement agreement said about security and privacy. From the perspective of a few weeks later, it looks to me like most of the hard-core details that matter are in fact not in there and will need to be worked out by the parties to the agreement.

This morning, I took some vacation from my vacation and went over to Berkeley to participate in the a conference called "The Google Books Settlement and the Future of Information Access". I was struck by a comment that Jason Schultz, Associate Director of the Samuelson Law, Technology & Public Policy Clinic at U.C. Berkeley School of Law made at least twice. He said it was important to get more assurances on privacy into the settlement agreement because "at least then we'll have something on paper that we can enforce". As I mentioned before, IANAL, so the mention of having anything on paper that I'll have to pay a lawyer lots of money to attempt to enforce something else doesn't have a lot of appeal to me. And I have questions.
  1. Can third parties enforce "pieces of paper"? Suppose Joe and Mary, instead of hiring a divorce lawyer, sign an agreement to have coffee together once a week at Starbucks, can Starbucks sue them if one or both of them breaches the agreement?
  2. Can third parties sue to enforce a class action "piece of paper" that has been approved by a court? Suppose a class of baseball widows files suit against the class of Yankees tickets holders and gets a settlement that offers them a cup of coffee at Starbucks with their spouses once a week. Can Starbucks file suit to get better compliance with the agreement?
  3. Do opt outs affect the ability of class members to enforce the class action agreements? If the Yankee widows can opt out of the settlement, can they sue to get better coffee agreement compliance even if they opt out?
  4. If you are an author concerned about privacy, do you have more or less leverage on the devilish details if you opt in or opt out of the settlement agreement?
  5. Why am I so focused on coffee?

Article any source

Friday, August 14, 2009

Tr.im's Brief Demise and the Privacy Implications of the Bit.ly Monopoly


In case you missed it, last Sunday the URL shortener Tr.im announced that it was going to close. Then, on Wednesday, they announced that they were going to keep the service open. I spent yesterday morning listening to the TechZing interview with Tr.im and Nambu founder Eric Woodward, which was interesting in a lot of ways. His perspective on having worked with a Chinese development team was quite interesting, and though not relevant for this post, it explains why development on Nambu (an OS X and iPhone twitter client) has been slow in coming- the Chinese developers left to get rich making iPhone apps. What I was more interested in was the discussion of business models for URL shorteners in specific and the Twitter ecosystem in general. It was the lack of any plausible business model for Tr.im that led to the decision to close the service. According to Woodward, there are only three plausible business models for a legitimate URL shortener:
  1. you can charge users
  2. you can sell advertising
  3. you can sell data that you generate
and given that Bit.ly has an inside track with Twitter and offers everything to users for free, all of the three business models were just not going to work for Tr.im.

Part of the problem that Tr.im has experienced is that the cost of running a URL shortener scales with the amount of usage. (In a good web business, you have a high fixed cost and very sublinear scaling of cost with usage) Woodward says that he has to spend an hour a day just dealing with spam, and this problem is getting worse. Why is spam a problem for URL shortening? It's because spammers (and presumably phishers, too) will use URL shorteners to hide links to porn, scams, malicious content, etc. Afflicted users will report the links to the URL shortener's ISP host (Tr.im uses Rackspace) and the URL shortener will be shut down unless the spamming links are turned off. Another problem that afflicts popular URL shorteners is the problem of popular twitterers. When Ashton Kutcher or Shaquille O'Neill tweets a link to their millions of followers, the URL shortener can suddenly be hit with 10,000 hits per minute. A read-only website can handle this traffic easily by spreading service over multiple servers, but a URL shortener such as Tr.im, architected to generate dynamic usage data by writing to a MySQL database, needs beefy hardware to avoid getting overloaded.

In addition to the problems highlighted by Woodward, I see three core difficulties for URL shortener businesses:
  1. It's really easy to build a small-scale URL shortener. A good web developer could probably build one in a day. Because of this low barrier to entry, casual users are forever going to be able to find free URL shorteners.
  2. There are plenty of illegitimate (or at least annoying) business models for URL shorteners. These involve stealing traffic, stealing "Google juice", putting interstitial advertising in links, framing links, etc. This attracts entrants who make it harder for someone wanting to run a non-annoying business to attract paying users.
  3. Links need to be reliable, because if your shortener fails, the user doesn't get the content they are trying to access In a lot of applications, links are meant to keep working forever. Reliability and persistence are expensive.
So basically, URL shorteners are a high cost, tiny revenue business.

Nonetheless, URL shorteners are very useful in today's 140 character world. Woodward had concluded, and I agree, that there is no room for more than one URL shortener business in the Twittersphere, and that Bit.ly has won. Bit.ly thus finds itself with an odd sort of natural monopoly. Of the three plausible business models for Bit.ly, it seems to me that generating and selling data is the only one that would maintain the monopoly, and thus the business. What kind of data might Bit.ly sell? I'll place my bet on "psychographic data". With its URL shortening monopoly, Bit.ly has access to a huge number of clicks. Bit.ly knows who I am, because I signed up for an account. Whenever I click a bit.ly link, my browser sends a cookie to Bit.ly which it could be using to track my interests, what I read. Aggregated over all the people who click Bit.ly links, the dataset of who clicked what could be very interesting to advertisers. Just as Google has the ability to tailor advertising to me based on my search history, Bit.ly could use a psychographic profile of me to help advertiser do targetting. Bit.ly has an interesting advantage over Google, however. Because so many of its properties rely on the user's perception of Google as a company that can be trusted with sensitive information, Google is quite limited in how far it can go in tracking users. In contrast, Bit.ly is not inhibited in this way. In fact, Bit.ly's ability to profile users could make it even more attractive to people putting links into tweets. Bit.ly could even provide this sort of profiling without violating its privacy policy which promises that it
... discloses potentially personally-identifying and personally-identifying information only to those of its employees, contractors and affiliated organizations that (i) need to know that information in order to process it on Bitly, Inc.‘ behalf or to provide services available at Bitly, Inc. websites, and (ii) that have agreed not to disclose it to others. Some of those employees, contractors and affiliated organizations may be located outside of your home country; by using Bitly, Inc. websites, you consent to the transfer of such information to them. Bitly, Inc. will not rent or sell potentially personally-identifying and personally-identifying information to anyone. Other than to its employees, contractors and affiliated organizations, as described above, Bitly, Inc. discloses potentially personally-identifying and personally-identifying information only when required to do so by law, or when Bitly, Inc. believes in good faith that disclosure is reasonably necessary to protect the property or rights of Bitly, Inc., third parties or the public at large.
Ironically, URL shorteners could also be used in ways that enhance user privacy from a different direction. As I discussed in my post on the semantics of redirectors, most URL shorteners use HTTP redirects. Although it's browser-dependent, either 301 or 302 redirects will result in the originating page being sent in the referrer header. "META refresh" redirects, on the other hand, can be used to wipe (or replace) the value of the referrer header. (Unfortunately, these can be also used annoyingly to cause "referrer spam".)

Redirectors deployed for purposes other than URL shortening also have market-share related privacy implications. For example, the dx.doi.org redirector which handles most DOI traffic could be a very useful vantage point for industrial or technology espionage. Because this redirector serves so many scientific article links, a spy agency might be able to monitor everyone in the world doing research on nuclear fission or anthrax weaponization, to give two examples.

In preparing my post on privacy mechanisms for Google Book Search, I was struck at the many directions that someone intent on privacy intrusion could take to collect potentially sensitive information. Part of the skepticism I expressed about being able to "sell" privacy comes from a feeling that privacy as traditionally thought of is pretty much a lost cause on the internet, no matter what Google, Bit.ly, or anyone does. Somehow, traditional concepts of privacy need to be recast into something that people still value as they use the internet.
Reblog this post [with Zemanta]

Article any source

Tuesday, August 11, 2009

Shibboleth, Google Book Search, and the Hello Kitty Diary

I don't think I can sell privacy. In fact, it's hard to think of any technology that has succeeded in the market place because of privacy attributes. Swiss banks don't count as technology; strong encryption succeeded in the market for its security attributes rather than privacy attributes. (The same is true of locks on houses- its true that they provide privacy, but people use them against thieves, not against snoops. OK, maybe curtains, but the if there's a trade-off between privacy and style in curtain technology, style usually wins. Even the Hello Kitty Electronic Password Diary does not appear to be a big commercial success.


In my post on privacy and Google Book Search, I alluded to technological solutions libraries could use to enhance patron privacy while also protecting against unauthorized access. I thought it would be useful to elaborate on this comment with some details. In general, there is no reason that privacy and security objectives can both be met in a properly engineered solution, other than the fact that it's hard to find someone willing to pay for the properly engineered solution. For example, I mentioned Shibboleth as a possible solution to providing security and privacy. Shibboleth is an open-source single-sign-on authentication system developed as part of the Internet2 project. It uses strong cryptographic techniques to delegate trust over a network, and in so doing, allows for significantly enhanced privacy.

Think about the situation where a company has licensed some content to a university. The licensor wants to make sure that only persons associated with the university are allowed to access the content. It doesn't need to know who the user is, it only needs to know that the user is properly entitled. The Shibboleth system allows the institutional user to sign in to an authentication point once using their institutional credentials, then any licensed resource can check with the central authentication point that the user is accredited by virtue of institutional affiliation. Shibboleth also allows users an institutions to disclose attributes to providers of their choosing. Attributes might include their name, preferred language, subject areas of interest, subgroup membership, etc.. Security is preserved because the institution still knows the identity of the users, and is enhanced because the Shibboleth system is designed to be much harder to defeat than competing solutions.

As far as I understand, Shibboleth would not significantlyonly slightly enhance privacy in the specific scenario created by Google Book Search, where users have to be tracked as to how much of individual books they have viewed. However, a system could be built that distributes information over a network. Here's how it would work:
  1. When the user is authenticated by the institution, a session id would be sent to Google. The session id tracks the user, but only the institution knows the identity of the user.
  2. When the user views a page in a book, Google sends a message to the institution to increment a named counter associated with the user. The name of the counter identifies a book, but only Google knows which book is associated with the counter.
  3. when the user asks to view another page, Google asks the institution for the page count associated with the book and the user, and grants access accordingly.
Such a system works to enhance privacy by storing separately the identity of the person reading the book and the identity of the book. Only if Google and the institution agree to exchange information can the reading history of an identified patron be revealed. This results in much stronger privacy even than we have in the print world. A government request for a patron's GBS reading habits would have to be made to two separate entities, probably in two different jurisdictions.

What is the likelihood that such a system can be created and adopted? On this score I am very skeptical. Who would pay for the enhanced privacy afforded by such a system? The success of a variety of Web 2.0 services seem to indicate that users are almost eager to give up privacy to gain the ability to communicate. As Randal Picker has discussed in a recent paper, consumers have significant incentives to give up their privacy to online advertising networks because doing so amounts to advertising by the consumer that results in a more efficient market. The history of Shibboleth can be used as an indicator of market behavior. Although it can provide enhanced privacy and strong security, these advantages have not been able to counteract implementation and usability costs and compared to competing technologies, and Shibboleth has not been widely adopted. When Peter Brantley raised the specific question of using Shibboleth for Google Book Search, Google's Dan Clancy commented that "Some institutions use Shiboleth and we will support this although most institutions prefer IP authentication". Google is known for putting a very high priority on usability, which is an area of significant weakness for Shibboleth.

On second thought, maybe the Swiss banks are onto something. Maybe the best target market for ultimate privacy is ultra rich people. Sergey and Larry, Warren and Bill, might I sell you a bit of privacy?
Reblog this post [with Zemanta]

Article any source

Friday, August 7, 2009

What the Google Books Settlement Agreement Says About Privacy

Here's what Google thinks I'm interested in:
  • Computers & Electronics - Enterprise Technology - Data Management
  • Computers & Electronics - Software - ... - Content Management
  • Entertainment - Movies
  • Finance & Insurance - Investing
  • Internet - Web Services
  • Lifestyles - Clubs & Organizations
  • Lifestyles - Parenting & Family
  • News & Current Events - Technology News
  • Reference - Libraries & Museums
  • Social Networks & Online Communities - Social Networks
If you want to find out what Google thinks you're interested in, go to http://www.google.com/ads/preferences/view and find out. Does anything there disturb you? Can you imagine items that might appear on your list that would disturb you, or that you wouldn't want to post on your Facebook profile or in a comment to this post?

Last Friday I participated in a workshop sponsored by Harvard's Berkman Center focusing on Google Books and Google's settlement agreement with authors and publishers. The meeting was very well tweeted, so I won't bother to summarize or comment on the blog, at least for now. In the afternoon, I participated in a breakout session on privacy issues facilitated by Marc Rotenberg from EPIC. Recently, the Electronic Frontier Foundation (EFF) has focused attention on the neglect of patron privacy in the settlement agreement. At the NYPL panel I attended, the view expressed by the participants was that though privacy was a very important issue, the settlement agreement was not the place to address privacy concerns. At the Harvard workshop, the opposite view was predominant.

In the EFF posting, I was struck by by the fact that they suggest that Google should be required to
allow users of anonymity providers, such as Tor, proxy servers, and anonymous VPN providers, to access Google Book Search
but they don't seem to expect libraries to be participating in the digital environment for books. In fact, I doubt that many libraries today view themselves as potential anonymity providers, despite the deep-seated respect for patron privacy that is part of the inherited culture of librarianship. I had wondered whether libraries would be able to use technological means, such as proxy servers, to ensure the privacy of their patrons who use Google Book Search. With some inspiration from the workshop, I've spent some time closely examining the agreement to see what it really says about privacy and what libraries might be able to do to enhance patron privacy.

Nowadays, electronic resources librarians can't help but focus more concern on monitoring for misuse of resources than on patron privacy issues. The obligation to do so is built into most license agreements for electronic resources. The following passage is from section 5.2 of the CLIR/DLF Model License:
Protection from Unauthorized Use. Licensee shall use reasonable efforts to inform Authorized Users of the restrictions on use of the Licensed Materials. In the event of any Authorized User makes an unauthorized use of the Licensed Materials, the parties may take the following actions as a cure:
  1. Licensor may terminate such Authorized User's access to the Licensed Materials;
  2. Licensor may terminate the access of the Internet Protocol (“IP”) address(es) from which such unauthorized use occurred; or
  3. Licensee may terminate such Authorized User’s access to the Licensed Materials upon Licensor’s request. Licensor shall take none of the steps described in this paragraph without first providing reasonable notice to Licensee (in no event less than [time period]) and cooperating with the Licensee to avoid recurrence of any unauthorized use.
Frequently libraries have to negotiate to get this language, as publisher licenses frequently have more burdensome requirements.

The sorts of things that actually happen, and librarians worry about, are of two types. The first is when a student with legitimate credentials "loans" them to a friend and in a week or two thousands of "friends" (often in another country) are using a resource through the campus proxy server. The librarians obligation is to identify and disable the rogue credentials. In the other scenario, a student or faculty member tries to use some type of downloading tool to download an entire journal or database for some sort of offline use. In this case as well, the offending user must be identified and told "don't do that". Publishers of electronic resources typically have monitoring tools (and bot traps and poison pills) in place so that they can detect such misuse and shut off a customer's access when this sort of thing occurs. A call to a support desk is typically needed to restore access. Publishers realize that these things happen, and that on a campus with 20,000 students, there are limits to how much librarians can control what their patrons do. I do not know of any case where legal proceedings or demands or compensation have resulted from such incidents, but I do know that one publisher cut off access to all of China for several months when a breach occurred there.

There clearly exists tension between a library's obligations to prevent unauthorized use and its obligations to protect the privacy of users. In the CLIR/DLF model license, there is a mutual obligation that balances the licensee obligation to control unauthorized use:
Confidentiality of User Data. Licensor and Licensee agree to maintain the confidentiality of any data relating to the usage of the Licensed Materials by Licensee and its Authorized Users. Such data may be used solely for purposes directly related to the Licensed Materials and may only be provided to third parties in aggregate form. Raw usage data, including but not limited to information relating to the identity of specific users and/or uses, shall not be provided to any third party.
The balance between providing for security against unauthorized use and confidentiality of user data is the practical determinant of the degree to which a patron can expect to have real privacy. To provide security against unauthorized of electronic resources, a library needs to generate logs for any proxy servers that it operates. To assure patron privacy, a library must be diligent to limit the retention of those log files and of any other records that might be used to identify and track users and their usage of particular resources.

The settlement agreement (available here) says very little about patron privacy. (In fact, the only users whose privacy is mandated are users with print disabilities who access Library Digital Copies in Fully Participating Libraries under the special access provisions of section 7.2(b)(i) of the agreement.) (I'm capitalizing terms defined in the agreement.) It says quite a lot about security, however, and thus many aspects of patron privacy will be effectively governed by the provisions for security. The use of proxy servers by libraries is implicitly mentioned in two places. In section 4.1(a)(iv) pricing bands are specified for government, public and school library subscriptions with the qualifier "no remote access without Registry approval", while higher education and corporate pricing bands are specified without the remote access qualifier. Remote access is most typically provided by libraries in higher education through the use of proxy servers. Additionally, the "Security Standard" set out in Appendix D to the agreement specifies that
Google shall use commercially reasonable efforts to authenticate individual End Users for access to Books in an Institutional Subscription by verifying that an individual is affiliated with an institution with an active subscription. Google’s efforts will be in partnership with the subscribing institutions in a manner consistent with, or otherwise equivalent to, generally accepted industry standards for authentication of use of subscriptions. Techniques used may include IP address authentication, user login, and/or leveraging authentication systems already in place at an individual institution.
Since the current "industry standard" is to allow users to authenticate through a proxy server against an institutional id/password service, it would seem that proxy servers would be permitted under the agreement, at least for higher education settings.

There is a specific security requirement set by the agreement that is likely to result in increased user tracking by Google. Google is required to make sure that each user cannot preview more than a certain number of pages of a book. Thus, Google must keep track of the books that a user has viewed, and stop the preview once the quota is reached. For the purposes of this requirement, Google is supposed to treat multiple users of a given computer as a single user. Assuming it is possible do so, this would have some odd consequences for computers in a library. A patron would be able to move from computer to computer and view more than their quota, but might not be able to view any pages from a book popular enough to have been previously viewed by another patron. In the current version of Google Book Search, cookies, not IP addresses, are used to track users, but a user is not required to log into the service at all unless they want to access personalization features. Google sets a 2-year cookie when you use the service, but the service can be used without cookies. To fulfill the terms of the settlement agreement, it appears to me that it's likely that Google would have to either require users to log into personally identifiable accounts, or to use IP addresses of individual computers to allow unidentified users access the service. Either way, libraries would be limited in their ability to use proxy servers to protect patron privacy (for example, by blocking cookies), and it's quite clear that what EFF has proposed with respect to anonymity providers is incompatible with the agreement.

I do not know of any resources currently licensed to libraries that are comparable to the post-settlement Google Book Search in the requirements for user tracking to prevent excessive uses, so it's not clear to me how much guidance is really given by the phrase "generally accepted industry standards for authentication of use of subscriptions". Authentication methods less widely deployed, such as Shibboleth may provide more patron privacy than use of cookies and IP addresses and/or proxy servers, while at the same time allowing Google to satisfy the terms of the settlement agreement. It is also likely that authentication technology, or modifications of existing authentication technologies, could be developed and specifically tailored to meet both the security requirements of licensors and the privacy requirements of libraries.

It is worth noting that the requirement for user tracking is not found in main part of the settlement agreement, but rather in "Attachment D", the Security Standard. Interestingly, the settlement agreement includes a provision for the Security Standard to be reviewed and revised every two years, by "Google, the Registry and up to a total of four (4) representatives on behalf of the Fully Participating Libraries" to allow for changes in Technology. Note that although the libraries are included because of their role in allowing use of library digital copies, there is a single Security Standard which applies to both Google and library-provided services. Thus there will be four library representatives who must agree to revisions in the security policy (and thus on the privacy that it allows) to be implemented in Google services. In theory at least, libraries could use the review of the Security Standard to introduce use of security/privacy technologies suited to the special characteristics of Google Book Search Subscriptions.

It is unclear what security requirements apply to the "Public Access Service" which would put free terminals (with paid printing available) in any US library that wanted it, because the settlement agreement treats the Public Access Service as something separate from the Institutional Subscription, while the Security Standard makes no mention at all of the Public Access Service. It seems possible that adding coverage of the Public Access Service to the Security Standard would also have to be addressed by the security review group that includes the library representatives.

In any case, it is clear that the power of the Book Rights Registry to review and approve the security implementation plans of Google gives it a great deal of leeway to set standards for patron privacy. Since the primary duty of the Registry is to serve rights-holders, its intrinsic motivation for protecting privacy would be only to see that privacy intrusions do not act to depress revenue significantly. Strong oversight by the court such as has been requested by the library associations may also promote attention to privacy concerns. Finally, it is likely that the Registry will need to pay close attention to state patron privacy laws. The library-registry agreements explicitly allow for state laws to trump any library obligations under the settlement agreement, so there can be no provision of the security standard that is incompatible with state privacy laws.

Google, as presently constituted, has every reason to be concerned about user privacy and guard it vigilantly; its business would be severely compromised by any perception that it intrudes on the privacy of its users. As Larry Lessig pointed out at the Berkman workshop, that doesn't mean that the Google of the future will behave similarly. Privacy concerns should be addressed; the main question has been how and where to address them. My reading of the settlement agreement is that it may be possible to address these concerns through the agreement's Security Standard review mechanism, through oversight of the Registry, and through state and federal laws governing library patron privacy.

And I am still not a lawyer.
Reblog this post [with Zemanta]

Article any source