Wednesday, February 29, 2012

Update on Inquirer Data

Well, I just got word that the Inquirer has decided to make their dataset on homicides in Philadelphia publicly available. Apparently they haven't settled on a general data policy, but this one is now accessible. You can find data on every reported homicide in Philadelphia between Jan 1, 1988 and December 31, 2011 here: https://www.google.com/fusiontables/DataSource?snapid=S4035208e94

Monday, February 20, 2012

Inquirer, Inquirer, let down your data!

So, I discovered last night that the Philadelphia Inquirer has put together a Google Fusion table containing a record for every homicide in Philadelphia county since 1988. I've used homicide data compiled by the Inquirer before to estimate the risk of homicide that normal Philadelphia residents have compared to UPenn affiliates. With 23 years of data, the possibilities to find all sorts of patterns are enormous. Homicide rate could be compared to economic indices, public policies, or climate even, and we could get some reliable results with a time depth like this!

But, the ability to export the data was turned off by the owner of the fusion table, by accident I assumed. I wrote to them about it, and apparently it is the Inquirer's policy to not let anyone access the data! They're concerned that someone might alter the data, and attribute it back the Inquirer. Here's the message I sent them when I heard about this.
I am a student at Penn, and that's why I'm interested in data generally. But I have no specific interest in the data related to my academic pursuits. I'm merely a concerned and interested Philadelphian who also has some quantitative know how.

I appreciate the sensitivity of the subject. In my own research, we spend a lot of time anonymizing interviews, and of course, it was a big issue with some of the Wikileaks data distributed by the NYT that it wasn't anonymized enough. However, is there precedent for altered data being hung around the neck of the original compiler? If there were an example case or two, your unease would make more sense to me. As it is though, since you are already maintaining the original data in a (relatively) publicly accessible way, it would be trivial for you, or anyone else, to demonstrate alteration or falsification of data attributed to the Inquirer.

The fact that you're already only distributing something which is publicly available from the PPD makes allowing public access to your compiled version even less risky. There are then two sources to turn to to verify the accuracy of data that someone attributes to the Inquirer.

My interest in this data spawns mostly from the fact that I'm a concerned Philadelphian with the necessary skills to analyze a data set like this. It looks like the Inquirer has done a great public service by compiling this data into a useful format from the various PPD reports. But it has only done so by a half measure so far, because the data is of no use when we can only look at the tables with our eyes. I'm also strongly influenced by the open data movement from within the research world. The best way to assert your confidence in your own research and analyses is to make the data openly available for anyone to recreate your results. Researchers who keep their data private are more and more looked upon with suspicion, and rightly so. The same goes for data journalism.

Moreover, there is a huge opportunity here for the Inquirer too. I am not the only person in Philadelphia who cares about data like this and knows how to analyze it. You have a forum to curate and display analyses and mashups contributed by your readers. The Guardian does something like this with their Data Blog http://www.guardian.co.uk/news/datablog, but frankly, the data sets they distribute are thin and uninteresting compared to what you could make available.

I hope you reconsider your data policy.
I'm frankly not too hopeful of a change of heart regarding making the data available. There's sure to be a lot more cases like this, of news organizations jumping onto the data journalism train, without really getting how it's supposed to work.

Friday, January 27, 2012

Distressing Numbers for Women

Sometimes I play with non-linguistic data sets recreationally. It's a totally valid hobby! I tend to gravitate towards data on the disparities between men and women, because gender equality is something that matters to me.

I've had this one data set for a while which I got from the Guardian Data Blog. It's 2006 data compiled by Unesco on men and women across a number of indicators. The ones of particular interest to me were student enrollment and estimated earned income. The student enrollment data is the percentage of potential students who are currently enrolled as students.

So, for each country for these two indicators, I calculated the ratio of Female/Male, to have one comparable measure. And then I took the log of the ratio, cause that's a good thing to do.

Before you look at the graph, make a guess. In countries with more gender equality in student enrollment, what do you think happens to gender equality in income?


The answer is nothing. And these are not all high income, high education countries either. These are global estimates, not just OECD countries.

On this graph, the red lines indicate total equality, a 1:1 ratio. What's especially striking about this graph is how many countries are cluster on the right of the red line. There are a lot of countries where more women are enrolled as students than men. But those countries have no better income equality on average than those countries with extreme education inequality!

This figure plots the density function (an estimate of how many countries are located at each point along the education dimension) and the cumulative density function (what percent of countries have at least that much equality or less).

In about 60% of the countries in the world, more women are students than men! The US is one of these. Maybe you've heard about it. They're calling it the "crisis of boys". Quite a crisis for boys, that on average they have about 90% the education, but 156% of the money.

I wonder what this means for the education panacea for world problems.

Disqus for Val Systems