Wednesday, June 20, 2012

Have you been in a Wawa's?

I would like to pre-empt all discussion here by saying that this blog post is strictly motivated by linguistics, and has no relevance to the US presidential election.

This video of Mitt Romney speaking about his experience in a Wawa has been floating around my newsfeed this week.


Full Disclosure: Wawa is my convenience market of choice.

What strikes me most about this video is the fact that Romney repeatedly said "Wawa's" even though the name of the store is just "Wawa" with no "s."  First, there's the linguistic issue of why this was such a natural mistake for Romney to make. Second, there's the sociolinguistic issue about why this particular mistake seems so egregious.

On the first point, there is clearly a strong tendency for store names to be formed in the possessive, indicating their ownership (or at least that's the origin). For example, "Macy's" was founded by Rowland Hussey Macy, "Wanamaker's" was founded by John Wanamaker, etc. However, not all stores which have names clearly formed in the genitive follow the ordinary orthographic rules for possessives.  For example, "Starbucks" is named after the Moby Dick character Starbuck, but the official name doesn't have an apostrophe. Similarly, JCPenney, which today isn't formed in the genitive, used to go by "Penneys" according to this logo from Wikipedia, also lacking the apostrophe.


Perhaps this is some kind of specialized "commercial genitive," I don't know.

At the same time, there are a lot of store names which are not formed in the genitive for some reason or another. One example someone brought up to me is "Nordstrom," which has no "s" even though it was founded by John Nordstrom. It's a little mysterious to me why this might be, except its original name was "Wallin & Nordstrom" (as in Carl Wallin and John Nordstrom), and coordination structures wreak havoc on everything. A similar story could be told for "Barnes & Noble."  In fact, the one kind of business that I know to be named after coordinated personal names are law firms (like "Dewey, Cheatum & Howe") which seem to never be formed in the genitive. There's also this blog post from Linguism which discusses the question of which store names get formed in the genitive and which don't, and he concludes that store names which are originally acronyms, like "Asda" and "Tesco" and foreign imports, like "Aldi" are less likely to be in the genitive.

At the same time, there is also a lot of asymmetric variation. People seem to be likely to form an officially non-genitive store name in the genitive, but not vice versa. How many of you would blink if someone said "I went shopping at Aldi's."? But no one would say "I went into a Starbuck."

Update: Ben Zimmer informed me on Twitter that the "Friendly Ice Cream" company officially changed their name to "Friendly's" perhaps because that's what all their customers called them anyway. According to this slideshow from the Boston Globe, the name change happened in 1989.

The point of all this is that Romney was wading into very muddy linguistic waters when he started talking about Wawa, and it's not surprising he screwed it up.

Which brings us to the second point: Why was saying "Wawa's" such a big deal? I just said that I wouldn't blink an eye if someone said "Aldi's" and that's basically the same kind of error. But, and I'm trying to speak here as a Philadelphian and Wawa devotee, not as a partisan hack, when I heard him say "Wawa's" my reaction was "Oh, he doesn't know how it works."

In some ways, my reaction was similar to how I feel when someone screws up the correct use of determiners in proper names. For example, if someone said to me "I looked it up on the Wikipedia," I'd immediately know they were uninitiated to the internet. Similarly, if someone said "they were uninitiated to Internet," I'd immediately know they were hopelessly ignorant.

I think what it comes down to is that where there is variation, there is complexity, and where there is complexity, the ability to successfully navigate complexity the right way is an important social signal that you are the right kind of person. Consider, for example, the needlessly complex language surrounding Twitter, and the communal paroxysm of self satisfaction when a politician says "I sent out a twitter to my followers," or refers to the service as "Tweeter."

I don't think that reaction, or the reaction to Romney saying "Wawa's," is fundamentally different from the dirty word in linguistics: prescriptivism. A lot of prescriptivism is specific discrimination against politically, economically and socially marginalized people, but a lot of it also comes out of nowhere, and just turns into a really complex game that people play for the sake of showing they can play it. So be cautious, fellow linguists, because today's "Wawa's" and "Tweeters" are tomorrow's split infinitives and passive voice.

Monday, June 18, 2012

Overplotting solution for black-and-white graphics

I'm working on producing some black and white graphics of data which has a lot of overplotting. There are three basic groups, which if I made the plot in ordinary full color ggplot2 would look like this (the code for the reverse-log x-axis is available in this gist, and the code for stat_ellipse() is available in this github repository).


For a black and white image, however, it's trickier. I don't usually find grey color scales to be sufficiently different for a plot like this, so I'd go for different point shapes. Unfortunately, the default shape scale in ggplot2 isn't very distinct in this case.
My first strategy to improve things was to add a custom shape scale, with alternating empty vs solid point shapes.
Better, but not great. All the overplotting of the empty point shapes creates this awful indiscriminate mash in the middle of the clusters.

My solution to this problem was to use filled points. While point shapes 1 and 5 in R correspond to an empty circle and an empty diamond, respectively, point shapes 21 and 23 correspond to a filled circle and a filled diamond, respectively, where the fill color and the border color can be different. So, I used shapes 21 and 23 instead of 1 and 5, and set the fill color to be white.
I think it's a big improvement. Here's one more iteration, filling the points with a light grey shade instead of white, just for some aesthetic appeal.

Thursday, May 17, 2012

On calculating exponents

In my post on the decline effect in linguistics, the question came up of how I've calculated the exponents for the Exponential Model in my papers. I think this is a point worth clarifying, but it's not likely to be interesting to a broad audience. You have been forewarned.

To recap as briefly as possible, in English, when a word ends in a consonant cluster, which also ends in a /t/ or a /d/, sometimes that /t/ or /d/ is deleted. This deletion can affect a whole host of different words, but the ones which have been of most interest to the field are the regular past tense (e.g., packed), the semiweak past tense (e.g., kept) and morphologically simplex words (e.g., pact), which I'll call mono. Other morphological cases which can be affected, and which I believe have occasionally and erroneously been categorized with the semiweak are no-change past tense (e.g., cost), "devoicing" (or something) past tense (e.g., built), stem changing past tense (e.g., found), etc. For the sake of this post, I'm only looking at the the main three cases: past, semiweak, and mono.

Now, Guy (1991) came up with a specific proposal where if you described the proportion of pronounced /t d/ for past as p, for semiweak as pj and for mono as pk, then j= 2, and k = 3. It is specifically whether or not  j= 2 and k = 3 that I'm interested in here. If you've calculated the proportions of pronounced /t d/ for each grammatical class, you can calculate j by log(semiweak)⁄log(past) and k by log(mono)⁄log(past). The trick is in how you decide to calculate those proportions.

For this post, you can play along at home. Here's code to get set up. It'll load the Buckeye data I've been using, and do some data prep.


So, how do you calculate the rate at which /t d/ are pronounced at the end of the word when you have a big data set from many different speakers? Traditional practice within sociolinguistics has been to just pool all of the observations from each grammatical class across all speakers.

So you come out with j = 1.91, k = 3.1, which is a  pretty good fit to the proposal of Guy (1991).

The problem is that this isn't really the best way to calculate proportions like this. There are some words which are super frequent, and they therefore get more "votes" in the proportion of their grammatical class. And, some speakers talk more than others, and they get more "votes" towards making the over-all proportions look more similar to their own. One approach to ameliorate this is to first calculate the proportion for each word within a grammatical class within a speaker, then for each grammatical class within a speaker, then within a grammatical class. Here's the code for this nested proportion approach.

All of a sudden, we're down to j = 1.34 and k = 2.05, and I haven't even dipped into mixed-effects models black magic yet.

But when it comes to modeling the proposal of Guy (1991), calculating the proportions is really just a mean to an end. I asked Cross Validated how to directly model j and k, and apparently you can do so using a complementary log-log link. So here is the mixed effects model for j and k directly.

The model estimates look very similar to the nested proportions approach, j = 1.38, k = 2.11.

What if we fit the model without the by-word random intercepts?

Now we're a bit closer back to the original pooled proportions estimates, j = 1.57, k = 3.19.

My personal conclusion from all this is that the apparent j = 2, k = 3 pattern is driven mostly by the lexical effects of highly frequent words. This table recaps all of the results, plus the estimates of two more model. One has just a by speaker random intercept, and a flat model, which looks just like the maximum likelihood estimate of the fully pooled approach, because it is.
Methodjk
Pooled1.913.1
Nested1.342.05
~Gram+(Gram|Speaker)+(1|Word)1.382.11
~Gram+(Gram|Speaker)1.573.19
~Gram+(1|Speaker)1.843.14
~Gram1.913.1

The lesson is that it can matter a low how you calculate your proportions.

Disqus for Val Systems