Sunday, April 13, 2014

Baby Naming Trends: Now With More Linguistics!

This animated graph about the rise in boys names ending in <n> has been making the rounds lately.


It comes from this blog post by David Taylor.

It's a really cool graph, but then, I tend to find analysis of baby names a bit frustrating because they almost always rely strictly on the written, or orthographic, forms of the names. It's not that the way people spell their children's names doesn't matter, but it's half of the puzzle. For example, I'm named after my grandfather. He was German (more specifically, a Donauschwob), so he spelled his name <Josef>, and pronounced the initial sound like <y>, which in the IPA is /j/. When naming me, my parents had a whole bunch of options. Would the pronounce my name like my grandfather did, or like most English speakers would? And how would they spell it? They wound up settling on the English pronunciation, and the German spelling. I've made a little diagram displaying a very partial set of options my parents had in choosing my name.


And of course, Sarah Jessica Parker played a woman named /sændi/ who spelled it <SanDeE☆> in Steve Martin's LA Story, so clearly the spelling of proper names is an important expressive dimension, but still just half the picture.

So, I decided to look at a bit more at popular linguistic structures in baby names. Hadley Wickham has already compiled the top 1000 baby names in the US per year since 1880 (https://github.com/hadley/data-baby-names), and Kyle Gorman has a nice python module that syllabifies CMU dictionary entries (https://github.com/kylebgorman/syllabify). So I put together some sloppy code to analyze it (https://github.com/JoFrhwld/names). The biggest weakness to my approach is the number of names which are not to be found in the CMU dictionary. 2525 out of the total 6782 names in the data (about 40%) aren't in CMU, so this post should be understood as being for entertainment purposes only.

One other thing that bugged me about the name final <n> plot is that it seemed kind of arbitrary to focus on the final letter of the name. I suspect that it's a real trend that people noticed eyeballing lists of names, but that it wasn't compared against other kinds of trends. I went ahead and labeled name initial and name final syllables, codas, onsets and rhymes as being special, but I'm not going to single them out.

Kicking things off, there's a graph of popular syllables between 1880 and 2008. To be included in the graph, a syllable had to be in the top 3 most popular in any given year. The y-axis is how many times more frequent the syllable is than if syllable selection were random. It's not frequency rated, that is, this is just the distribution over names that have that syllable, not babies.
It's a bit chaotic, I know. It's a time like this that I wish I'd learned a little JavaScript so I could make an interactive version with brushing. Here's another version where each syllable gets a facet. They're ordered by their decreasing maximum ratio.
It looks like at the syllable level, name final /nə/ and /li/ for girls are both long time favorites, as well as more popular syllables than any boy's name final /n/ syllable. The most popular boy's name final /n/ syllable looks like it's always been /tən/, but maybe it's flagging a bit compared to the recent surges in /sən/ and /dən/. It also looks like popularity in syllables is pretty evenly split between name initial and name final syllables. For both boys and girls, some kind of initial between /e/ ~ /ɛ/ ~ /æ/ is pretty popular, but I can't be sure what's going on there, because the CMU dictionary has the same entry for both <Aaron> and <Erin>.

But maybe the reason boy's name final /n/ isn't shining through like you might expect is because of phonological reasons. A boy's name ending in a word final syllabic /n/ is necessarily going to pull the preceding consonant into the syllable with it. Looking at the plot above, it's not likely that the preceding consonant is totally random either, cause we've only got /t, s, d/ (all coronals) and vowels preceding the /n/. But for the hell of it, here's the same kind of plot as the ones above, but this time with syllable rhymes.
There's a lot less volatility in the rhymes data, probably because there's fewer different kinds of syllable rhymes. Complex rhymes don't seem to be that popular ever. We've mostly got vowels from open syllables, and syllabic consonants. At any rate, the popularity of name final /n/ for boys is pretty clear, taking over from /i/ (from names like Billy and Jonny). The boy's trend towards name final /n/ seems to be about on par with the trend for girls names to end in /ə/.

I'd like to play around with this data a bit more if I get some time. It occurred to me that you could come up with a few different ways of generating popular names from different eras by randomly sampling popular syllables, or by estimating transition probabilities between syllables and going on a random walk that way.

All my code and the data are up on github, if anyone else wants to play around with it: https://github.com/JoFrhwld/names.

Thursday, January 23, 2014

When did Americans start spelling it "color"?

Usually I wouldn't apologize for a lapse in posting, since I think an obligation to apologize acts as block to actually making another post. But in this case, my hiatus is (loosely) related to the topic.

In September, I defended my dissertation (check it out here if you're so inclined), then hopped on a plane and immigrated to Scotland to start a job as a lecturer in Sociolinguistics at the University of Edinburgh!

The sun was in my eyes. I was feeling very excited.
It's been a fun experience exploring the new cultural landscape. As is usually the case, the differences are more noticeable and surprising than the similarities, but I'm managing more or less. I'm crossing streets with confidence that I know which way the cars are coming from, participating in rounds at pubs (though I'm still not sure I'm doing that right), and thanking the bus driver, which is nice enough to do in the States, is apparently more obligatory here.

One thing I was told by another recent US immigrant to the UK wis that that students appreciate attempts at accommodation to British spelling. So, I'm giving that a shot too, although I'm sure I slip up a bunch, since I'm a very poor copy editor. However, out of passing curiosity, I decided to look at the historical trends in these spelling differences in the Google ngrams, and the patterns seemed interesting enough to warrant this blog post.

First, looking at the American English data for <color> and <colour> there's a very nice and clear cross over from to around 1845, which is about 20 years after Webster's 1828 dictionary, which according to Wikipedia is what we have to blame these differences on. Once you start trying to add more <-or ~ -our> to the graph, it gets chaotic fast, so instead of plotting out each word, I'll plot out the percent of <-or> spelling using Google ngram's handy arithmetic functions. (I've also included color/color, just to anchor the top of the range at 100%) I don't think I should have been, but I was a bit surprised with the uniformity with which the <-or> spelling replaced the <-our> spelling across all of these words. The trends seems to kick off around the 1820s, consistent with blaming it on Webster's dictionary, and increased till reaching its plateau around 1860.

But of course, <-or ~ -our> spelling isn't the only difference between British and American systems. The next set of consistent spelling differences involve <-er ~ -re>. Here are those words plotted out, with <color ~ colour> and <humor ~ humour>left in there as a representative items of the <-or ~ -our> set. So, it seems like there is a similar uniformity within the <-er ~ -re> words (maybe saber and theater are lagging behind) but the <-re→-er> replacement is offset from the <-our→-or> by about 60 years or so.

Of course, I shouldn't have been surprised, because I know how a little bit about language change, but it was fun to see this thing that I think of being a uniform "American Spelling" is actually the result of multiple changes that didn't happen all at the same time.

Just for fun, I took a look at what these patterns look like in British English. So, it looks like there might be a bit of a creep of American spellings into British English, but interestingly, the particular alternations aren't differentiated. So while the end product of "American Spelling" appears to be the result of an accumulation of different changes, the borrowing of American spelling into British English is being done holistically.

Thursday, June 20, 2013

The relationship between Kanye, Rap and African American English

This started out as an update to my post on Kanye West's song "I am a God," but wound up being nearly as long as the post itself, so I'm separating it out. But look over that first, or maybe keep it open in a separate tab.

A concern has been conveyed to me that I may have been equating rap as a lyrical form, African American English, and Kanye West in a problematic way. I wasn't really clear about my assumptions about how these three things are related, so I'll try to clarify.

So first, I definitely don't want to imply that the conventions of what is possible and not possible in rap is equivalent to the grammar of African American English. Rap is strongly identified as an African American art form such that people frequently malign hip hop as code for AAE, but as a linguist I know better than to draw similar equivalencies. As a lyrical and musical form, rap has its own conventions which aren't the same as AAE grammar. This must be the case, because speakers of other dialects and languages can produce songs which are clearly identifiable as rap!

I am assuming that Kanye West is a speaker of AAE as it spoken in Georgia, and that his phonology, as he acquired it, generates a bunch of representations. That's why I tied in the Labov, Cohen & Robins (1968) reference, to try to emphasize that these ∅ coda Go(d)s weren't just a quirk of Kanye, but rather reflective of larger dialectal trends in which Kanye is a participant. What does it matter? It's just more interesting if the reasoning I proposed here could generalize beyond just Kanye.

Next, I assume that Kanye has a personal filter, partially due to his personal taste, partially due to the conventions of rap, whereby he decides whether two words work as a rhyme. My reasoning is that if we understand Kanye's filter, and can see what comes out the other end, then we can make some assumptions about what went into it in the first place. Importantly, Kanye's rhyme filter is not AAE. In AAE, the set of words {God, Go(d), massage, ménage, garage, restaurant, croissants} definitely aren't perfect rhymes, but they all passed through Kanye's filter as rhyming equivalent.

My reasoning is as follows. All of these words passed some metric of Kanye's rhyming filter as being equivalent enough. The zero coda variant "Go(d)" is a product of Kanye's AAE phonology. If we can figure out what Kanye's filter is, then we can know something about the phonological status of "Go(d)",  which can tell us something about this feature of AAE phonology. Excluding "Go(d)", all of the final syllables of the words in this rhyming scheme have 1) a low-back vowel 2) a coronal obstruent. The coronal obstruents vary quite a bit in the sub-coronal place of articulation, their manner, and in their complexity. So, I want to conclude that Kanye's filter requires matching on the vowel quality, and the major place of articulation of the coda, but not its manner or complexity.

The zero coda Go(d) is an outlier in this pattern, unless we conclude that the missing /d/ actually counts as being there. Whether or not the missing /d/ counts as being there has more to do with Kanye's phonology than his rhyming filter, and since I'm assuming Kanye's phonology comes form AAE, this could also be a property of the phonology of many AAE speakers. So, I want to conclude that it is probable that for many AAE speakers, when they produce zero coda Go(d), the /d/ still counts as being there.

Now how gone is the /d/? That's where the phonetics comes in, and the answer seems to be "very."

This is just a very general view of how to use rhyming verse to figure out something about the phonology of any language or dialect. A speaker of some language has some phonology which generates forms, and then they have a rhyming filter to see what works. By working out what the properties of the filter might be, you can try to reconstruct what the properties of the phonology is.

Disqus for Val Systems