Friday, February 25, 2011

Hand Coder Praat Script

I've written a Praat script for general hand coding of segmental variation, relying upon forced alignments produced by P2FA or FAAValign.

The most recent version of the script is available here:
https://github.com/JoFrhwld/FAAV/raw/master/praat/handCoder.praat

Background and documentation below.

Background


Recently, we hosted here at Penn a workshop on New Tools and Methods for Very-Large-Scale Phonetics Research. It was definitely my kind of workshop. The talks and the posters were all very high quality, and very interesting.

One tool that was featured rather prominently was the Penn Phonetics Lab Forced Aligner (P2FA). This tool takes as input a recording of speech, a transcription of the speech, and returns a word and phone level alignment of the transcription to the audio (please see the P2FA page for more details).

Of course, once you have a large corpus of time aligned transcription, the ideal thing to do is an automated analysis of the acoustic data. In fact, this is the goal of the FAAV project, which focuses on analyzing vowel formant data. The most recent version of our code to automatically analyze vowels is hosted here: https://github.com/JoFrhwld/FAAV/tree/master/extractFormants

However, for most purposes, there doesn't already exist an automated method for acoustic analysis. For example, if you wanted to study -ing ~ -in variation, or TD deletion, you would have to first build a classifier, which would require some hand coded data anyway.


Documentation

So, I've written an interactive Praat script that allows you to rather flexibly define segments to search for, narrow down the search context to specific word and segmental contexts, and define segmental contexts to exclude, as well as a list of stop words. Given an audio file, and the output of P2FA or FAAValign, the script will search for the specified contexts, play them, and allow you to enter a code. It will then write your code along with other important information about the token which can be used for analysis in and of itself, or as training data for a classifier.

Setup
Open a Long Sound file and a Text Grid into Praat. These two objects must have the same name. Next open handCoder.praat. To run the script, select Run>Run.



Defining the Search
A dialogue box will open, allowing you to define segments to search for, and refinements of the search context. The default settings are for coding TD deletion.
You can understand these settings this way:
  • Search objects with these names.
  • Send output to this file.
  • Search for T and D.
  • Restrict the search to word final contexts.
  • The segment must be preceded by a consonant.
  • No restriction on following context.
  • Exclude segments preceded by R.
  • Exclude segments followed by T, D, TH, DH, JH, and CH.
  • Exclude AND.
  • Play a window of 3 words preceding and following the word the segment is in.
  • There is no default code
These are what the settings from -ing, or str- coding would look like.


Coding
As the script runs, it will play segments of the audio surrounding segments which meet the search criteria. Then, the coding window will open. It contains two fields: one for codes, and one for comments. After entering codes and comments, hitting enter, or clicking on Continue will move along to the next segment.

Output
The output of this script is a tab delimited file with the following pieces of data for each segment.
  • Object name
  • Segment of focus
  • Word position of the segment
  • Code from the coding field
  • Time of segment start
  • Time of segment end
  • Word of focus
  • Word start
  • Word end
  • Preceding segment
  • Preceding segment start
  • Preceding segment end
  • Window duration
  • Vowels per second in the window
  • Comments

Feedback

Please feel free to contact me with any comments or question. You can find my e-mail on my website: http://www.ling.upenn.edu/~joseff/

Thursday, January 20, 2011

Language Census Data

Hat tip to Mr. Verb for pointing out that the American Community Survey collects data on language spoken at home. I've downloaded their pre-compiled detailed spreadsheet, but you can generate custom tables broken down by various geographic granularities, along with all sorts of other demographic information at the American FactFinder website.

The data I downloaded had two data columns (excepting the margin of error estimates): Number of Speakers and Number who spoke English "Less than very well". Here are the top five languages other than English spoken at home, along with the English only numbers for comparison. The proportion column represents what proportion of all speakers surveyed each language represents.

Language Number Of Speakers Proportion
English only 225,488,799 0.804
Spanish 34,183,622 0.122
Chinese 1,554,505 0.006
Tagalog 1,444,324 0.005
French 1,304,758 0.005
Vietnamese 1,204,454 0.004

"Chinese" represents people who wrote down "Chinese" as another language spoken at home. They also have reported numbers for Mandarin and Cantonese separately, but there's no way to apportion "Chinese" responses to one dialect or the other.

I wondered whether there was a relationship between how many speakers of a language there were, and how many of those speakers spoke English less than very well. You might think that the larger the available speech community for a non-English language, the less need for English there would be.


The answer (at the national level mind you) looks to be "maybe a little," but there's a lot of variation.

Wednesday, January 12, 2011

Grammar-phobia -or- judging a book by its cover.

I was walking through the University bookstore today to pick up beginning of the semester office supplies, when this book caught my eye:

Grammar Sucks: What to Do to Make Your Writing Much More Better.

Right off the bat, the cover art is goofy. I am highly incredulous that there is a native speaker of English that does not know the comparative form of good is better, and if there were some speakers who regularly said or wrote gooder, it must be part of their native dialect, not a habitual speech or writing error. Morphological rules like that are part of the natural language system, and are not learned in school along with arbitrary orthographic conventions, like comma rules (which I actually never mastered).

The back of the book only got worse.


Do you suffer from grammar phobia because...
  • You're so used to IMing, you've forgotten how to write a normal sentence. :-)
  • You've started thinking in rap lyrics.
  • Last time you gave a report, your handout got you laughed out of the room.

I won't start off by pointing out that the elided sentence at the beginning is clearly a question, and the three continuations underneath end in periods, not question marks, because that'd just be catty. ...oops.

What made this book seem blog-worthy to me is the not-so-subtle coded language used to refer to those speakers who the book cover authors (maybe not the book authors) feel are culpable for the degradation of... I don't really know what. Let's take them in turn.

So used to IMing

Some people are really bugged by text message abbreviations, like "c u l8r", probably because they don't understand the difference between the arbitrary orthographic system we use to encode spoken language and native linguistic competence. But my guess is that what really bugs a lot of them, and who this bullet point is really aimed at, is youth. While technology use like text messaging and instant messaging is diffusing up into older age groups, and the earliest adopters are getting older, the use of electronic communication like this is still solidly identified culturally as an activity of youth. Just how many sunday morning news stories have you heard about people who send thousands of text messages a month? Who do they always showcase? Teenaged girls.

Thinking in rap lyrics

This is just blatant. Ok, of course not all rappers are black, but it is an art form that is so solidly identified with the African American community, more so than texting with youth. And, of course, they're not really talking about "rap lyrics," they're talking about AAVE (African American Vernacular English). What an offensive and transparently coded throwback to the linguistic inferiority of African Americans!

But, let's take them at their word. Maybe you have grammar phobia because you're thinking in rap lyrics. Do you mean, like, you're freestyling in your head all the time? Do you mean you're kind of like this guy?

You mean, all your thoughts have flow, and rhyme, are creative, and drop properly formed Spanish imperative verbs? To the book cover authors: you fucking wish. I mean, I wish I could do that.

Laughable Handouts

This isn't coded language for a demographic as far as I can tell, but coupled with the first two lines, it makes a clear point. If you are young, and black (and your hat's real low), you're not worthy of social respect, or economic achievement.

* * *


Needless to say, I went on to go buy my office supplies, and didn't read the body of the book. I can't really tell you if it gave any good advice that made any sense. This book is just another case where supposed discussion of language isn't really about language. It actually ties in nicely with my previous post on how people discuss language in terms of morality. Here, the book cover authors are laying blame on the same groups of people that are accused of leading moral decay: youth, and racial and ethnic minorities.

Disqus for Val Systems