Friday, July 24, 2015

The healthy and unhealthy vocal fries

There has been much discussion in the news media lately about the phenomenon known as "vocal fry" and its use among English-speaking women in the United States. Vocal fry refers to the irregular vibration of one's vocal folds and it is normally produced with low pitch. In an interview with Terry Gross, Susan Sankin, a speech-language pathologist stated that vocal fry is harmful to one's vocal folds. In a follow-up piece on 7/23/15 on NPR, she maintains this view, stating
...I have heard ENTs say that it can cause damage. And for a lot of the languages where it's a habitual pattern - as you develop from a young age, that's how you're training and using your vocal cords. And I think when you start to fall into that pattern later on, I think that it can cause some damage. Again, I'm not a doctor, so I can't say that I've looked at people's vocal cords and I've seen it, but I have heard ENTs say that they do notice that it can cause damage. And sometimes the jury is out on that as well.
Just what is behind this notion that vocal fry may be damaging for one's vocal folds? After all, what we're calling "vocal fry" is used in many languages to contrast meaning among words, just like one might contrast the words 'heed' and 'hid' by their vowel sounds. It is also ubiquitous throughout the languages of the world to mark boundaries between phrases. How can something that is so common be considered a vocal pathology?
To answer this question, it's necessary to first make a distinction between speech articulation and speech acoustics. Speech articulation involves what you do in your oral cavity to produce speech sounds. Speech acoustics involves what sounds you hear that convey a linguistic message. Phonetics involves the study of both these things and phoneticians are interested in understanding how certain articulations produce certain acoustic characteristics. One can more easily investigate this relationship for sounds with un-hidden articulations. For instance, the 'p', as in 'pan', is made with the lips. One can see them close when this sound is produced and observe silence in the acoustic signal while one's lips remain closed. 
The same thing is not true for the vocal folds though. When it comes to the vocal folds, it's often a rather messy business to investigate what they are actually doing. They're quite small (just about 1 - 2.5 cm in length, depending on one's sex) and taking a video recording of them moving during speech involves inserting a small camera attached to a wire through one's nostrils to hang near the upper portion of one's pharynx (throat) and peer downward. As you might imagine, many people object to having foreign objects inserted into their noses.
One way around this is to just look at the acoustic signal and interpret what the configuration of the vocal folds must be. People don't object nearly as much to being recorded as to having wires inserted into their noses. Moreover, plenty of other articulations have consistent acoustic consequences. For instance, lowering one's tongue and jaw during speech changes the acoustic resonances of the oral cavity in a rather consistent manner. So, the theory goes, one can rely on the acoustics of the speech signal to tell us what the speech articulators are doing. So far, so good.
While this method is fairly robust, there's something problematic about it with the vocal folds. What is called "vocal fry" involves irregular vibration of the vocal folds (see below, taken from a previous post). In the figure here, one notices the irregular vocal fold vibrations on the right. Each glottal pulse is individually stronger (has higher amplitude) but the timing between each is erratic. To quote a well-known linguist, this voice quality sounds like "a stick being dragged along a fence."

But, to return to our main interest, what is the articulation that gives rise to this acoustic pattern. The term "vocal fry" refers not to the articulatory configuration, but to one's perception of the acoustics. As it turns out, there are many things that can produce the type of vocal fold vibration that we observe above. Much like a wheel that is fastened too tightly, if one constricts the larynx (where the vocal folds sit), it is harder for the vocal folds to vibrate regularly. Since the vibration of the vocal folds requires consistent airflow from the lungs, if one runs out of breath at the end of a sentence, the vocal folds also do not vibrate so regularly.
For people who have developed vocal fold nodules, brought on by laryngeal cancer or other pathologies, the vocal folds also do not vibrate so regularlyClearly, the same acoustic pattern matches a number of different articulatory configurations. Yet, all of this irregular vibration is described with a cover term, "vocal fry." 
So, if one were to observe vocal fry in different speakers, what could one conclude? While there is independent evidence for the health of speakers in a clinical setting, the notion that vocal fry is pathological is a case of the symptom getting confused with the cause. Since we rely on the acoustic signal to tell us about articulation, we associate the presence of a certain characteristic of the acoustic signal with an articulatory pathology. In other words, vocal fry must be pathological, right? No, in fact this is a classical logical error (affirming the consequent).
Research on the production of voice quality across languages has shown that speakers use a number of different configurations to constrict the larynx and produce what is known as "vocal fry." Acoustically, and only acoustically, these might appear similar to pathologies that produce irregular vibration of the vocal folds. Yet, the cause of the irregular vibration is different. The articulation of the vocal folds is difficult to examine. So, researchers have assumed aspects of their configuration on the basis of what the acoustic signal says. Yet, this only works insofar as there is not a one-to-many association between the acoustic signal and the articulatory mechanism involved. 
The problem is, we do have a many-to-one relationship when it comes to voice quality. Thus, one can not just infer on the basis of one part of the acoustic signal what articulation is involved. Speech-language pathologists, like Susan Sankin, might heed this before they label "vocal fry" as damaging to one's vocal folds. It's not the voice quality that is damaging, but this misunderstanding of cause and effect.
What does this mean for the young women whose vocal fry is singled out as being unhealthy and damaging for their careers? It's the attitudes and knowledge about women's voices that needs to change, not the voices themselves.

Monday, July 6, 2015

Being cooperative is not evidence of confirmation bias

A few days ago, the New York Times posted a piece which argued that confirmation bias is a common failure of human thinking. Confirmation bias is the idea that one tends to interpret new facts in terms of one's existing preconceptions.

The author of the study, David Leonhardt, discusses confirmation bias by way of a mathematical example where the reader is asked to guess the rule determining the sequence "2, 4, 8" by testing additional examples. Thus, one can type in sequences like "4, 8, 16" or "10, 95, 387" and see if they follow the same rule as the sequence "2, 4, 8." If one enters a sequence like "4, 8, 16" into the boxes in Leonhardt's article, one receives a confirmation that it also follows the same rule as that which produced "2, 4, 8."

So, just what is this rule? Leonhardt states:

"...most people start off with the incorrect assumption that if we’re asking them to solve a problem, it must be a somewhat tricky problem. They come up with a theory for what the answer is, like: Each number is double the previous number."

The true rule, Leonhardt explains, is not that each number is double the previous number, but rather that each subsequent value is greater than the preceding value. That people assume the former rule is taken as evidence for confirmation bias. As stated, "Not only are people more likely to believe information that fits their pre-existing beliefs, but they’re also more likely to go looking for such information."

However, it strikes me that there are other, rather sensible reasons that people will assume the former rule that Leonhardt does not consider. One is found among the the well-known maxims of conversation, created by the famous philosopher of language, Paul Grice. These maxims, well-known to any introductory linguistics student, state that conversation is guided by constraints of quantity, quality, relation, and manner. As a default, we assume that speakers will give only enough information, be truthful with it, be relevant to the topic, and be clear, respectively. When speakers deviate from these expectations, we are annoyed with the conversation. In such cases, we might state "He was long-winded." or "He kept going off on tangents." Our ability to follow these maxims demonstrates our cooperation within a conversation. Hence, they fall under what Grice terms the cooperative principle.

Grice's maxim of quantity states that one should not make his/her contribution more informative than required. Thus, when someone asks for directions to a particular room in a building, one does not expect the speaker to provide instructions on how to open a door nor the history of certain rooms that the listener will likely pass. If additional details are provided, our interpretation is that they must somehow be relevant (incidentally, another maxim). So, just what might Grice have to do with Leonhardt's example here?

Consider the initial example that he provides: 2, 4, 8. The reader's expectation from this example is that it is as informative as necessary. If the author chose a sequence where each subsequent value is double the previous value, then this must be relevant to the question. After all, the expectation is that the author has provided this information and it must be important. When we hear that the rule is, "Haha!", not what we assumed, our reaction is one of surprise. Why provide this particular example if any random sequence, like 1, 5/3, 9, would have sufficed?

Providing too much information in this way would seem to be a case of conversational deception. The listeners/readers are led astray believing that the author was following the maxim of quantity and relevance when, in fact, the example was intended to be overly informative. So, are 78% of those who participate in this particular task guilty of confirmation bias? Perhaps some are, but Leonhardt would be wise to consider that most people are guided not by the expectation that the problem is tricky, but rather by the expectation that the author's example does not provide too much information. An entirely different outcome would be produced if the example were as simple as "1, 2, 3."

*Incidentally, the rule "Each number is double the previous number." is just a more specific case of the rule "Each number is greater than the previous number." The first entails the second.


Wednesday, March 4, 2015

More on Triqui reduplication

Back in December, I posted a quick summary of this interesting pattern on Triqui verbs to Facebook. The post read as follows:

===
So, I discovered a new pattern today. There is a process of partial reduplication on Triqui verbs that indicates something like an emphatic, but non-specific third person. The reduplicant is partially fixed, always containing tone /4/ and a coda /ʔ/, but the vowel identity (including nasality) is taken from the stem. The stem has to be marked with the non-specific 3rd person (which is marked by tone/glottal deletion).

/βĩ³/ 'to be/exist' > bĩ³-ĩʔ⁴
/a³βi³²/ 'to leave' > a³βi³-iʔ⁴
/tʃeh⁴/ 'to walk' > tʃe³-eʔ⁴
/a³tah³/ 'to say' > a³ta³-aʔ⁴   
===

Well, as it turns out, the pattern doesn't quite just signify an emphatic. It appears to be a way of marking a wish/blessing on someone, e.g. "May they do X." Thus, to state "May they eat" or "Que comerían", the form is /tʃa²=aʔ⁴/ and "May they run" would be /ku²nã²=ãʔ⁴/. The interesting thing about these new forms is that both of the verbs occur in the potential aspect. The unmarked, progressive/habitual form for 'run' is /u⁴nãh³/ and the perfective form for 'eat' is /tʃa⁴³/. So, it is not clear if this reading of these forms just derives from the fact that they are in the potential aspect. This reading also appears to differ from the hortative, e.g. /tʃa²=yũʔ¹/ 'Let's eat!' (Incidentally, this is different from the regular /tʃa²=ũh⁴/ 'We will eat ~ We are eating.')

So, the mysteries now are: (a) Is there some semantic consistency when reduplication is used on verbs with varying aspects? (b) Might we be able to tease apart just what this means with unmarked aspect?

On a related note, one of the things I've come to realize about Triqui clitic morphology in doing documentation is that there is a very productive system for marking vocatives (or perhaps just 'terms of address') in the language. I used to think that only certain words had distinct, lexicalized vocative forms, e.g. /nni³/ 'mother' vs. /nnãh⁴³/ 'mother!', but it turns out that most names, and even clitics can undergo this process. This is done by a process that looks eerily like the reduplicative system above. You take the word, e.g. /tʃu³be³/ 'dog' and just change it's tone to /4/ and add a coda /ʔ/, e.g. /tʃu⁴beʔ⁴/ as in something like 'Hey, dog, how are you?'. One can do this with proper names too, e.g. 'Enrique' is /li⁴ki⁴³/ normally, but would be /li⁴kiʔ⁴/ when referred to directly. I have a feeling there is some relation between emphatic forms (maybe this is the good term for it after all?) and direct terms of address, but just how the expression of a wish/blessing fits in here is a mystery.

Incidentally, I learned the above form for 'dog' through a story about a magical dog who learned how to grind corn and make tortillas for a man who went to farm. When the man came into the house, he had some conversations with this dog and had to address it directly.

Saturday, January 3, 2015

Ode to 2014

Here's to that email you never answered,
the message you never got,
still flagged and beeping,
haunting in being unknown
to all but the sender.
Nagging and distracting
like an unscratched lotto ticket
or a joke, unheard because you arrived too late.
"You had to be there"
but you were unavailable
as the message sits,
bored.

Whatever emergency we are waiting for
we wait too long,
with false flashing yanking us
to action-reaction,
seeing and skimming and sitting
and revealing to us no less than we expect;
minutiae.

The attack move to click
to see or hear what might be there
gives only a hasty rush
as we unfurl objects of glitter
to finally know no more than we did before.
Must a hamster run in its wheel
just because it's there?

Monday, August 4, 2014

If airlines were restaurants...

If airlines were restaurants...

We'd stand in line to get in,
but we could pay more to be put at the front of the line.
We'd be seated at a table of screaming children,
but we could pay more to get our own table.
We could also pay more to sit near a window
or in a "premium" seat near the fire escape.
But, just think, if we called in ahead of time to reserve a seat,
we wouldn't deal with such things.
...but we'd pay more to even do that.

We might wait for hours for our food
or suffer the injustice of being told
"no food today"
And we might have to go wait in another line.
Because it's not the cook's fault
that it's raining.

We could get a frequent diner card
and come in with a coupon
to be told that the terms had changed
because the money we once spent
at this fine dining establishment
was in the past.

If airlines were restaurants,
we might not visit them
or just decide to occasionally
visit the only other restaurant
in our town.

We might just make our own lunches
or forage for food in the streets
where the restaurants have ensured
a slow food movement.

Just don't complain to the chef or the wait staff.
You might not be allowed to come back.

Friday, June 6, 2014

Vocal fry doesn't harm your career prospects, but not being yourself just might

In a recent article in PLOS One, authors Anderson et al. find that, vocal fry is harmful to a woman's career prospects ("Vocal Fry May Undermine the Success of Young Women in the Labor Market"). At face value, the findings are surprising. When we perceive speech, we pay attention to very subtle cues and it's always a surprise when certain things that we hear are shown to have some hidden influence on our attitudes towards people. Moreover, anything to do with the influence of one's voice on employment prospects is similarly notable in a market where competition for jobs remains fierce.

Resultingly, this article has garnered some attention in the media recently (as in this Atlantic piece and this Marketplace piece) with the conclusion that women now need to police how they speak for fear of being perceived as untrustworthy by an employer. Yet, on closer inspection, it turns out that the police might not be needed at all. The original study contained quite serious flaws in its design which, when considered carefully, prevent us from making any conclusions about which specific acoustic characteristics sounded "untrustworthy" to the listeners who participated.

The design of the study was relatively straightforward. A group of 800 people, via an online system (Qualtrics), listened to speakers produce the sentence "Thank you for considering me for this opportunity." Some of these sentences were produced with vocal fry, which, in contrast to normal voice, involves temporal irregularity in the vibration of the vocal cords (folds) and lower overall pitch (see Figure 1 below). To a listener, the vocal folds sound like a stick being dragged along a fence, where one can hear individual vibrations or pulses of the vocal folds. The listeners were asked to evaluate speakers based on whether they were trustworthy, competent, educated, hireable, and attractive. The expectation in a study like this is that listeners might have different attitudes towards those sentences with vocal fry than they would towards sentences without vocal fry. The big issue here is just where the authors got the voices with vocal fry.

Figure 1: Example of regular (modal) vocal fold vibration and irregular vocal fold vibration (vocal fry) within the latter half of the word "opportunity".

When linguists, phoneticians, or speech scientists want to study whether an acoustic characteristic in someone's voice influences how listeners perceive them, they often will record a person and then modify those aspects of the person's voice which they wish to test. This process, called resynthesis, allows one to carefully control the acoustic dimensions in the signal and requires some knowledge of speech acoustics and digital signal processing. Certain aspects of one's voice are harder to modify than others. As it happens, vocal fry is one of these hard-to-modify characteristics. (I'll leave the more detailed question of why it is hard to resynthesize vocal fry and voice quality, more generally, out of the discussion for now.)

Fortunately, there is a solution. Just as one might buy two types of apples to compare their flavors, we can look for speakers who just happen to produce more vocal fry in their speech and compare them to those who do not produce it. If one were to play the speech of these two groups to listeners (and potential employers), listeners might have different attitudes about one of the groups. This is, in fact, what Yuasa (2010) did in her study of creaky phonation. Yet, importantly, the authors of the study here did no such thing. Rather, they recorded speakers producing normal utterances and then trained them to produce an utterance with greater vocal fry. As a consequence, the speech contained in all of the vocal fry stimuli is actually speech where speakers are attempting to imitate a voice with vocal fry. There are several reasons why this is problematic, but the first is perhaps the most obvious: most people are not particularly accurate at imitating someone else's speech. If you ask the average person to "talk like a Texan", they might (or might could) try to imitate something that they believe to be an important characteristic of Texas speech. Yet, to most listeners, especially those from Texas, they would sound like a caricature of an actual Texan.

As it turns out, this is the rub. While the speakers in the study here insert creak at various places in their speech, its real use in natural speech is more carefully controlled. Previous studies which look at vocal fry, particularly Redi and Shattuck-Hufnagel (2001), find that it is rather restricted. It tends to occur in locations in phrases and utterances where we might expect low pitch. Vocal fry is disconnected from these locations of low pitch in the imitated speech here. Rather, the speakers seem to produce a very flat, robotic voice when imitating vocal fry. The typical intonation for the stimulus sentence is something like "THANK you for conSIdering me FOR this OPorTUnity", where the syllables in caps reflect higher pitch levels than the surrounding ones.

This is not the only way in which the imitated speech sounds unnatural, however. With one exception (speaker 5), each of the imitated sentences produced by female speakers is also longer than the corresponding non-imitated sentence for that speaker, as shown in the table below:

Sentence duration (vocal fry) Sentence duration ("normal")
Speaker 1 2.91 2.25
Speaker 2 2.90 2.84
Speaker 3 2.69 2.19
Speaker 4 2.33 2.07
Speaker 5 2.15 2.37
Speaker 6 2.57 2.43
Speaker 7 3.24 2.57

These differences do not appear to be restricted to particular words either. As seen in Figure 2 (below), almost all words were longer in the imitated speech than in the natural speech. The longer duration here, in comparison with the shorter natural sentences, may have the quality of sounding stilted to the listener.

Figure 2: Duration of words in the sentence "Thank you for considering me for this opportunity" spoken by seven female speakers across two conditions: vocal fry (left) and normal voice (right). Note that the durations for all words are longer in the vocal fry condition.

A related problem in the study is the authors' acoustic analysis of the speech signal. The calculation of pitch in the speech signal requires determining how well successive vocal fold vibrations correlate with one another. When the vocal folds are vibrating normally, such a correlation is possible, but when vocal fold vibration is too irregular, as in vocal fry, it is impossible to calculate pitch accurately. However, an acoustic analysis program may still try to calculate possible (erroneous) values. Anderson et al. argue that the pitch in the vocal fry sentences is universally lower than that in the natural sentences, but they neither controlled nor mentioned how pitch was calculated during durations of vocal fry. In fact, the pitch on the expression "Thank you", which contained no vocal fry in any of the utterances, had universally lower pitch in the vocal fry sentences than in the normal sentences. This suggests that the speakers may simply be lowering pitch across the entire imitated sentence, rather than simply adding vocal fry. Finally, no quantitative acoustic estimation of actual vocal fry (such as jitter, shimmer, cepstral peak prominence, etc.) was ever included in the authors' study. Yes, you heard that right - in a study relating vocal fry to listener attitudes and hireability there was no actual estimation of whether the stimuli differed with respect to the test variable.

Taken together, these observations suggest that the speakers in the study simply attempted to lower their overall pitch level while imitating vocal fry rather than simply including more vocal fry. The increased effort involved in the imitation also made their utterances longer. These two acoustic differences, among others, would seem to contribute to the speakers sounding unnatural when imitating vocal fry. So, when listeners judge the female speakers with vocal fry as sounding "untrustworthy", there is a good possibility that they are simply making such a judgment based on the speaker not sounding like herself. The better lesson that one might take home instead here is that one's job prospects are harmed if you try to talk (or act) like someone who you are not.

References:
Anderson, R. C., Klofstad, C. A., Mayew, W. J., and Venkatachalam, M. (2014) Vocal fry may undermine the success of young women in the labor market. PLOS ONE 9(5): 1-8.

Redi, L. and Shattuck-Hufnagel, S. (2001) Variation in the realization of glottalization in normal speakers. Journal of Phonetics 29:407-429.

Yuasa, I. P. (2010) Creaky voice: a new feminine voice quality for young urban-oriented upwardly mobile American women? American Speech 85(3):315--337.

Tuesday, October 22, 2013

Visualizing vowel spaces in R: from points to contour maps

Typically when linguists wish to examine the vowels of a language, they plot the vowels in an F1xF2 space, which approximates a relative articulatory position of the vowels. Now, there are certainly problems with this approach (lack of F3, possible nasal formants, dynamic movement). Yet, despite these drawbacks, visualizing vowels this way is relatively standard and has the advantage of being understood by a wide audience. In R, there are several methods one might use to plot vowels in a space like this. I will discuss three here, two of which are clearly less than ideal and another which I am in the process of learning still. I will be relying on a data set of Arapaho vowels from elicitation sessions from three speakers. Given the nature of the data, I had to analyze a different number of vowels per speaker, so that one speaker is over-represented (1290 vowels) and two others are underrepresented (600-700 vowels each).

I am interested in visualizing the quality differences between long and short vowels in the language. Arapaho has short, long, and extra long vowels, though I only really have enough data to analyze long and short vowels, so I am sticking to that. I am looking at the monophthongs /i, ɛ, ɔ, u/, which are realized as more centralized variants when they are short. Here is a sample of my data:

Vquality Label Length Vaccent seg_Start  seg_End Duration Time       F1       F2       F3
8         o    v2  short       H  19985.83 20088.10  102.277    2 673.0124 1116.783 2887.543
9         o    v2  short       H  18887.63 18990.38  102.753    2 682.5495 1204.757 2636.614
10        o    v2  short       H  10048.30 10152.09  103.789    2 679.4601 1077.850 2910.295

I am leaving out several columns here (including speaker, word, etc.), but all data is coded for vowel, written i, e, o, u, and for length (short, long).

1. The first way to visualize a vowel space given data like this is to use R's plot function. The default here is to built up a plot by adding individual elements. With the data formatted the way it is, it would be necessary to create several subsets and plot each separately. For instance, if we restrict ourselves to just the long vowels, one could do the following:

> long <- subset(formant_data, Length=="long")
> u.long <- subset(long, Label=="u")
> i.long <- subset(long, Label=="i")
...

The result of this will be 4 different data frames, each of which could be plotted separately as points, as follows:

> par(mar=c(5, 4, 2, 2))
> plot(F1~F2, data=i.long, ylim=c(1000, 200), xlim=c(3000, 600), pch=1, col="red")
> points(F1~F2, data=e.long, pch=2, col="blue", add=T)
> points(F1~F2, data=o.long, pch=3, col="green", add=T)
> points(F1~F2, data=u.long, pch=4, col="black", add=T)
> legend("top", horiz=TRUE, c("i", "\u025B", "\u0254", "u"), pch=c(1, 2, 3, 4), col=c("red", "blue", "green", "black"), x.intersp=0.8)

This produces the following:

Fig 1

This looks good enough to plot observations, but what if one wants to get an idea about where the averages lie? It might be easy to imagine an average "i" here, but the other vowels seem somewhat dispersed throughout the vowel space (the short vowels are even more so). So, this is harder. 

2. One solution is to plot the vowel data using the vowelplot() function in the vowels package. This package can both draw circles around the vowel area and compute mean values for each of the vowels. However, it requires the user to reformat his/her data to fit a template used by the package. Depending on one's data organization, this can be cumbersome. The format of their template is a data frame of 9 columns, which includes: speaker_id, vowel_id, context, F1, F2, F3, F1_glide, F2_glide, and F3_glide. Plotting requires fewer commands, but the options within each command are limited. If we try the following:


> vowelplot(long, color="vowels", ylim=c(1000, 200), xlim=c(2600, 600)); it produces:

Fig 2


This figure is interesting insofar as it color codes the vowels and divides up the space by speaker. It does all this with a single command too. However, there is no way within the package to avoid dividing the data into speakers (as you might notice, I did not specify this in the plot command) when showing individual data points. 

Another advantage of this package is the ability to display vowel ellipses around the plotted data. We can do this with the following command added on:

> vowelplot(long, color="vowels", ylim=c(1000, 200), xlim=c(3000, 500))
> add.spread.vowelplot(long, ellipsis=TRUE, labels="vowels")


Fig 3

Ack! Yes, this is indeed very ugly. Unfortunately, the vowelplot() package always assumes that you want individual speakers. One could simply eliminate speaker differences in the data frame and replot the data with single ellipses. Alternately, one could plot the ellipses without the observations. I won't go into how to do this here. Instead, I will show one of the interesting advantages of using the vowelplot package: the ability to extract and plot mean formant values. So far, I have not plotted the short vowels because the degree of overlap would have been particularly large. I will do so here though:

These commands calculate the mean values for the first two formants among the short and long vowels.
> vlong <- compute.means(mono.long)
> vshort <- compute.means(mono.short)

These commands plot the data:
> vowelplot(vlong, label="vowels", ylim=c(800, 300), xlim=c(2400, 800), title="Vowel space in Arapaho")
> add.spread.vowelplot(vshort, labels="vowels")

This produces the following:

Fig 4


This looks much cleaner, but averages always look cleaner. The vowels with the dots represent the long vowels, while the vowels without the dots represent the short vowels. You'll notice that the short vowels are more centralized than the long vowels, though the back low vowel doesn't really change in quality. So far, so good.

Yet, as phoneticians (or as linguists/speech scientists), we are often more interested in the distribution of the data than the average values. Yet, if we plot ellipses here, it looks just as chaotic as Figure 3. This is because an ellipse contains two mutually perpendicular axes about which the ellipse is symmetric. These axes are the two dimensions (F1 and F2) which position the vowel in the vowel space. However, actual observations are not elliptically symmetrical around the center. Thus, ellipses might tend to overestimate the actual degree of overlap in a vowel space.

3. One solution to using the vowelplot package is to use ggplot2(). This very modern plotting software allows us a larger set of data visualization techniques. One way that I might think about plotting the distribution of my vowel data is with a two-dimensional contour map. Contour maps include three dimensions, with density as a "higher" point. They rely on kernel density estimation (KDE), which is a non-parametric way to estimate the probability density function of a random variable. Given that a prior distribution is not assumed, they also have the advantage of non-symmetry. As far as I know, I have not seen these applied to vowel spaces before. We can plot our data as follows:

> f.plot <- ggplot(formant_data, aes(x = F2, y = F1, color=factor(Vquality))) + geom_density2d(aes(label= factor(Vquality))) + scale_y_reverse() +  scale_x_reverse() + ylim(900, 200) + xlim(2800, 500)+ theme_bw() + scale_color_hue(name="Vowel quality", breaks=c("i", "e", "o", "u"), labels=c("i", "\u025B", "\u0254", "u"))
> f.plot 

This produces the following figure:



What this figure reveals that is so often left out of vowel plots is a clearer sense of the concentration of observations. One can observe a somewhat bimodal distribution for /u/, one concentrated with an F2 around 1000 Hz and another with an F2 around 1500 Hz. These probably reflect differences among speakers, but they may also reflect a difference of context (there is substantial alveolar fronting). If we wish to plot both short and long vowels, we can do so by using a facet_wrap() function. 

f.plot <- ggplot(form.mono2, aes(x = F2, y = F1, color=factor(Vquality))) + geom_density2d(aes(label= factor(Vquality))) + scale_y_reverse() +  scale_x_reverse() + ylim(900, 200) + xlim(2800, 500)+ theme_bw() + scale_color_hue(name="Vowel quality", breaks=c("i", "e", "o", "u"), labels=c("i", "\u025B", "\u0254", "u")) + facet_wrap(~Length)
> f.plot 

This produces the following:


Now, we observe not only the tightness of observations around the median for the long vowels, but the asymmetrical ways in which the vowel space changes as a function of length. There are clear realizations of short /i/ which match those of long /i/ in quality. However, there are also a larger number which encroach into the center of the vowel space (though significantly more along the F1 dimension).

The advantage of using ggplot2() to show this data is that one can represent most of the observations and simultaneously observe non-linearities in the shape of the distribution. Outliers are more naturally excluded since they do not contribute to the estimated density function. By contrast, ellipses simply expand to symmetrically cover the entire space of the observations (or at least a space determined by symmetries inherent to normal distributions).

I think I am a fan of this method for vowel visualization, but I am unsure if this is the right way to go about things. Thus, any commentary is welcome.