Tag Archives: Text-to-Speech

How Do They Say That? Using Text-to-Speech for Foreign Language DX Identification

By Don Moore

More of Don’s traveling DX stories can be found in his book Tales of a Vagabond DXer [SWLing Post affiliate link]. If you’ve already read his book and enjoyed it, do Don a favor and leave a review on Amazon.

When I first discovered shortwave radio way back in 1971, right from the start I was both a listener (looking for interesting content) and a DXer (trying to hear as many different places as possible). But my DXing had a problem in that I only spoke English. Sure, lots of countries had international broadcasters with programs in English. But how was I going to log the many that didn’t?

In July 1972, I ordered a sample copy of FRENDX, the bulletin of the North American Shortwave Association. It was the first time I had seen a DX bulletin, and I learned a lot. The most eye-opening thing I found out was that other North American DXers were routinely identifying and logging radio stations using languages that they didn’t themselves speak. I was genuinely surprised by this. But if they could do it, then I could, too.

I began by going after Spanish-language broadcasters. I had started studying Spanish in school, so I knew the pronunciation rules even if I didn’t yet understand much. Some station names, like Barquisimeto and Bucaramanga, were even fun to say. I soon branched out to DXing stations in Portuguese, Arabic, and French and eventually to other languages like Russian and Indonesian. One thing that helped was that in those days the World Radio TV Handbook included transcriptions of standard ID announcements from many stations, such as in the image below of Radiodiffusion Television Algerienne in the 1972 WRTH. But whatever the language, the difficulty was always the same. What my English-centered brain thought the words in a book should sound like didn’t always match how the native speakers of the language actually said it.

The anecdote, of course, was experience. Over time I developed an ear for the various languages used by lots of DX stations. Even so, long words could still trip me up. Sure, my knowledge of Spanish made Barquisimeto and Bucaramanga easy. But words like Tanjungkarang and Manokwari in Indonesian and Itatiaia and Goiânia in Portuguese still gave me headaches.

The Reasons Why

In the late 1980s I went back to school to get a master’s degree in Applied Linguistics with the goal of teaching English as a Second Language. One of the first classes I had to take was phonetics – the study of how our mouths make the sounds that turn into words. It’s more complex than you might think. The International Phonetic Alphabet has 107 symbols for distinct vowel and consonant sounds. I had to learn to audibly recognize each of them. (Just don’t ask me to do it now.) Each of those sounds can be further modified in various ways, and it’s those sounds with modifications that produce the unique phonemes that make up each language. Standards of measurement vary, but the best estimates put the total number of phonemes in all the world’s languages at around eight hundred. Most languages only use around forty. English has 44 phonemes.

So that station name spelled out on a printed list may not only sound different than you expect, but it likely also contains sounds that you don’t even know. And when the sounds get chained into words, there’s the issue of which syllable gets the stress (consider photograph versus photography). Many languages help by using accent marks when stress differs from the standard. As words get chained into sentences, the connections also affect pronunciation, such as how “got to” becomes “gotta” in colloquial English. Other languages do the same thing. And I’m not even going to go into the use of tones in languages like Mandarin, Vietnamese, and Thai.

If you think English is an easy language, you are wrong. In terms of how the pronunciation of a word matches the spelling, English is one of the worst. There are historical reasons for the inconsistencies, but about 20% of English words are pronounced significantly differently from how they are spelled. And we don’t use accents or other diacritic marks to indicate when the sound differs from what’s standard. That makes English a difficult language to learn to speak and to understand. My linguistic DX heroes are any non-native English speakers DXing American medium wave stations. Local people do not pronounce places like Louisville (Kentucky) or DuBois (Pennsylvania) the way you would expect.

So there are all types of reasons why that station identification may not sound like we think it should sound.

But There’s a Solution …

I didn’t lead you all the way here just to complain about how complicated it all is. We now have a solution.

Text-to-speech is the process of presenting a computer program with a word or phrase and having it generate the sounds. For most of us, our first experience with this probably happened two or three decades ago when calling an automated phone system that used a mechanical voice to read back unique information that couldn’t be prerecorded by a human. It was pretty bad, and text-to-speech of the era deservedly got a bad reputation.

But the good news is that, like so much technology, text-to-speech has gotten a lot better. Some of it is so good that it’s become difficult to distinguish computer-generated speech from real human speech. Of course, the ways that can be misused also make it scary, but that’s a topic for other forums. Dozens of companies provide text-to-speech software in dozens of languages. And most of them have websites where you can try out their software for free.

I do a lot of DXing of stations in languages that I don’t understand or know well. And sometimes I hear what seems to be an ID but doesn’t match what I would expect from the spelling of any listed station. Then I go to one of those text-to-speech websites and plug in the station name to create a sample audio file. I compare that file with what is being said in my DX recording to see if they match.

I’ve been using this method for several years with Brazilian medium wave stations, Russian airports, and Indonesian marine stations, among others. I’ve confirmed numerous IDs this way that otherwise I wouldn’t have figured out, or at least not have been as comfortable that I had gotten the logging right. A couple of times it’s even saved me from making mis-logs on similarly spelled station names on the same frequency.

Most text-to-speech generators offer multiple voices, including both male and female. I always pick the gender that matches that of the ID I’m comparing to. For the most widely spoken languages, some companies offer multiple dialects, like Nigerian English or Argentine Spanish. So if you’re trying to ID a Colombian station, pick Colombian Spanish, not Argentine.

And, very important, be sure to use the correct spelling when pasting in the station name to check. That means including all the accent marks, tildes, umlauts, cedillas, etc. They are there to indicate how the word is pronounced, and if you omit them you will not get the right result. To make sure I have the correct spelling with all the proper diacritic marks, I usually copy/paste from Google Maps, a Wikipedia article, or some other webpage.

Finally, be aware that some local pronunciations are so unusual that the speech engines still may not get it right. I grew up near DuBois, Pennsylvania, and not one of the text-to-speech sites I tested said “DuBois, Pennsylvania” the way the people who live there do. They call their town “DO-boys.” To any native French speakers reading this, I sincerely apologize for what we Pennsylvanians have done to your language.

Links

A quick way to find a TTS engine is to do a web search for “<language name> text to speech”. Here are a few of the better ones.