ITM assignment

profilelord mots
speechtechmag.com_managing_text-to-speech_expectations.pdf

Keyword Search...

Speech Technology Magazine Home

Subscribe to Magazine

Current Online Edition

Past Issues

Web Events

STM eWeekly

Speech Awards

SpeechTech Blog

White Papers

Sponsored Videos

RSS Feeds

Conferences

SpeechTEK

CRM Evolution

Customer Service Experience

Online Reference Guides

Buyers Guide

Vertical Markets Guide

Advertising Info

Advertising Opportunities

2017 Media Kit

Editorial

2017 Editorial Calendar

Submission Guidelines

About Us

About Us/Contacts

Other Information Today Sites

EContent

DestinationCRM

Faulkner

Information Today

KMWorld

OnlineVideo.net

Smart Customer Service

Streaming Media

Managing Text-to-Speech Expectations Being clear up front and providing options will make the technology an easier sell.

By Gregory Pulz - Posted Sep 27, 2013 Print Version Page1 of 1

The quality of text-to-speech (TTS) engines has increased dramatically in the past few years. Users of the technology have gone from working with the robotic sound of the "drunken Swede" as the cutting edge of TTS to a voice that could mimic Marilyn Monroe's. Intonation, prosody, and inflection have all become much more natural. We are approaching a norm where a computerized system can carry on a totally realistic conversation with a human.

The work being done by companies involved in research and implementation on both hardware and software for TTS has led to excellent products in the marketplace. TTS is being used in numerous solutions for a vast range of business users, from some of the smallest establishments to Fortune 500 corporations. And these companies continually strive to offer the best solutions available with the technology at hand, but always with an eye towards greater innovation in the near future.

In the voice application industry, we are almost at the stage when a computer can carry on a free-form conversation with a synthesized voice. We have rightly touted the new functionality of our TTS engines. When clients hear of these advances they end up expecting unconstrained, fluent speech from the present TTS engines for their specific application needs.

Our challenge as interactive voice response (IVR) system designers is to manage the expectations of the client when using TTS technology. One of the main things to do is to employ TTS to benefit application performance and not merely to showcase a whiz-bang type of technology. One can manage a client's expectations by the judicious use of the technology. Whereas TTS output is not yet indistinguishable from natural, recorded speech, it is extremely effective for use in IVRs to compliment recorded speech.

As a general rule, IVR designers should use recorded speech whenever possible simply because it sounds natural. However, the decision to use TTS versus recorded speech is usually based on a consideration of costs and benefits. Although recorded speech sounds more natural, TTS can be used effectively and more cheaply in specific situations, such as playing back items from a large list or from a list that changes frequently. In deciding on the use of recorded speech versus to TTS, one needs to always keep in mind the balance between the cost and the quality. Sometimes a solution that provides TTS-quality speech is perfectly acceptable, and sometimes recorded speech is necessary.

If the TTS is intelligible, it might be good enough, depending on the task at hand and the customers who will be hearing it. Certain business spaces might put greater emphasis on the intelligibility of the message than on the naturalness of the speech. For example, for pharmacists, doctors, or other professionals working with a drug- naming application, the main concern for the specialized vocabulary is intelligibility. Intelligibility of the message in this business space is vital, whereas naturalness of the speech is not.

Secondly, when using TTS for playback, it would be best not to switch back and forth between recorded speech and TTS, even if the TTS and the recorded speech are in the same talent's voice. The caller can discern the break points between the switch, and these can become distracting or even irritating. Also, if the TTS engine has the capability, one might consider letting the TTS engine provide playback for some information within a prompt where one would normally use recorded speech. This is most effective with items such as numbers or letters that are adjacent to the TTS playback.

Consider the strategy my colleagues and I employed in an application we worked on recently. We designed a store location playback function wherein the application would play a list of stores with general information. We presented each list item with a leading number, such as "One. The x store. Two. The y store…” and so on. We first recorded the numbers and played them back prior to the TTS readout of the general store information. In practice, this layout was less than optimal: The break between the recorded speech and the TTS playback was noticeable, as was the transition.

The solution was to let the TTS engine provide both the playback of the item number as well as the general store information. The TTS engine had been tuned to accurately play back numbers, so we were able to use that capability to provide a playback of information that was smooth and natural. So don't be hesitant to use TTS where it can perform well within an application.

While these tips are useful, I think the most important thing in managing a client's expectations about TTS is to set up a prototype using the TTS engine early in the development process. You need to meet the client's expectations head on. Clients sometimes come in with incorrect assumptions as to what TTS can and will do. In the prototype, you can have the client listen to some scenarios of what a caller would hear in the finished application. This gives a realistic view of how the TTS playback would be integrated into the complete customer experience for the application in question.

If possible, one could even allow the client to use a demonstration system of the TTS engine. They could enter business-specific information that could be played back within the application; information that would not be familiar to you as a designer. The drawback here is that the client would be hearing the TTS playback separate from the other recordings within the total application. That is why it would be best to create a prototype that will show both playback of recorded speech and TTS in a realistic setting.

It would be best to have the parameters of the demonstration system be the same as, or close to, the parameters in the production system. For example, one needs to take into consideration such items as the audio sampling levels, filtering, system database access and playback through a phone setup. Even if these items can be closely matched to the production system, the client would need to be reminded that the demo might not reflect the true fidelity of the IVR in a production environment. There could be other variables that cannot be controlled in a demo

fidelity of the IVR in a production environment. There could be other variables that cannot be controlled in a demo of the system.

When I've created demos in the past, clients generally would want to hear how the TTS would play back some of their most complex business terms. If the engine did not pronounce these terms well, we would start a dialogue on how the terminology should be pronounced, as well as the pronunciation of other, more common terms within the business space. The end result would be a set of specific recommendations to massage the playback of information within the application. A useful example of this is in the medical profession, as I mentioned previously.

Other business areas with specialized vocabularies would benefit from the up-front identification of the specific pronunciation deficits of the generic TTS engine. The goal would then be to create an exception dictionary of specific verbiage (drug names, city/state, etc.) that would augment the generic engine. The TTS engine could then provide specific updates for the pronunciation of items within that specialized business space.

When possible, it is also a good idea to check the performance of your chosen TTS engine with other engines that have been used in similar situations. Such comparisons can provide insights into how the chosen TTS engine stacks up against others, as well as provide opportunities for enhancing the capabilities of the TTS engine that you are using. If you find any applications, prompts, vocabulary, or pronunciations that are currently difficult for the TTS engine, provide this feedback to the IVR developers. They could use that feedback to address and most likely improve the performance of the TTS engine, both in the short term for the existing application as well as the long term for future scenarios that would need TTS capabilities.

TTS has definitely come a long way. It will continue to evolve to sound more natural and realistic. But we still need to temper clients' expectations on what TTS is good for. Be sure to determine up front the client's level of familiarity with TTS. Doing so will enable you to create an effective sales pitch on when to use TTS. Early, collaborative involvement will then help to ensure that their expectations about the technology remain realistic.

Gregory Pulz has been involved in user experience engineering for more than 25 years. He is a principal member of the technical staff at AT&T and has been involved with designing IVR applications for internal and external organizations of AT&T for more than 17 years. He can be reached at [email protected].

Print Version Page1 of 1

Copyright © 2007 - 2017, Speech Technology Media, a division of Information Today, Inc. PRIVACY/COOKIES POLICY

0 Comments Sort by

Facebook Comments Plugin

Oldest

Add a comment...