Human or bot: how to create someone who talks to you on the phone

What are the ways to voice a bot? 

There are two main ways to voice a bot: use speech synthesis or

In the first case, you will get a pleasant, correct, but most often emotionless voice acting, in the second case, lively, natural speech.For example, if a voice assistant is initially positioned as a robot, then a little mechanical speech will not beIf the goal is to turn a bot into a full-fledged call center employee, then it is more rational for the company to use a pre-recorded speech to process questions. 

One of the most famous methods of speech synthesis evenamong non-programmers — TTS. The computer simply reads and voices the text. A striking example is the synthesis of Speech Services by Google, which can be heard in the Google Translate translator. The service's Russian speech is very easy to distinguish from natural speech, which is the main problem with such synthesis. To place the necessary accents, they use SSML — speech synthesis markup. But even with it, it is difficult to achieve a sound like a professional announcer. Therefore, we began to develop a high-quality alternative.

Why do we need speech synthesis with variables 

Businesses are actively using voice capabilitiesbots and aims to make their speech more pleasant and personalized towards customers. Personalized communication increases engagement and loyalty: 70% of companies using advanced personalization have seen 200% or more ROI on their efforts.

To make the bot's voice realistic, the companyrecord the voice of a speaker, actor or contact center operator. The downside is that all lines are static. This is suitable, for example, for instructions on how to cancel an order or inform about current promotions. But business messages contain a lot of changing data - variables.

To make the bot's voice realistic, companies record the voice of a speaker, actor, or contact center operator. The downside is that all the lines are static.

For example, to confirm an appointment at a clinic:Patient's full name, doctor's full name, date and time of appointment. In the delivery case: client's full name, order number and its composition, amount and delivery time. Each entry must include individual parameters of the client or his request. And then file gluing, audio mixes or simply robotic voiceover come into play.

The essence of gluing is that it creates"applications" from numerous recordings of the announcer. In case of delivery, it is necessary to record about 1.5 thousand frequently encountered streets, city names, house and apartment numbers. This is a very labor-intensive task, which requires considerable effort and time from the announcer. After recording, the audio files must be glued together in the correct sequence. A professional sound engineer can make smooth transitions between the cuts and place pauses, but the joints will still be heard, and large volumes of source data make such work unmanageable. We tried this option, but realized that it is not suitable for solving complex problems.

We know that some companies try to find a speaker with a voice similar to the standard SpeechKit voiceover from Yandex or Google Speech. But in this case, the inserts are also audible.

We tried to integrate it into the voiceover recordingsTTS. That is, we did not record all the names of streets and cities in advance, but generated the necessary phrase using TTS based on the announcer's voice and inserted it into the pre-recorded announcer's line. The result was a mixture of emotional, lively speech of a person and a robot.

Whatever we try from different instruments,we did not get a good result. As a result, we came to develop our own algorithm, which we called hybrid synthesis. This technology allows you to quickly change phrases in the voice recordings for the voice bot, it is enough to pass the replacement to the algorithm. At the same time, the synthesized speech copies the intonation and emotions of the speaker, sounds natural and does not stand out from the context. Thus, you can voice any variables that were not in the original speaker recording, as well as test new scenarios.

TTS

How Hybrid Fusion Works 

The first stage is working with the announcers.To cover the diversity of the Russian language, it was necessary to record 10 hours of speech in a professional studio. These are phrases from books, news, and frequently encountered data: numerals, names, addresses, and cities. This way, there is enough data for training, after which it is possible to start creating hybrid synthesis models for each speaker.

For each robotic calling project, the announcer records templates of phrases, for example: “Hello,Sergei Petrovich! We will deliver your order to your address tomorrowTverskaya street, 3With14 to 18hours.Would it be convenient for you to receive it?" The hybrid synthesis model can handle tasks such as replacing Sergei Petrovich with Anna Vasilievna, and the address and time with 24 Bolshaya Zelenina Street from 10 a.m. to 2 p.m. The result is a new replica with synthesized variables: "Hello,Anna Vasilievna! We will deliver your order to your address tomorrowBolshaya Zelenina street, 24With10 to 14hours. Will it be convenient for you to receive it? " For each call, a separate task is created and executed in the customer base. 

Voiceover

High sound quality is achieved due to the fact that the neural network uses the intonations and emotions of the speaker from the example.

How to set up hybrid synthesis 

We offer several ready-made modelshybrid synthesis with female and male voices. Template calls work on the JAICP and JAICF platforms, calls from bots created in other services are possible via API. Setting up a script from scratch takes several hours. One synthesized replica costs 12 kopecks, voicing a template for a project is 3,000 rubles per hour of the announcer's work.

The set of ready-made voices will expand.There is also an option using Just AI hybrid synthesis with the creation of a model for an individually selected announcer or who is the official voice of the company.

hybrid synthesis

Hybrid synthesis technology allows you to personalize IVR and robotic calls for the purposes of NPS surveys, questionnaires, reminders, upsale and support for loyalty programs.

Read more:

A black hole in the galaxy proved Einstein right. The main thing

Space destroys bones and changes their structure: scientists do not know how people will fly to Mars

Astronomers have found planets that are different from Earth, but suitable for life