Genesys Cloud - Main

 View Only

Sign Up

  • 1.  How are you testing and comparing TTS voices in Genesys Cloud?

    Posted 9 days ago

    Hi everyone,

    We're currently building a Voice Bot POC in Genesys Cloud to understand the caller's reason for contacting us and route them to the appropriate queue based on the detected intent.

    As part of the POC, we'd also like to compare the different TTS voices available and understand what others are using in production.

    How are you testing and comparing different TTS voices? Have you built a separate test flow to listen to the same prompts across different voices, or found an easier way to compare them without continually changing the voice and republishing the flow?

    I'm thinking of creating a simple test flow using a Data Table with a set of sample phrases, so we can work through them and listen to the different voices one by one. Has anyone done something similar, or is there an easier way to compare voices in Genesys Cloud?

    We're currently using Genesys Enhanced TTS with Amazon Polly Olivia NTTS, but we're also interested in hearing whether others are using Enhanced TTS, Generative TTS, or third-party TTS providers and what influenced your choice.

    Also, for any Australian customers here, which voice have you found sounds the most natural for an Australian audience?

    Would love to hear what others are doing and any lessons learned.

    Thanks!


    #ConversationalAI(Bots,VirtualAgent,etc.)

    ------------------------------
    Phaneendra
    Technical Solutions Consultant
    ------------------------------


  • 2.  RE: How are you testing and comparing TTS voices in Genesys Cloud?
    Best Answer

    Posted 9 days ago

    What I normally do during the implementation and ramp-up phase is something very similar.

    When I need to present and compare TTS options with a customer, I usually create one separate test flow for each of the main voices available within what the customer has contracted/licensed.

    I replicate the same basic flow and configure a different TTS voice in each copy. All of them use the same Data Table with the same sample phrases, so the comparison is consistent.

    For the phrases, I normally include scenarios where TTS differences or limitations are easier to notice, such as:

    • Dates and times

    • Currency values

    • Long numbers

    • Phone numbers

    • Acronyms

    • Proper names

    • Decimal values

    • Longer sentences or prompts

    If there are spare DIDs available in the organization, I assign one DID to each test flow. This makes the customer session very easy because they can simply call each number and compare the voices directly.

    If there are no numbers available for testing, I use the option to call the flow directly from within Genesys Cloud. In that case, I provide the customer with the exact flow name to test. Since I normally have one flow per voice, I usually include the TTS/voice name in the flow name as well, which makes it easier to identify which one they are listening to.

    The main reason I started doing it this way was to prepare everything before the customer testing session.

    Instead of changing the TTS voice during the meeting, republishing the flow, testing it, changing it again, and repeating the process, all the flows are already published with their respective voices.

    It was the easiest way I found to anticipate that work and make the actual testing session with the customer much smoother.

    So I think your Data Table approach makes a lot of sense. The only thing I would add, especially if you are planning to compare several voices live with the customer, is having one pre-published test flow per voice, all using exactly the same set of phrases.

    That has worked well for me during ramp-up and customer validation sessions.



    ------------------------------
    Raphael Poliesi
    ------------------------------



  • 3.  RE: How are you testing and comparing TTS voices in Genesys Cloud?

    Posted 9 days ago

    Hi Raphael,

    Thanks! This is really helpful. I like the idea of having one pre-published flow per voice and using the same Data Table and sample phrases across all of them. That should make the comparison much easier.

    One other question if you happen to know: when using TTS and speech recognition/STT within a Dialog Engine Bot Flow, is there an additional usage cost associated with either of these?

    I've been looking through the Genesys documentation, but some of the pricing/included usage information seems to refer specifically to Virtual Agent/Agentic Virtual Agent or AI Guides, so I'm not completely clear on how TTS/STT usage is charged for a standard Bot Flow.

    Would be interested to know how this works in practice.

    Thanks again!



    ------------------------------
    Phaneendra
    Technical Solutions Consultant
    ------------------------------



  • 4.  RE: How are you testing and comparing TTS voices in Genesys Cloud?

    Posted 9 days ago

    In a standard Genesys Dialog Engine Bot Flow, there is already the normal Bot Flow voice usage, but the native speech services work a little differently.

    For TTS, Genesys Enhanced TTS is included without additional charge when used within Dialog Engine Bot Flows. So, in your case, using Amazon Polly Olivia NTTS through Genesys Enhanced TTS should not add the normal per-character Enhanced TTS charge while it is being used inside the Bot Flow.

    Outside of that scenario, Enhanced TTS does have character-based pricing, and more advanced voices such as NTTS are normally priced higher.

    For STT, I haven't personally used it in this type of implementation yet. From the documentation, Genesys provides its Enhanced STT engines as part of the Bot Flow experience, and I couldn't find a separate usage charge documented for the native engines beyond the normal Dialog Engine Bot Flow usage.

    If you use a third-party STT integration, such as Microsoft Azure or Google Cloud STT, additional BYOT usage charges may apply.

    These are the references I found:

    Genesys Enhanced TTS pricing
    https://help.genesys.cloud/articles/genesys-enhanced-tts-pricing/

    TTS engines overview
    https://help.genesys.cloud/articles/tts-overview/

    STT engines overview
    https://help.genesys.cloud/articles/speech-to-text-stt-engines-overview/

    Genesys Cloud pricing hub
    https://help.genesys.cloud/articles/genesys-cloud-pricing-hub/

    Since this involves billing and licensing, I think it would also be worth validating it with your Genesys CSM or opening a question with Customer Care, just to make sure the pricing for your specific subscription and configuration is confirmed.



    ------------------------------
    Raphael Poliesi
    ------------------------------



  • 5.  RE: How are you testing and comparing TTS voices in Genesys Cloud?

    Posted 9 days ago

    Thank you Raphael, this is really helpful and clears up my confusion around the TTS pricing. We're currently using Polly Olivia NTTS through Genesys Enhanced TTS in the Dialog Engine Bot Flow, so it's good to know there isn't an additional per-character TTS charge in this scenario.

    I'll also check with our Genesys team regarding the STT/bot usage pricing for our subscription. Thanks again for taking the time to look into this!



    ------------------------------
    Phaneendra
    Technical Solutions Consultant
    ------------------------------



  • 6.  RE: How are you testing and comparing TTS voices in Genesys Cloud?

    Posted 9 days ago
    Hello everyone,
    Great to hear about your Voice Bot POC in Genesys Cloud. This is a very relevant topic, and I'd be happy to share my experience.
    Based on what I've been working on, your idea of a dedicated test flow with sample phrases is the most practical approach. I have built a similar solution, and it has proven to be the most efficient way to listen to different voices without the tedious cycle of changing and republishing the main flow.
    Here are a few specific considerations I've learned along the way that you might find valuable.
    Your thought to use a Data Table is a very good one. This is a common and effective pattern. I've used a Data Table as a simple "playbook" for test phrases, which allows me to call them dynamically in a test flow.
    The documented best practice for comparing TTS engines is indeed to create a test flow in Architect, as you are planning. This approach has the major advantage of letting you change the TTS engine and voice under the Supported Languages settings for the flow, publish it once, and then listen to the sample text. You can easily update the test text and use a **Play Audio** action to hear the output. This is much faster and more structured than modifying your production flow for each comparison.
    A crucial point to remember when testing is that voices can interpret punctuation and numbers differently. For example, one engine might read "5.0" as "five point zero" while another might say "five period zero". This is especially important for your use case of routing, where you'll likely deal with account numbers or other identifiers. Test your specific data formats thoroughly.
    You are currently using Genesys Enhanced TTS with Amazon Polly Olivia (NTTS), which is a solid, high-quality choice. However, for context and future planning, be aware of a significant change: Genesys Enhanced TTS is transitioning to primarily use AWS Amazon Polly Neural Voices. After August 5, 2026, Google and Microsoft Azure voices will no longer be natively available via Enhanced TTS, with only a few exceptions for specific languages.
     
    Beyond Enhanced TTS, two other key options exist:
    This is the basic, free TTS engine. It's suitable for testing but is often described as having lower quality and a more "robotic" sound compared to other options.
    This is a newer offering from Genesys designed to deliver even more natural, human-like speech. This might be worth exploring for your POC if you want the most advanced voice quality.
    For the Australian market, while I don't have specific data, the most likely candidates for a natural-sounding Australian English voice are the Amazon Polly neural voices. Polly offers an Australian English voice, "Olivia" (NTTS), which you are already using. The other major option is Microsoft Azure, which also offers Australian English neural voices, but their native availability via Genesys Enhanced TTS is ending soon. If you are considering a third-party provider to keep the Azure option, be aware that you'll need to manage a separate integration and associated costs.
     
    I hope this information helps you with your POC. The test flow approach with Data Tables will give you a clear and controlled way to evaluate the best option for your callers.


    ------------------------------
    Camila Meneghini
    ------------------------------