This is a great question.
In my experience, the first step is always to validate the prompts in a test flow before moving them to production, since pronunciation and overall quality can vary depending on the selected TTS engine, voice, language, and even the specific phrases being synthesized.
For multilingual environments, I also think it's important to have native speakers review the prompts, especially for abbreviations, numbers, dates, proper names, and industry-specific terminology. Those are usually the cases where pronunciation issues become more noticeable.
One limitation worth keeping in mind is that SSML isn't supported in call flows, secure call flows, or in-queue flows, which reduces the ability to fine-tune pronunciation in those scenarios. SSML is only available with supported third-party TTS providers in bot flows.
For business-critical messages, another approach is to use pre-recorded audio whenever pronunciation must be absolutely consistent.
I'm also curious to hear how other organizations handle this. Do you have a formal QA process for TTS content, or do you rely mainly on business users and native speakers to validate prompts before deployment?
If you're looking for implementation guidance beyond community best practices, your Customer Success Manager (CSM) or Genesys Customer Care may also be able to provide recommendations based on your specific use case and deployment architecture. Although... that's usually the process-unless you happen to be the CSM! 😄
Hope this helps in some way. Best regards!
------------------------------
Raphael Poliesi
------------------------------