I have been a little quieter in the community lately, even though I would really like to be more active here.
The last couple of months have been incredibly busy, but also full of learning. Together with the team, I have been involved in several Genesys Cloud projects across agentic AI, Virtual Agents, Agent Copilot, Knowledge, Quality, Analytics, and automation.
Across these initiatives, our team is currently working with 11 Virtual Agents. Nine are already in production, while two are still being tested or developed. They support different journeys across retail, financial services, healthcare, loyalty programs, and internal training.
A big part of my recent work has been understanding what happens after a Virtual Agent goes live. We have been reviewing real conversations, following customer journeys, measuring retention and resolution, evaluating response quality, and investigating situations where the standard metrics do not show the complete picture.
In one of the operations we are currently monitoring, retention has remained close to 85%. One finding that really caught my attention was that some conversations initially classified as abandonment had already achieved their intended outcome before the customer left. It was a good reminder that the technical end of a conversation does not always represent the actual customer outcome.
We have also been improving how we evaluate Virtual Agent quality by looking at operational results, conversation content, quality evaluations, and agent behavior together. Recent quality baselines have remained above 89%, helping us better understand where the experience is consistent and where we still have opportunities to improve.
Another area I have been exploring is AI Scoring. In a recent analysis of more than 27,000 evaluated questions, the agreement between AI and human reviewers reached approximately 93%. The overall number was interesting, but the differences were even more valuable. They helped us identify which questions depend more heavily on human interpretation and where evaluation criteria may need to be clearer.
Knowledge has also taken up a good part of my time. We have been working with large knowledge bases, reviewing content quality, identifying duplicated or conflicting information, improving training phrases, and observing how content changes affect the answers generated by Virtual Agents.
One of the most interesting recent projects involves using a Virtual Agent as a training assistant. Instead of interacting only with customers, the agent helps employees explore procedures, understand operational content, and receive guidance during learning activities. It has been great to see how the same technology can support completely different experiences.
Alongside these projects, I have also been contributing to internal tools, dashboards, documentation, and handovers that help quality teams, knowledge managers, workforce teams, analysts, and support teams work with these solutions more effectively.
That is basically why I have been less active here. Most days have been split between testing, analyzing production behavior, documenting findings, discussing improvements, and learning something new from real interactions.
One of my biggest lessons from this period is that putting an AI agent into production is only the beginning. The most interesting part comes afterward, when real customers start interacting with it and challenge many of our initial assumptions.
I have collected quite a few lessons and unexpected findings over the past few months, and I hope to start sharing more of them here.
What have you been working on lately? I would love to hear about it.
#ConversationalAI(Bots,VirtualAgent,etc.)------------------------------
Mateus Nunes
CX Manager at Solve4me Solucoes em Tecnologia Ltda
------------------------------