Discussão muito interessante.
Juan, a última informação que você compartilhou me parece especialmente importante: vocês utilizam auto-answer, mas, quando o problema ocorre, a interação passa a tocar para o agente e o teste "Able to reach Genesys services via DNS" falha.
Nesse cenário, eu trataria o Not Responding mais como uma consequência do problema do que como a causa principal. A investigação de conectividade/DNS que vocês estão realizando parece ser um caminho importante para encontrar a causa raiz.
Também gostei bastante da automação apresentada pelo Phaneendra como forma de reduzir o impacto operacional enquanto a investigação continua. Eu apenas adicionaria alguns controles para evitar que a automação acabe mascarando um problema persistente:
-
limitar a quantidade de resets automáticos por agente dentro de uma janela de tempo;
-
registrar horário, usuário, site/divisão e quantidade de ocorrências;
-
após atingir um limite, parar de retornar automaticamente o agente para Idle e gerar um alerta para investigação;
-
correlacionar os horários das ocorrências com os diagnósticos WebRTC e logs da infraestrutura.
Assim conseguimos separar bem duas responsabilidades:
Mitigação: recuperar rapidamente a operação.
Observabilidade: preservar informações suficientes para identificar e corrigir a causa raiz.
Esse tipo de automação com Trigger + Workflow é um ótimo exemplo de como o Process Automation pode ajudar a reduzir impacto operacional sem substituir a investigação técnica do problema.
Juan, seria muito interessante saber o resultado final da análise do seu time de infraestrutura quando vocês identificarem a causa.
__________________________________________________________________________________________________________________________________________________________________
Very interesting discussion.
Juan, I think the latest detail you shared is particularly important: your queues use auto-answer, but when the issue occurs, the interaction starts ringing for the agent instead, and the "Able to reach Genesys services via DNS" diagnostic test fails.
In this scenario, I would look at Not Responding more as a consequence of the underlying issue rather than the primary cause. The connectivity/DNS investigation your IT team is already performing seems like an important path toward identifying the root cause.
I also really like the automation Phaneendra shared as a way to reduce the operational impact while the investigation continues. I would just add a few guardrails so the automation does not unintentionally hide a persistent problem:
-
limit the number of automatic resets per agent within a defined time window;
-
record the timestamp, user, site/division, and number of occurrences;
-
once a threshold is reached, stop automatically returning the agent to Idle and generate an alert for investigation;
-
correlate the occurrence timestamps with WebRTC diagnostics and infrastructure logs.
This creates a useful separation between two responsibilities:
Mitigation: quickly restore the operation.
Observability: preserve enough information to identify and fix the root cause.
This type of Trigger + Workflow automation is a great example of how Process Automation can help minimize operational impact without replacing the technical investigation of the underlying issue.
Juan, it would be very interesting to hear the final outcome from your infrastructure team once the root cause is identified.
------------------------------
Matheus Mendonca
------------------------------