On this page
Key Points
- According to Userflow, low-risk answers should escalate below 0.4 plus one corroborating signal, medium-risk changes should ask the user to confirm below 0.7, and high-risk actions stay gated or human.
- An ICLR 2024 study by Xiong and colleagues found LLMs overconfident; EverHelp's illustration is that a bot claiming 90% certainty can be closer to 75% right.
- California's AB 1609, the Right to Human Customer Service Act, requires businesses with more than $500 million in gross annual revenue to let chatbot users request a live human agent.
The handoff moment: where an AI agent's confidence threshold decides whether a person takes over.
Quick Answer
A handoff confidence threshold refers to the score below which an AI support agent stops answering and hands the conversation to a person, the line between human-in-the-loop and autonomous handling. I'd set it near 0.7 for general questions, raise it for regulated topics, and keep high-risk actions human. Models overstate their own confidence. Tune the floor per intent from real transcripts. Some requests skip the dial entirely: under California's AB 1609, large businesses must let a customer who asks for a person reach one.
Every AI support agent already runs on a handoff threshold, whether or not anyone set it on purpose. The real question is whether yours was chosen for your risk or inherited from a vendor default.
A useful way to see the choice comes from outside support. According to John Capobianco of Itential, AI agents in network engineering operate at three layers of autonomy: "human in the loop," where the agent proposes a config and asks before applying it; "on the loop," where the human monitors loosely while the agent makes tickets and sends email recaps; and "full autonomous."
Support conversations sit on the same three layers. The confidence threshold decides which layer each one lands in. Set it well and routine questions run on their own while risky ones wait for a person. Set it badly and the agent either answers what it should have escalated or escalates what it could have closed.
That is why I treat handoff as a handoff dial rather than a switch. For buyers, the practical test is simple to state: does a platform let you set that dial per intent and risk tier, or does it offer one global slider?
Below, I set out the range I'd start with, why the raw confidence number runs high, the signals and hard rules that should override it, and where I expect the practice to head over the next 12-24 months. One position on the dial is already fixed for some teams. In California, AB 1609 means large businesses cannot leave an explicit request for a person to the score.
An AI support agent should hand off when its confidence in a grounded answer falls below a floor you set by risk, not when a vendor default says so.
For general questions, I'd start near 0.7. Most teams treat handoff as a switch. It works better as a dial. A handoff confidence threshold is the score below which the agent stops answering and routes the conversation to a person. Set it low and wrong answers leak through; set it high and people field questions the agent could have closed.
According to Userflow, whose escalation guidance was published on September 24, 2026, the right starting point depends on what an answer can break: low-risk answers, medium-risk changes and high-risk actions each get a different rule, and the riskiest get no confidence cutoff at all. The same guidance treats a model's stated confidence as weak evidence on its own. That is why I treat the dial as retrieval confidence, checked against other signals, rather than the model's self-report.
The stakes are also shifting from commercial to legal. California's AB 1609, the Right to Human Customer Service Act, turns an explicit request for a person into a compliance question for large businesses rather than a tuning choice. The published evidence does not yet show how accuracy changes by confidence bucket for support agents, so treat the range below as a starting point to test against your own transcripts.
Three questions this article answers
- What confidence threshold should my AI support agent use before handing off to a human?
- Can I trust the confidence score my AI chatbot reports?
- Which signals and rules should trigger a human handoff regardless of the score?
Start with the first. The number you pick there only works if the next two hold up.
What confidence threshold should you start with?
Start general support answers at a 0.7 confidence floor after a grounding check, raise it to 80-90% or higher for regulated topics, and give high-risk actions no dial at all.
An analysis of 12 sources shows no agreement on a single universal number, but clear agreement that the floor should rise with risk. According to Userflow, low-risk answers should escalate below 0.4 plus one corroborating signal (based on the Rasa default), medium-risk changes should ask the user to confirm below 0.7 (based on a Rasa NLU example), and high-risk actions get "None; always gated or human" (based on Salesforce per-action guidance). Practitioner guidance from the support outsourcing side lands in the same place: a 60-70% floor for general support and 80-90% or higher for regulated topics.
I call this the risk-tiered dial: one threshold per risk tier, each with its own floor, instead of one global slider. It is also the first setting I'd check when comparing help desk software with live chat. Zazachat's reviews are written and edited by our editorial team, led by Daniel Calloway, and in my view a handoff dial the buyer cannot inspect is a gap any review should name. A common misconception is that handoff is a yes/no switch. The reality is that it works like a ladder, and each rung carries a different cost of being wrong.
| Risk tier | Typical examples | Starting floor | What happens below it |
|---|---|---|---|
| Low risk (read and guide) | Read-only answers and how-to guidance | 0.4 plus one corroborating signal | Clarify first, then hand off |
| General support answers | Grounded answers drawn from the knowledge base | 60-70%; start at 0.7 | Hand off with full context |
| Medium risk (reversible change) | Updating a profile field, resending an invite, toggling a notification | 0.7 | Ask the user to confirm in chat |
| Regulated topics | Legal, medical and payment questions | 80-90% or higher | Route to a qualified person |
| High risk (irreversible action) | Refunds, deletions, permission changes, anything writing to billing | None | Always a person or a per-action approval gate |
Why start at 0.7 rather than the bottom of the band? A wrong low-risk answer, in Userflow's words, "costs a follow-up question." A wrong general answer costs more, because the customer may act on it. In my view, starting at the top of the general band and loosening it only after transcript review is the safer direction to tune from. Loosening a strict floor is a reversible experiment. Winning back a customer who acted on a confident wrong answer is not.
Why does the grounding check come before the number?
The sequence matters as much as the setting. The confidence floor is a second gate, applied only after the answer has been grounded against the knowledge base, and an answer with no supporting passage should escalate whatever its score. Userflow defines the low-confidence trigger the same way: "No matching article retrieved, plus one corroborating signal."
- Hard rules first. Explicit human requests and high-risk actions exit before any score is read.
- Grounding second. No retrieved passage means no autonomous answer.
- Confidence floor third. Only grounded answers are compared against the tier's threshold.
The first gate is no longer only a design choice. According to Quarles & Brady, California's AB 1609 requires businesses with more than $500 million in gross annual revenue that make a customer service chatbot available to people in California to let customers request a live human agent. For those businesses, a request for a person cannot be scored away.
The broader principle is to escalate when the expected cost of continuing exceeds the cost of involving a person, which is exactly what a tiered floor encodes. In practice, the dial is a ladder, not a switch. The takeaway: the more damage a wrong answer can do, the higher the floor sits.
What will matter most for AI handoff in the next 12-24 months?
Over the next 12-24 months, handoff will move from one global cutoff to risk-tiered thresholds, with explicit requests for a person and high-risk actions bypassing the score entirely.
Three signals carry most of the weight in my view. Each is visible today, even if it has not yet reached most vendor settings pages.
| Prediction | Weak signal today | Why it matters | Source |
|---|---|---|---|
| Risk-tiered cutoffs replace one global slider | Published escalation guidance now ties starting thresholds to risk level, pairing low-risk answers with a corroborating signal and gating high-risk actions entirely. | A single cutoff either over-escalates simple questions or lets risky actions through. Per-intent or per-tier settings become a buying criterion. | Userflow, AI agent escalation criteria (September 2026) |
| The model's self-reported confidence stops being the dial | Research on verbalized confidence finds models overconfident, with accuracy in each band landing below the figure they state. | Buyers should ask how a platform calibrates its score and which retrieval signals feed it. An uncalibrated score passes more wrong answers than its setting implies. | Xiong and colleagues, ICLR 2024 |
| An explicit request for a person becomes a hard override | In logged WhatsApp conversations, customers who asked for a person got the same rigid menu, while design documents promised automatic escalation on frustration. | Practitioner escalation guides already treat an explicit request for a person as a trigger in its own right. Teams need a separate, auditable path for it. | Customer support practitioner discussion (June 2026) |
One nuance keeps the second prediction from going too far. Self-assessment is useful when it is situational. In one clinical example, an AI reports that its confidence in a diagnosis is lower than usual because of an atypical presentation, and it strongly recommends human review. That is the behavior I'd want from a support agent: a flag tied to why the case is unusual, not a bare number.
Two developments would change this forecast. If models' stated confidence came to track their real accuracy, a single cutoff could work again. And if platforms began publishing accuracy by confidence bucket, the tuning debate would move from judgment to evidence.
What most buyers miss: the threshold number is not where most handoffs fail. Benchmarking shows first-contact resolution runs about 19% lower for customers who are transferred, and 74% of customers are frustrated when they have to repeat information they've already given. A perfectly tuned dial still loses if the person on the other end starts from zero.
Why can't you take the confidence score at face value?
Because models overstate their own certainty. A threshold read at face value passes more wrong answers than the number implies, so the score needs calibrating before the dial means anything.
Every range in the previous section assumes the score is honest. In practice, it often is not. According to Userflow, which cites Xiong et al. (2023), GPT-4's stated confidence across question-answering benchmarks produced an average AUROC of 62.7 percent for vanilla verbalized confidence. AUROC measures how reliably a score ranks correct answers above incorrect ones, so a middling result means the score is a loose sorter rather than a precise gate. In other words, the model's self-reported number tells you something, but not enough to hand it the handoff decision on its own.
The miscalibration also runs in one direction. According to EverHelp, an ICLR 2024 study by Xiong and colleagues found that LLMs tend to be overconfident when they state their own confidence, with accuracy in each confidence band landing "well below" the number the model reported. EverHelp's illustration is blunt: a bot claiming 90% certainty "can be closer to 75% right." What this means for the dial is specific. If you set your floor against raw self-reported scores, part of what clears the bar is less accurate than the setting promises.
This is why I treat handoff as a retrieval-confidence dial, not a self-report dial. The score worth gating on reflects whether the agent found and grounded a source for its answer, rather than the model's own sense of certainty. Low-confidence retrievals are where I expect wrong answers to concentrate, and that is the zone a higher floor removes first.
What does a mis-set threshold cost in each direction?
Both directions carry a real cost, and they land on different people.
| Setting | What happens | Who pays |
|---|---|---|
| Too loose (low floor) | The agent answers questions it half-understands and passes overconfident wrong answers | The customer first, then the brand's credibility |
| Too tight (very high floor) | Simple, well-documented questions reach people who did not need to see them | The support team, through queue load and slower replies for everyone |
| No working exit | The bot keeps answering even after the customer asks to leave it | The customer, who escalates in frustration or abandons the channel |
The loose side is the one I would worry about most. In a survey of 1,000 US and Canadian chatbot users, 62% named the chatbot misunderstanding them as the root cause of escalation, ahead of underperformance at 22% and emotional triggers at 16%. Misunderstanding is exactly what a low floor lets through. The damage compounds quickly: 80% of users lose confidence after one or two incorrect responses. The takeaway: a loose threshold does not avoid escalations. It delays them until trust is already spent.
The opposite failure is quieter but no less damaging. In logged WhatsApp conversations, a user typed exactly "talk to a human" and "I need to speak to a person." Each time, the bot returned "the same rigid menu of options" and never transferred, and the user's frustration escalated until they insulted the bot or left the channel. Published design documents and patents describe "advanced concepts for detecting user frustration and automatically escalating to a human." The actual logs show different behavior. In practice, a handoff rule is only as good as its behavior in live transcripts, which is where I would test any threshold before trusting it.
So the recommended range holds, but only when the number feeding it is honest and the exit behind it works. Getting both right is an implementation problem, and it is the one the next section takes on.
How do you implement and tune the handoff dial?
Pair the score with corroborating signals, let hard rules override it, and build a transfer that carries state. Then tune each intent from real transcripts.
A single number should never decide alone. According to Petronella Technology Group, which organizes AI-to-human handoff logic into seven named trigger categories, the first is low-confidence or ambiguous intent: the agent can't confidently classify the request, or multiple intents compete. That second clause is the one most setups miss, because a leading intent can clear the floor while a rival sits right behind it.
Which signals should back up the score?
I'd recommend pairing the score with four corroborating signals. Any one of them should be able to pull a passing answer back to a person.
- Competing intents: a near-tie between two intents means the top score overstates certainty, so treat it as a low-confidence result.
- Missing required parameters: if the answer depends on an order number, plan or account detail the agent has not collected, a confident retrieval is answering the wrong question.
- Contradictory cues: a customer who confirms a fix and then repeats the problem is telling you the answer did not land.
- Fallback spirals: repeated rephrase or fallback turns mean the conversation is looping, and the next turn should be a handoff.
In practice, the score sets the floor and the signals decide the close calls.
Which rules should bypass the score entirely?
Some requests should never reach the dial at all. I treat three as hard routes.
- An explicit request for a person. According to Quarles & Brady, California's AB 1609, the "Right to Human Customer Service Act," requires covered businesses to disclose chatbot use and to let customers request a live human agent, with a good faith effort to connect them within 15 minutes. Outside its coverage, the logic still holds: a customer who asks for a person has already answered the handoff question.
- Requests that need legal, financial or policy sign-off. Petronella lists these as their own trigger category. An agent cannot grant an exception it has no authority to grant, however strong its retrieval.
- High-risk actions. As set out earlier, these stay gated or human whatever the score says.
For everything in between, the cleanest decision rule is cost-based. Hand off when the expected cost of continuing exceeds the cost of involving a person.
What makes a transfer actually carry the conversation?
A handoff that announces itself but moves nothing is worse than no handoff. Practitioners describe bots that tell the customer a human agent is on the way, after which nobody joins. The stated root cause is structural: the escalation logic was built to send a message, not to hand off state. The workflow ends at the "transferring you" reply, with no mode change, no context passed, and no task created for a human agent.
The fix mirrors the failure. A working transfer does three things:
- It switches mode, so the bot stops replying once the handoff fires.
- It passes the transcript, the detected intent and any collected parameters to the agent.
- It creates a routed task in the human queue, so the conversation has an owner.
Finally, tune per intent rather than globally. Start each new intent at the conservative end, review its transcripts, and lower the floor only where answers hold up. Autonomy should be earned one intent at a time. The takeaway: a dial you can set per intent or risk tier is worth more than a smarter default.
Outlook - next 12-24 months
Where chatbot handoff thresholds are heading
Forecasts on how support teams will set, calibrate, and override the confidence cutoff that sends a chatbot conversation to a human agent.
What changes in handoff thresholds next
Read each forecast against your own risk tiers and staffing before changing where your bot hands conversations to people.
For crisis, self-harm, and other high-risk situations, providers will route to people or crisis resources by category rather than by confidence score. High-risk actions will be treated as always gated or human, whatever the model's certainty.
Large businesses serving California customers will treat a request for a human as a hard override that bypasses confidence scoring. AB 1609 requires companies above $500 million in national revenue to disclose chatbot use and give customers a path to a live agent.
Support teams will set separate handoff thresholds by risk tier instead of one global slider. Low-risk answers will escalate below about 0.4 confidence, general support floors will sit in the 60-70% range, and high-risk actions will always be gated or handed to a person.
Vendors will stop using a model's own stated confidence as the handoff dial. Instead, they will combine retrieval results with corroborating signals, because verbalized confidence overstates accuracy.
Handoff rules will give more weight to low intent confidence and competing intents than to response delays. Misunderstanding is the leading reason customers leave a bot for a person.
Many support teams will find that failed transfers, not mis-set thresholds, do the most damage. They will invest in mode switching and in passing conversation state and context to human agents, rather than in further tuning of the cutoff.
Low-confidence indicators Published guidance now ties starting thresholds to risk level. It cites a Rasa default of 0.4 for low-risk answers and recommends a 60-70% confidence floor for general support. California's AB 1609, the Right to Human Customer Service Act, applies to businesses with more than $500 million in gross annual revenue. Meanwhile, logged WhatsApp conversations show bots returning the same rigid menu to users who typed that they wanted to talk to a human. Research cited in vendor guidance found an average AUROC of 62.7 percent for GPT-4's verbalized confidence. A bot claiming 90% certainty can be closer to 75% right. In a Botpress survey of 1,000 US and Canadian chatbot users, 62% named the bot misunderstanding them as the root cause of escalation. Only 8% said they escalate quickly because of delays. Practitioners report bots that announce a transfer and then end the workflow there, with no mode change, no context passed, and no task created for a human agent. In Pichowicz et al., 2025, researchers tested 29 mental-health chatbot agents against escalating suicidal-risk scenarios. Only three supplied the correct regional emergency number without extra prompting.
Sources behind the handoff forecasts
Each public source below is paired with the specific line on thresholds, calibration, or escalation law that a forecast relies on.
| Source | What it states | Forecasts it backs |
|---|---|---|
| Synthetic Relational Force - The Human Anchor Protocol [Substack / Newsletter] | Pichowicz et al., 2025: 29 mental-health chatbot agents were tested against simulated escalating suicidal-risk scenarios. “It does not yet prove the best thresholds, timing or handoff methods for every population and setting, but that uncertainty is part of the protocol, not an…” | Safety and high-risk cases bypass confidence entirely |
| AI Agent Human Escalation Criteria - Userflow [Web source] | High-risk actions: "None; always gated or human" (based on Salesforce per-action guidance). “Confidence is a necessary gate and an unreliable one.” Low-risk answers: escalate below 0.4, plus one corroborating signal (based on the Rasa default). Userflow cites Xiong et al. (2023), which tested GPT-4's stated confidence across question-answering benchmarks. It found an average AUROC of 62.7 percent for vanilla verbalized confidence. |
Safety and high-risk cases bypass confidence entirely Risk-tiered cutoffs replace one global threshold Self-reported model confidence gets discounted |
| Human-AI Collaboration: The Partnership Imperative [Podcast] | [11:33] Speaker 1's example: an AI reports that its confidence in a diagnosis is lower than usual because of an atypical presentation, and strongly recommends human review. | Safety and high-risk cases bypass confidence entirely |
| New Customer Service Requirements for AI Chatbots - Quarles [Web source] | Coverage threshold: A "large private business" is one with more than $500 million in gross annual revenue nationally that provides goods and services to customers and makes a customer service chatbot available to a person in California. “Fifteen minutes to connect with a human being.” | Explicit requests for a person override the score |
| If a customer asks to speak to a human, should the chatbot [Community / Forum] | In one logged WhatsApp case, a user typed exactly "talk to a human" and "I need to speak to a person." Each time, the bot returned "the same rigid menu of options" and never transferred. “The honest answer most ops teams won't say out loud: if someone explicitly types "talk to a human," the bot should transfer. Full stop.” | Explicit requests for a person override the score |
| AI-to-Human Escalation: When Support Bots Need an Agent - Soon [Web source] | The customer explicitly asks for a person. “A good escalation is not a failure.” | Explicit requests for a person override the score |
| Customer Service Escalation Process: AI Handoff Guide - EverHelp [Web source] | 60-70% for general support. “Despite escalations carrying a bad reputation, the handoff itself isn't the problem.” ICLR 2024 study (Xiong and colleagues): LLMs tend to be overconfident when they state their own confidence. Accuracy within each confidence band lands "well below" the number the model reported. Zendesk 2026 CX Trends report: 74% of customers are frustrated when they have to repeat information they've already given. 81% want the next rep to continue exactly where the last one left off. |
Risk-tiered cutoffs replace one global threshold Self-reported model confidence gets discounted Handoff mechanics matter more than the number |
| 50+ Customer Support Chatbot Statistics in North America for 2026 [Web source] | Root causes of escalation: the chatbot misunderstanding the user (62%), underperformance (22%), and emotional triggers (16%). “Chatbot users have a low tolerance for repeated friction. Across all common interaction issues, the threshold for escalating to a human is consistently low.” | Intent confusion, not wait time, drives escalation |
| AI Customer Care Human Escalation Triggers When to Hand Off [Web source] | Low intent confidence: the agent can't confidently classify the request, or multiple intents compete. “Good triggers reduce escalations that don’t need to happen, prevent long frustrating loops, and protect customers when the conversation involves risk, complex…” | Intent confusion, not wait time, drives escalation |
| Your AI agent says "transferring you to a human" and then [Community / Forum] | Stated root cause: the escalation logic was built to send a message, not to hand off state. The workflow ends at the "transferring you" reply, with no mode change, no context passed, and no task created for a human agent. “the escalation logic was designed to send a message, not to actually hand off state.” | Handoff mechanics matter more than the number |
What would move the handoff cutoff
These scenarios, from better-calibrated models to looser enforcement of human-access rules, would shift or reverse the forecasts.
Not without caveats
75 rests on the firmest ground here, while 61 is the call we would revise soonest.
- If three developments would change these forecasts.
- If models' stated confidence came to track their real accuracy, a single cutoff could work again.
- If California's AB 1609 were narrowed or weakly enforced, the pressure for hard human-request overrides would fade. And if AI agents proved track records near 995 correct out of 1,000 tickets, teams could move toward fuller autonomy with far fewer handoffs.
What should you change about your handoff threshold first?
Move it off the vendor default, tie it to risk, and raise the floor for general answers, because the weakest retrievals are where wrong answers gather.
The case rests on one uncomfortable fact. A model that calls itself nine-in-ten sure can be right only about three times in four, so a dial read at face value is always looser than it looks. According to Userflow, the practical response is tiered: the more an answer can break, the less the score alone should decide.
The number is the easy part. A threshold nobody has checked against real transcripts is a guess, and a transfer that drops context wastes even a well-chosen one. For large businesses serving California, AB 1609 already takes one decision away from the dial entirely. An explicit request for a person goes to a person.
So here is the next step I'd take. Pull last month's AI-answered conversations, group them by confidence score, and have a reviewer mark each answer right or wrong. The first bucket where the wrong marks start piling up is where your floor belongs.
Summarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What else do support leaders ask about AI handoff thresholds?
These are short answers to the common questions about AI handoff thresholds, from where to set the floor to when a customer's request overrides it entirely.
When should my AI support agent hand off to a human?
Hand off when confidence in a grounded answer drops below the floor for that risk tier, when a corroborating signal fires, or when the customer asks for a person. Any one of the three is enough.
What is a handoff confidence threshold?
A handoff confidence threshold is the score below which an AI agent stops answering and routes the conversation to a human. I treat it as a dial you tune per intent, not a switch you set once.
Should every intent use the same threshold?
No. According to Userflow, low-risk answers, medium-risk changes and high-risk actions each need a different rule. What this means in practice: a password-reset answer and a refund decision should never share one cutoff.
Does asking for a human override the confidence score?
Yes, and it should. Under California's AB 1609, large businesses must honor that request, so route it straight to a person rather than weighing it against the score.
Can I rely on the model's own confidence rating?
Not on its own. Models tend to be overconfident when they rate themselves, so a raw self-score passes more wrong answers than it suggests. Pair it with retrieval checks.
What counts as ambiguous intent?
Ambiguous intent means the agent can't confidently classify the request, or multiple intents compete. Treat it as a handoff trigger even when the leading score passes.
Will a higher threshold flood my team with escalations?
It will add some volume, and the public evidence does not yet size that increase. My advice is to raise the floor one intent at a time and log escalations before and after each change.
Read next
Right article, wrong answer: where AI support agents break
45% of AI queries produce errors even from correct sources. Learn to diagnose retrieval vs generation failures and fix your AI support agent today.
Read
AI support agents pay off past a 60% deflection rate
AI support agents cover their cost when true deflection reaches 60%. Learn how reopen rate, ticket routing, and QA monitoring determine whether your…
Read
Your AI agent's accuracy is set by knowledge base quality
AI support agent accuracy is determined by knowledge base quality. See how RAG grounding, governance, and retrieval config drive correct answers - not…
Read