What Is AI Answering Machine Detection?
A plain-language explainer, starting from zero, of what AMD does, why outbound dialers need it, and how the AI approach differs from the timing-based logic most dialers ship with.
The problem AMD exists to solve
An outbound dialer places a call and, a few seconds later, something picks up. It might be a live person saying “hello?” It might be a voicemail greeting. It might be a fax tone, a disconnected-number message, or dead air. The dialer has to decide, in well under a second, which of those it is hearing, and then decide what to do next: hand the call to a waiting agent, play a pre-recorded voicemail drop, or hang up.
That decision is Answering Machine Detection, usually shortened to AMD. It sits between the moment the far end picks up and the moment the dialer routes the call somewhere. Get it right and agents spend their time talking to humans. Get it wrong in one direction and a live prospect gets silence and hangs up. Get it wrong in the other direction and an agent sits listening to a voicemail greeting, burning paid time on a call that was never going anywhere.
Dialers need AMD because agent time is the most expensive resource in an outbound campaign. Answering the “human or machine?” question automatically, and quickly, is what makes predictive and progressive dialing modes possible at all: the system can dial ahead of agent availability specifically because it plans to filter out the calls that never needed a human.
How traditional, signal-based AMD works
The AMD built into most PBX and dialer software (Asterisk’s stock AMD app, and by extension ViciDial, FreeSWITCH, and Issabel deployments that use it) does not listen to audio the way a person does. It measures a small number of acoustic properties of the call signal and applies fixed thresholds and timers to them.
- Silence detection. It watches for gaps of silence after the call connects. A human answer tends to produce a short greeting followed by a pause waiting for a response. A voicemail greeting tends to run on longer before any pause.
- Energy/word-count heuristics. It counts how many distinct bursts of speech energy occur in the first few seconds and how long each one lasts. A short “Hello?” looks different, on paper, from a 15-second scripted greeting.
- Beep detection. Many voicemail systems end their greeting with a distinct tone before recording starts. Detecting that tone is one of the more reliable traditional signals, when it is present and undistorted.
- Fixed timers. All of the above run inside a hard-coded analysis window, commonly 2.5–5 seconds. If the decision isn’t made by then, the system has to guess or default.
This approach was designed for narrowband, uncompressed analog telephony. It struggles on modern calls because mobile networks, VoIP codecs, and carrier transcoding all reshape the very things it measures: silence gets filled with comfort noise, greetings get compressed and clipped, and beep tones get distorted or dropped entirely. A live human who pauses half a second too long, or a voicemail greeting that happens to have a short opening line, can each fall on the wrong side of a fixed threshold.
Why this shows up as errors, not just imprecision
Because the traditional approach is a handful of rules with numeric cutoffs, it fails in the same two directions no matter how the cutoffs are tuned: it sometimes calls a live person a machine, and it sometimes calls a machine a live person. Tightening one threshold to fix one kind of mistake reliably makes the other kind of mistake more common. That tradeoff is fundamental to the rule-based approach, not a bug that better tuning removes.
How AI/ML-based AMD is different
An AI-based AMD system replaces the fixed rules with a trained audio classification model. Instead of checking a few hand-picked measurements against thresholds someone chose in advance, the model has learned statistical patterns from many examples of what human answers and machine answers actually sound like across a wide range of call conditions.
- It uses more of the signal. Rather than reducing the call to silence gaps and energy bursts, a classifier typically works from a spectral representation of the audio (something like a mel-spectrogram) that preserves far more of the acoustic detail a person’s ear would use.
- It generalizes across conditions. A model trained on calls across many codecs, carriers, and greeting styles is not relying on one fixed timer or one fixed tone; it is weighing many learned patterns at once, which tends to be more robust to a compressed or clipped connection.
- It outputs a confidence score, not just a verdict. Instead of a bare yes/no, a classifier typically produces a probability, e.g. “92% likely voicemail.” That score can be used to set an operating point: a campaign that cannot tolerate dropping live calls can require a higher confidence threshold before treating an answer as a machine, at the cost of catching fewer voicemails.
- It is still not infallible. A trained classifier does not eliminate the false-positive / false-negative tradeoff described above — it can shift where that tradeoff sits, and it can be more consistent across noisy real-world audio, but every AMD system, rule-based or learned, sits somewhere on that same curve.
The tradeoff that never goes away: false positives vs. false negatives
Whatever detection method is used, there are exactly two ways to be wrong, and they cost the campaign in different currencies.
False positive: a live human marked as a machine
The system decides “machine” and the dialer responds accordingly, typically by playing a voicemail drop or hanging up. If the answer was actually a live person, that person hears silence, a click, or a canned message that doesn’t fit the conversation, and hangs up confused or annoyed. The call is lost outright. This is the more expensive mistake for most campaigns: it destroys a contact that was already reached, not just wastes time getting to one.
False negative: a machine marked as a human
The system decides “human” and routes the call to a live agent. If the answer was actually a voicemail, the agent listens to (or talks over) a greeting that was never going to respond, then has to hang up and move on. No contact is lost, but paid agent time is spent on a call that had zero chance of converting. At scale, across thousands of calls a day, this is a direct hit to how many live conversations each agent can have per hour.
Because these two error types trade off against each other, there is no single “best” setting for an AMD system independent of what a specific campaign values. A campaign selling to a small, valuable list where every live contact matters will usually tune toward fewer false positives, accepting more wasted agent time on voicemail. A high-volume campaign optimizing strictly for agent talk-time efficiency may accept a slightly higher risk of dropping a live call in exchange for filtering more voicemail before it reaches an agent.
Where AMDY fits
AMDY is an AI/ML-based AMD classifier built to drop into ViciDial, Asterisk, FreeSWITCH, and Issabel deployments in place of the stock signal-based detector, without requiring a rebuild of the dialer configuration. It listens to the answer audio, classifies it, and returns a verdict with a confidence score fast enough to fit inside a normal call-answer window.
We do not treat “AI-based” as a synonym for “error-free.” The tradeoff described above is inherent to the problem, not specific to any one implementation. For a rigorous look at how accuracy claims for AMD should be defined and measured — and an honest account of what has and has not been formally benchmarked for AMDY — see the accuracy methodology page.
To see live, sourced statistics from AMDY’s own detection logs (human/machine split, false-answer supervision by line type, and more), see the features page. For plans and pricing, see the pricing page.