Accuracy Methodology
How “accuracy” should be defined and measured for answering machine detection, and a plain statement of what has and has not been formally benchmarked for AMDY today.
Why “accuracy” is harder to define than it sounds
“99% accurate” is a claim that only means something relative to a specific test: a fixed set of calls, a known correct answer for each one, and a stated way of counting mistakes. Without those three things, an accuracy percentage is not measuring anything — it is a marketing number dressed up as a statistic. Rigorously, accuracy for AMD requires:
- A ground-truth dataset. A collection of real call recordings where a human listener, independent of the detection system, has confirmed for each one whether it was actually answered by a live human or a machine (voicemail, IVR, fax, etc.). This label has to come from something other than the detector being tested — using the detector’s own verdicts as the “truth” it is graded against is circular and proves nothing.
- A holdout set the model has not seen. If any of the labelled calls were used to train or tune the classifier, testing on them again inflates the result. A valid benchmark needs a held-out sample the system had no part in shaping.
- A stated sample size and composition. A number computed from 50 calls made on one carrier over one afternoon is not the same claim as one computed from 50,000 calls across many carriers, codecs, and greeting styles over months. The sample size and how it was drawn both belong in the reported number.
- A confusion matrix, not a single percentage. A single “accuracy” number can hide a system that is excellent at one kind of mistake and poor at the other. Reporting true positives, true negatives, false positives, and false negatives separately is what lets someone judge whether a system fits their risk tolerance.
False positives and false negatives, defined precisely
False positive
A call where a live human actually answered, but the system classified it as a machine. The practical consequence is a dropped or mishandled live call: a voicemail drop plays to a real person, or the call is disconnected. This is a lost contact, not just lost time.
False negative
A call where a machine (voicemail, IVR, etc.) actually answered, but the system classified it as a human. The practical consequence is an agent connected to a call that was never going to become a conversation, spending time listening to or talking over a recorded greeting before disengaging.
Any complete accuracy report states both rates separately, because a system tuned to minimize one will generally raise the other, and which one matters more depends on the campaign (see the tradeoff discussion on the what is AMD page).
What a confidence score actually means
A classifier that outputs, say, “87% machine” is reporting the model’s own estimated probability that the correct label is “machine,” based on how similar the call’s acoustic features are to the machine-labelled examples it was trained on versus the human-labelled ones. It is not a measurement of how often the model is right in general — that is a separate quantity (calibration) that itself has to be checked against ground truth. A model can produce confident-looking scores that are poorly calibrated if it has never been checked against labelled, independent data. The score is useful for setting an operating threshold (how confident the system must be before acting on a “machine” verdict), but the score itself is not evidence of overall accuracy.
What a rigorous AMD benchmark requires
- A labelled holdout set of real outbound call answers, hand-verified by a human listener who did not have access to the detector’s own output while labelling.
- A sample large enough, and drawn across enough carriers, codecs, and call types, to support a statistically meaningful confusion matrix rather than a handful of anecdotes.
- Reporting broken out by segment where behavior is known to differ — at minimum by codec/carrier path (mobile vs. landline vs. VoIP transcoded) and by line type, since compression and transcoding are exactly the conditions that stress an AMD system.
- A stated false-positive rate and false-negative rate, not a single blended accuracy figure, so a reader can judge fit for their own risk tolerance.
- Disclosure of whether any of the test data overlapped with data used to train or tune the model.
Where this leaves AMDY today
AMDY’s detection logs give us a large volume of the model’s own verdicts on real production calls, and we publish live, sourced statistics from that data on the features page — things like the human/machine split and false-answer supervision by line type. What that data does not give us is ground truth: none of those verdicts have been independently checked by a human listener against what actually answered the call. That means we can describe what the detector says about itself, but we cannot currently compute a validated accuracy, precision, recall, or false-positive rate, because doing so honestly requires the labelled holdout set described above, and we have not built and run that benchmark.
Building that benchmark, on our own criteria above, is on our roadmap: a hand-labelled holdout set across multiple carriers and codecs, evaluated with a full confusion matrix and reported with sample size and methodology alongside the numbers. Until that exists and is published, we are not going to quote an accuracy percentage we cannot source, and we would treat any AMD vendor’s unsupported “99% accurate” claim with the same skepticism.