FloorTone › Agent audio quality metrics
If it isn't measured, it's an opinion
Agent Audio Quality Metrics
The short version: measure agent audio at three levels — what the network did (your platform's MOS-style scores), what a human hears (QA defect codes), and what it cost you (repeats, corrections, callbacks). The single most common mistake is treating the first level as the whole answer: a call can score a healthy MOS and still contain a barking dog, because MOS-style network metrics measure transport quality, not background noise. This page gives you all three levels, including a defect-code scheme you can paste into your QA form this week.
- Level 1 · The networkMOS-style transport scores: packet loss, jitter, latency
- Level 2 · A human earQA defect codes A1–A8, tagged on scored calls
- Level 3 · The costrepeats, corrections, callbacks
Level 1: what the platform measures (the network)
Every serious CCaaS platform reports per-call transport quality — usually a MOS-style estimate (the Mean Opinion Score convention runs 1 to 5, with anything around 3.5–4 traditionally considered acceptable for business calls) derived from packet loss, jitter and latency. Use these numbers for what they are: a clean read on whether the pipe degraded the call. When Level 1 is bad, no audio software fixes it — that's a network ticket.
What Level 1 cannot see is everything acoustic. The codec faithfully transmits the floor chatter, the traffic, the kitchen; the network metrics stay green because the network did its job perfectly. If your dashboards say "quality: fine" while clients say "your agents sound like a food court," both are telling the truth about different layers. That gap is exactly where noise-cancellation software lives.
Level 2: what a human hears (QA defect codes)
The fix for "audio quality is subjective" is a shared vocabulary. Here is the defect-code scheme we'd add to any QA scorecard — eight codes, one per distinct failure mode, designed so two evaluators tag the same call the same way:
| Code | Defect | Typical root cause |
|---|---|---|
| A1 | Sustained ambient noise behind agent | Floor density, WFH environment — see the WFH standard |
| A2 | Transient noise events (bark, doorbell, slam) | Unpredictable household/floor events; mute reflex + software layer |
| A3 | Third-party speech audible to caller | Open-floor cross-talk; the case for voice isolation |
| A4 | Agent voice low, distant or muffled | Mic distance drift, wrong input device selected |
| A5 | Distortion or "underwater" artifacts | Double processing — two suppression layers fighting |
| A6 | Echo or feedback | Speaker bleed, device conflicts |
| A7 | Agent asked caller to repeat 2+ times | Inbound noise burden — the two-way problem |
| A8 | Neighbouring agent's call audible in recording | Seating density; zoning fixes from the open-floor guide |
Three usage rules. Tag codes on every scored call, not just failed ones — prevalence data is the point. Never attach discipline to A1/A2 in an agent's first offense window; defect codes diagnose systems, and most audio defects have system causes. And run one calibration session before you start, so "sustained" and "transient" mean the same thing to every evaluator.
Level 3: what it cost (outcome metrics)
Level 3 converts audio quality into operational currency:
- Agent repeat-requests per call (A7 count) — inbound noise cost, in seconds.
- Caller repeat-requests per call — outbound clarity cost. QA can tag both from the same recording.
- Data-capture corrections — misheard digits and spellings caught later; each one is rework or a callback.
- Audio-attributed transfers and callbacks — the expensive tail. Add "audio" as a reason code in your transfer disposition list; you can't count what has no code.
We walk the arithmetic from these counts to a yearly cost figure in background noise and handle time — that page is the business-case companion to this one.
Running an honest before/after
The standard design: two weeks of baseline with your current setup, two weeks with the change (built-in suppression enabled, or a dedicated layer piloted — sequencing in the ops guide), same queues, same evaluators, codes tagged throughout. Honesty rules we'd hold ourselves to: don't claim victory on a few dozen scored calls — call small samples "directional" and keep counting; don't change headsets, seating and software in the same fortnight if you want to know which one worked; and pre-register the decision rule ("we buy if A1+A3 prevalence halves") so the renewal debate is settled by numbers you chose before you saw them.
- Two weeks · baselinecurrent setup, codes tagged throughout
- Two weeks · with the changesame queues, same evaluators
- Decideby the rule you wrote down before the data came in
Packaging it for a client QBR
If you're a BPO, Level 2 and 3 data is client-facing gold: a defect-prevalence chart trending down, three before/after audio clips (with consent and PII handled), and the repeat-request delta on the client's own program. That pack answers the audio clauses now appearing in contracts — see BPO client audio requirements for what buyers are writing in, and bring receipts instead of adjectives.
Measurement questions
Our platform shows MOS 4+ but clients complain. Who's right?
Both. MOS-style scores rate the network path; your clients are rating the acoustics. Add Level 2 defect codes and the complaint will localize — usually to A1/A3 on specific rows or specific remote agents.
Can speech analytics auto-detect these defects?
Increasingly, partially — noise levels and cross-talk are becoming machine-taggable, and vendor analytics suites (Krisp lists "Speech Analytics" in its call-center line, for instance) overlap with this scheme. Automate what you can verify; keep human calibration for what you can't.
How many calls do we need to score?
Enough that the numbers stop wobbling week to week — for most mid-size queues that's a few hundred tagged calls per condition, not thirty. Until then, report ranges and say "directional." Overselling a small sample is how measurement programs lose the floor's trust.