TL;DR: Your call quality dashboard and your bad call can both be telling the truth, because every number on that dashboard is an average and voice fails on the excursion. ITU-T Y.1541 scores networks per one minute interval and says any minute observed should meet the objective, and RFC 3550 defines the jitter number your tools report as a smoothed mean deviation with a one sixteenth gain, so a two second burst is flattened before you ever see it. The published bounds are real and worth memorising: Y.1541 class 0 allows 100 ms mean IP packet transfer delay, 50 ms of delay variation and a packet loss ratio of 1 in 1000, class 1 relaxes delay to 400 ms, and class 5, which is what the open internet sells you, is literally “U” for unspecified on every parameter. ITU-T G.114 puts 150 ms mouth to ear at essentially transparent interactivity and 400 ms one way as the planning ceiling. The trap is that those two numbers measure different paths: 400 ms is network only, 150 ms includes your codec, your packet loss concealment and your jitter buffer. Y.1541 does the arithmetic in public, and it is tighter than anyone expects. A 100 ms class 0 network plus a 50 ms endpoint lands on exactly 150 ms, while the same network with a 60 ms jitter buffer instead of a 50 ms one lands at 180 ms. That is the whole job. Enlarging the jitter buffer to hide jitter spends delay budget you do not have, so you are not fixing both, you are choosing which one the rep hears. Scope the incident across user, endpoint, site, network, direction, destination and time window before you touch a setting, test one variable at a time with a rollback, and remember that QoS only governs queues you own. RFC 4594 puts voice in the Telephony class on Expedited Forwarding, but it also requires call admission control and says markings from untrusted end user devices get re-checked at the network edge, which is why the DSCP tag your softphone sets on a home network changes nothing. Then hand the ticket to whichever of the four owners, endpoint, local network, internet provider or voice provider, owns the excursion you actually captured.
A rep tells you the last three calls were choppy. You open the call quality dashboard, and it is green. Average jitter 9 ms. Average latency inside the budget. Packet loss a rounding error. So you tell the rep to restart the app, you close the ticket, and three days later the same rep is in your office with the same complaint and a deal that stalled on a discovery call nobody could hear.
Both of those readings can be accurate at the same time. That is the part most troubleshooting guides skip, and it is a measurement problem rather than a disagreement.
Voice is unusual among the things that run on your network. It is constant rate, it is inelastic, and it does not retry. RFC 4594, the IETF’s guidelines for DiffServ service classes, puts it plainly: the payloads in the telephony class “do not react to loss or significant delay in any substantive way.” A file download notices congestion and backs off. A call does not notice anything. It keeps sending 50 packets a second into a path that cannot carry them, and the rep hears every one that does not arrive on time.
So what is the dashboard actually telling you? The average. Which is the wrong statistic here. A call does not fail across a fifteen minute window. It fails for two seconds, in the middle of a sentence, while the buyer is explaining their renewal date. The useful question is never “what is our average jitter.” It is “what was the worst two seconds, and who owns the path it happened on.”
So troubleshoot VoIP call quality problems the way the standards bodies actually wrote them down. The thresholds are published, they are in force, and they are specific enough to tell you which of four owners takes the ticket. What follows is that method, the arithmetic underneath it, and the reason your dashboard keeps lying to you with a straight face.
Why the VoIP call quality dashboard says fine while the call was bad
Two documents explain the gap, and once you have read them the green dashboard stops being mysterious.
Start with the jitter number itself. Almost every softphone, SBC and monitoring tool reports jitter from RTCP receiver reports. RFC 3550 does not leave that calculation to the vendor. It specifies it, and it uses MUST: “The jitter calculation MUST conform to the formula specified here in order to allow profile-independent monitors to make valid interpretations of reports coming from different implementations.” The formula is J(i) = J(i-1) + (|D(i-1,i)| - J(i-1))/16, where D is the difference in relative transit time between two consecutive packets.
What does that one sixteenth actually do? It makes the number stable and slow. Jitter here is not the maximum. It is not the 95th percentile. It is a running smoothed mean deviation, and each new sample moves the estimate by a sixteenth of the gap. A sudden burst of real delay variation has to persist for many packets to pull the reported number anywhere meaningful, and a two second disaster inside a twelve minute call is mathematically guaranteed to disappear into a low average. RFC 3550 is explicit that this is the design intent, calling one sixteenth the gain parameter that “gives a good noise reduction ratio.”
Noise reduction is exactly what you do not want when you are hunting a two second event. The statistic was built to be stable for congestion control. You are using it as a fault detector. It is not one.
Now the second document. ITU-T Y.1541 defines network performance objectives for IP based services, and it sets its bounds with an explicit measurement window: “An evaluation interval of 1 minute is suggested for IPTD, IPDV, and IPLR and, in all cases, the interval must be recorded with the observed value. Any minute observed should meet these objectives.”
Any minute observed. Not the daily mean, and not the monthly SLA report. If one minute in your business day blows the objective, that network did not meet the objective, and the rep who was on a call during that minute is the one who finds out. When a provider shows you a monthly availability figure and calls it call quality evidence, they have answered a different question than the one you asked.
So before anything else, change what you collect. You are looking for excursions in a window, not averages across a day. Get per call records with timestamps, get the worst case rather than the mean, and keep the clock aligned between your phone system and your network tooling so an audio complaint at 10:42 can be laid against a network event at 10:42.
The three VoIP call quality numbers and where they are actually bounded
Here is the part the vendor guides skip. There are published bounds. They are not folklore, they are in force, and they come with the measurement conditions attached, which is what makes them usable.
ITU-T Y.1541 Table 1 defines QoS classes for IP networks. Class 0 names your use case directly in Table 2: “Real-time, jitter sensitive, high interaction (VoIP, VTC).” Class 1 is the same application with “less constrained routing and distances.” The objectives:
- Class 0: upper bound on mean IP packet transfer delay 100 ms, delay variation 50 ms, packet loss ratio 1 x 10-3.
- Class 1: delay 400 ms, delay variation 50 ms, packet loss ratio 1 x 10-3.
- Class 5: “U” on every single parameter, where Y.1541 defines U as unspecified and warns that “performance with respect to that parameter may, at times, be arbitrarily poor.”
Class 5 is the one to sit with. Y.1541’s own note on Table 2 says any application listed there “could also be used in class 5 with unspecified performance objectives, as long as the users are willing to accept the level of performance prevalent during their session.”
That is the ITU describing the ordinary public internet. It is the honest version of what a rep on home broadband runs calls across. Nobody committed to anything. A commitment is the thing you are actually shopping for when you buy a managed path. If you have not bought one, the correct expectation is not a number. It is a shrug.
One way delay and the 150 ms that is not what you think
ITU-T G.114 is the delay recommendation, and it carries two numbers that are constantly collapsed into one. They measure different things.
The first: “Regardless of the type of application, it is recommended to not exceed a one-way delay of 400 ms for general network planning (i.e., UNI to UNI).” UNI to UNI means user network interface to user network interface. Network only. Your handset is not in that measurement.
The second: “Although a few applications may be slightly affected by end-to-end (i.e., ‘mouth-to-ear’ in the case of speech) delays of less than 150 ms, if delays can be kept below this figure, most applications, both speech and non-speech, will experience essentially transparent interactivity.”
Mouth to ear. That one includes everything. The microphone, the codec’s packet formation time, the network, the de-jitter buffer at the far end, packet loss concealment, and the speaker. So when someone runs a ping, sees 40 ms, and concludes they are comfortably inside the 150 ms budget, they have compared a round trip network measurement against a one way end to end budget and gotten the answer wrong twice.
And notice what G.114 does not say. It does not say 151 ms is a failure. It says below 150 ms interactivity is essentially transparent, and it points you at the E-model in ITU-T G.107 to estimate the effect of delay alongside other impairments. Delay degrades conversation gradually, in talkover and collision and that half beat of hesitation that makes a buyer think the rep is unsure. There is no cliff. There is a budget, and it gets spent.
Jitter and the 50 ms bound that is a quantile, not an average
Y.1541’s delay variation bound is 50 ms, and the definition matters more than the number. It is an “upper bound on the 1 – 10-3 quantile of IPTD minus the minimum IPTD.”
Translate that. Take every packet’s transfer delay over the interval. Find the value that 99.9 percent of packets come in under. Subtract the fastest packet. That spread must be under 50 ms. This is not a mean. This is explicitly a tail measurement, which is exactly the statistic RFC 3550’s smoothed jitter estimate is not.
Your dashboard reports the smoothed mean deviation. The standard bounds the 99.9th percentile spread. Those are different numbers describing the same network, and the first one can sit at 9 ms while the second one is blowing 50 ms several times an hour. If you only have the first, you do not have the measurement the objective is written against, and you should say so when you escalate rather than pretending the green number is evidence of anything.
Packet loss and the 0.1 percent line
The Y.1541 bound is a packet loss ratio of 1 x 10-3. One packet in a thousand. The recommendation says where that came from: “The class 0 and 1 objectives for IPLR are partly based on studies showing that high quality voice applications and voice codecs will be essentially unaffected by a 10-3 IPLR.”
Essentially unaffected at 0.1 percent. That is a much tighter line than the one percent figure that circulates in vendor blogs, and the difference matters because of how loss arrives. Codecs conceal loss by synthesising a replacement for the missing frame, and concealment works well on isolated losses and badly on consecutive ones. Half a percent spread evenly across a call is often inaudible. The same half a percent arriving as four back to back drops is a word the buyer did not hear.
So when you capture loss, capture whether it was scattered or bursty. “We saw 0.4 percent loss” is not actionable. “We saw 0.4 percent, and it arrived as three bursts of six to nine consecutive packets” tells a network team where to look.
The VoIP call quality budget is spent before anything breaks
This is the section that changes how you argue with providers, and it is just arithmetic that ITU-T published in an appendix.

Y.1541 Appendix VII works a reference VoIP endpoint all the way through. G.711 coder, 20 ms RTP payload, 60 ms jitter buffer at the receiver, packet loss concealment. It then totals the endpoint delay: 40 ms of packet formation, 30 ms in the jitter buffer, 10 ms of packet loss concealment. Eighty milliseconds, before a single packet enters the network.
Now add a class 0 network at its 100 ms bound. The recommendation states the result: “the total average delay for the user-to-user path is 100 + 80 = 180 ms.”
Put that against G.114. The transparent interactivity figure is 150 ms. A perfectly conformant class 0 network, the best class the ITU defines, carrying a textbook endpoint, lands at 180 ms. Over budget, with nothing broken anywhere.
Appendix VII then shows the only way back under. Shorten packet formation to 20 ms and shrink the de-jitter buffer to 50 ms, and endpoint delay drops to 50 ms. “The class 0 path IPTD and customer installation delays sum to a 1-way mouth-to-ear transmission time of 150 ms, satisfying the needs of most applications.”
Exactly 150. Not comfortably inside it. On it.
So where is the headroom? There is none. That is the finding.
Sit with what that means for your actual job. The standard fix for jitter is a deeper jitter buffer, because more depth absorbs more delay variation before a packet arrives too late to play. It works. It is also the most expensive thing you can do to your delay budget. The two reference endpoints above differ by exactly that choice. A 60 ms buffer costs 30 ms of mouth to ear delay. A 50 ms buffer costs 25 ms. That decision, plus packet formation time, is most of the gap between landing at 180 and landing at 150.
So you are not fixing jitter and delay. You are trading one against the other, and the trade is already close to maxed before you arrive. When a rep reports choppy audio and someone deepens the buffer, the choppiness improves and the conversation starts feeling slightly off, with more talkover and more of those half beat collisions. Nobody connects the two tickets, because they were filed three weeks apart by different people.
One more detail from the same appendix, because it is the most commonly miscalculated number in this whole subject. “It must be noted that a de-jitter buffer’s contribution to mouth-ear delay is based on the average time packets spend in the buffer, not the peak buffer size.” A 60 ms buffer does not add 60 ms. It adds about 30, the centre of the buffer. If you have been adding full buffer depth to your delay budget, your budget has been wrong in the pessimistic direction, and if you have been adding nothing, it has been wrong in the direction that gets a deal lost.
Practically, this is why capacity planning and call quality are the same conversation. Concurrency, codec choice and buffer depth all draw on one budget, which is also the argument for sizing your calling capacity deliberately rather than discovering the ceiling during a busy Tuesday.
Match the VoIP call quality symptom to the number it violates
“Bad quality” is not a symptom, it is a category. Before you touch anything, get the rep to tell you what each side heard and in which direction. The symptom maps to a different number, and the number maps to a different owner.
Choppy or robotic audio
Broken words, metallic artefacts, brief dropouts. This is loss or late arrival, which at the receiver amount to the same thing, because a packet that misses its play-out slot is discarded exactly like one that never came. Ask whether it was continuous or arrived in bursts. Bursts during a backup window or a busy call block point at contention. Continuous degradation on one rep’s calls and nobody else’s points at that rep’s endpoint or last hop.
Delay, talkover and collisions
The conversation keeps stepping on itself and both parties start over-apologising. This is the G.114 budget, and it is the symptom most likely to be misread as the buyer being distracted or the rep being nervous. Test across several destinations. Delay that follows one destination is a routing or far end issue; delay on every call from one site is yours.
Echo
Who hears the echo? Ask that before anything else, because echo is usually not the network. It is acoustic coupling most of the time, a speaker feeding a microphone, or a headset issue, or a hybrid somewhere in the path. Establish who hears it, because that single question splits the diagnosis. If the far party hears their own voice, the reflection is happening at your end. If your rep hears an echo, it is happening at theirs. Lower the volume, move from speakerphone to a headset, and swap the headset before anyone opens a router.
One way audio
The call connects, one side hears nothing. This is almost always signalling, NAT or firewall behaviour rather than quality, which is why it belongs in a different bucket than the three above. Document the failing direction precisely. Do not start opening ports or changing NAT traversal settings on a hunch, because the blast radius of a firewall change is much larger than the ticket you are working.
Dropped calls
Capture duration. Calls that end at a repeatable interval are a session timer or a mid-call signalling problem, not an audio quality problem, and they get solved somewhere completely different. Note whether audio disappeared before the disconnect or the call simply ended, and keep failed call setup in its own category.
Degradation only in a busy window
Quality that falls apart between 9 and 10 and is fine at 3 is the most tractable version of this problem, because the pattern itself is the evidence. Compare the window against concurrent call count, backup jobs, software update rollouts and site headcount. A speed test run at 3pm proves nothing about 9am, and running one is the most common way this investigation gets closed too early.
Worth separating out explicitly: none of this is about the quality of the conversation. If the audio is clean and the calls still are not landing, that is a coaching question and it has its own instrument in a call quality scorecard. Audio fidelity and conversation quality share a word and nothing else, and conflating them sends the ticket to the wrong team for a week.
Scope a VoIP call quality problem before touching a setting
Scope is the cheapest diagnostic you own. Every dimension you pin down deletes a whole class of causes, and it costs nothing but discipline.
- Users. One rep, one team, or everyone?
- Endpoints. One headset, one browser, one softphone build, one desk phone model?
- Locations. One home network, one office, one branch, or all of them?
- Networks. Wi-Fi, wired, hotspot, VPN?
- Direction. Inbound, outbound, or both?
- Destinations. One number, one region, one carrier, or everything?
- Timing. Constant, random, or a window you can name?
The combinations resolve almost immediately. One user everywhere they go is an endpoint or an account. Everyone at one site is that site’s network or circuit. Every site at the same moment is a shared service, and that is the one where you stop testing and start collecting evidence for the provider. One destination from everywhere is a route, and no amount of QoS on your side will touch it.
Then test one variable at a time. Wired against wireless, swap the headset, same rep on a different approved network, same network with a different rep. One change per test, written down, with the result. The instinct under pressure is to change four things at once because the rep has a demo in twenty minutes, and the cost is that you fix it without learning which change fixed it, which guarantees you are back here next month.
Wi-Fi is where most remote VoIP call quality problems live
If your reps are distributed, start here, because the probability mass is here and because it is the segment nobody owns.
Wi-Fi adds variables that wired links do not have: signal strength, co-channel interference, roaming between access points mid call, airtime contention with every other device in the house, and retransmissions that show up as delay variation rather than loss. Put a test device next to the access point and repeat the call. Then repeat it on Ethernet. If wired is consistently better across several calls, you have localised the problem to the wireless segment and you are done arguing with the voice provider.
Check what else is on the link. Video meetings, cloud sync, operating system updates, a game console downloading a patch. Upload saturation matters more than people expect, because the advertised number everyone quotes is the download figure, and a call needs its upstream path every bit as much as its downstream one. A household uploading a backup can degrade calls while a speed test still reports a healthy connection.
Two honest caveats. A mobile hotspot comparison is useful but not conclusive, because it changes the carrier and the entire path at once, so it tells you “not this link” rather than “this link.” And home networks are not yours. You can recommend, you cannot enforce, and the practical answer for a distributed team is usually wired by default for anyone whose job is calls, plus the evidence to show when it is not the network at all. That evidence discipline is the same one that makes tracking remote rep calls defensible, and it is worth building once.
What QoS does for VoIP call quality and exactly where it stops
QoS is the most over-promised item in this entire subject, so it is worth being precise about what it is and what it cannot reach.
RFC 4594 defines a Telephony service class and recommends it for traffic that “require[s] real-time, very low delay, very low jitter, and very low packet loss for relatively constant-rate traffic sources.” The marking is Expedited Forwarding, and the recommended handling is a priority queue. That part is well specified and it works.
But read the conditions the same document attaches, because they are where most deployments quietly fall apart.
First, admission control is part of the design, not an optional extra. RFC 4594 requires that “the bandwidth in the core network and the number of simultaneous VoIP sessions that can be supported needs to be engineered and controlled so that there is no congestion for this service.” It adds that “the call admission procedure should have verified that the newly admitted flow will be within the capacity of the Telephony service class forwarding capability.”
So a priority queue with no cap on entry does not protect calls. It moves the congestion inside the priority queue. Now every packet in the fight is somebody’s conversation.
Second, your markings are not trusted. “Packet flow marking (DSCP setting) from untrusted sources (end user devices) SHOULD be verified at ingress to DiffServ network.” That sentence is the reason a softphone setting its own EF marking on a home connection accomplishes nothing at all. The marking survives exactly as far as the first device that has an opinion about it.
Third, it stops at your boundary. RFC 4594 contemplates preservation across providers only “at peering points (between two DiffServ networks) where SLAs are in place.” No SLA, no honoured marking. And then Y.1541 closes the loop from the other side by defining class 5 as unspecified on every parameter, with performance that “may, at times, be arbitrarily poor.”
Put those together and the scope of QoS is clear. It governs the queues you own, at the choke points you control, which in practice means your office LAN and your own uplink. It does nothing about the twelve hops after that. This is not an argument against configuring it. It is an argument against expecting it to fix a problem that is happening somewhere you do not administer, and against the forum advice that tells you to disable SIP ALG or open a firewall range because it worked for a stranger with a different topology.
So if you do change configuration, change it like an engineer. Document the current state, read the current vendor guidance rather than a four year old thread, change one variable, define the rollback before you make the change, test inbound and outbound after each one, and confirm your security posture survived. Router, firewall, VLAN, QoS policy, SIP ALG, NAT and session timer changes belong with qualified network staff, because the failure mode is not a worse call, it is an outage or an opening.
Check the VoIP phone and endpoint before blaming the network
A meaningful share of these tickets never touch the network at all, and endpoint tests are the fastest ones you have.
Substitute, do not theorise. Known good headset on the same machine, on the same network. Then the original headset on a different machine. Then the original machine on a different approved connection. Each swap isolates one layer, and the sequence matters more than any individual test.
Check the obvious things that break silently. Confirm the correct microphone and speaker are actually selected, which goes wrong constantly after an operating system or browser update picks a new default device. Watch CPU and memory during a call, because a machine pinned at full load drops audio frames in a way that looks exactly like network loss and will send you chasing packets for a day. Verify the softphone, browser, firmware and desk phone versions are the ones your organisation has approved, and back up settings before updating anything.
One tell worth knowing: endpoint problems follow the person across networks, and network problems follow the location across people. If a rep’s audio is bad at the office and bad at home and bad on a hotspot, stop testing networks.
Confirm the VoIP call quality fix against the original conditions
Did one clean test call prove anything? Almost nothing. Treating it as proof is how the same ticket gets reopened three times, with three different owners.
Repeat the original scenario. The same direction, the same destination, the same network, the same time window, and more than one call. If the problem only ever appeared between 9 and 10, a clean call at 2pm is not evidence. If it only affected outbound calls to one region, inbound tests are not evidence either.
Then compare against what you captured at the start, which is the entire reason for capturing it. If wired testing is what improved things, confirm the improvement holds across several calls before you write down that Wi-Fi was the cause, because intermittent problems are extremely good at going quiet for a day on their own. Record the change, the result, and the rollback path, so the next person who sees this symptom starts where you finished instead of starting over.
Hand the VoIP call quality ticket to the owner who can act on it
Four owners are possible. Which one is it? The evidence you just collected decides, and that is the whole point of doing the work in order.

- The endpoint owner, when the symptom follows the person or the device across networks.
- The local network owner, when it follows a site or a link, and when wired consistently beats wireless.
- The internet provider, when it follows your circuit, correlates with a busy window, or shows loss and delay variation beginning at a hop you do not administer.
- The voice provider, when it follows a destination, a region, a carrier path or a direction, independent of your sites and endpoints.
Send the evidence, not the complaint. A report that survives the other team’s first question carries all of this:
- Timestamps with time zones.
- Call direction, with source and destination handled per your data policy.
- The affected user, location, device, application and connection type.
- The precise symptom, and which party heard it.
- Whether anyone else was affected.
- Call identifiers or diagnostic references.
- Latency, delay variation and loss observations, with the measurement window stated.
- Results of the Ethernet, Wi-Fi, device swap and alternate network tests.
- Any recent network, software, firmware or configuration change.
Two things to be careful with. Call recordings can help, but only where collection, access, retention and sharing are lawful and permitted by your policy, and they never substitute for timestamps and call identifiers. And state your measurement window explicitly, because without it your numbers are averages and you have just handed over the same ambiguity you started with.
The managerial version of all this is short. Stop measuring call quality as a monthly average, start capturing per call worst case with aligned clocks, test one variable at a time, and route by evidence. Do that and the argument with the provider stops being a matter of opinion, which is the only state in which it ever actually gets fixed.
Frequently asked questions about VoIP call quality problems
What jitter and latency numbers should I actually aim for?
Use the published ones. ITU-T Y.1541 class 0 is the class whose stated application is real time, jitter sensitive, high interaction VoIP. It bounds mean IP packet transfer delay at 100 ms, delay variation at 50 ms on the 1 – 10-3 quantile minus the minimum, and packet loss ratio at 1 x 10-3. ITU-T G.114 puts 150 ms mouth to ear at essentially transparent interactivity, and 400 ms one way as the network planning ceiling. Keep straight which path each number measures. The 150 ms figure includes your endpoint. The 400 ms figure does not.
Why does my monitoring show low jitter when calls are clearly choppy?
Because the two of you are measuring different things. RFC 3550 defines the reported jitter value as a smoothed mean deviation updated by one sixteenth of each new sample, which is deliberately insensitive to short bursts. Y.1541’s bound is a 99.9th percentile spread evaluated per one minute interval. A smoothed average can sit low while the tail repeatedly blows the objective, and the tail is what the rep hears.
Will a faster internet plan fix VoIP call quality problems?
Usually not, because voice needs very little bandwidth and a great deal of consistency. A call is a constant low rate stream, and RFC 4594 notes these payloads “do not react to loss or significant delay in any substantive way,” so they neither back off nor recover. More capacity helps only when the actual cause was saturation, which is why you check upload utilisation during the affected window rather than running a speed test afterwards.
Does QoS fix call quality over the public internet?
No. It governs the queues you control, and nothing else. RFC 4594 expects DSCP markings from end user devices to be re-checked at the network edge, and contemplates preservation across providers only where a peering SLA covers it. Y.1541 class 5, which describes ordinary internet service, sets no objective on any parameter at all. It warns that performance “may, at times, be arbitrarily poor.” So configure QoS at your own choke points. Buy a committed path when calls genuinely need one.
Will a bigger jitter buffer fix choppy audio?
It will help the choppiness and cost you delay, and the budget is tight enough that the trade is real. Y.1541 Appendix VII works it through: a reference endpoint with a 60 ms jitter buffer carries 80 ms of endpoint delay and lands at 180 ms mouth to ear over a conformant class 0 network, while trimming to a 50 ms buffer and 20 ms packet formation brings the endpoint to 50 ms and the path to exactly 150 ms. The buffer’s cost is the average time packets spend in it, roughly half its depth, not its full size.
How do I tell a network problem from a headset problem?
Follow the symptom. Endpoint faults travel with the person and the device across every network they use. Network faults stay with a location or a link and affect whoever is sitting on it. One rep with bad audio at the office, at home and on a hotspot is not a network problem, and swapping the headset is faster than any capture you could run.
Sources
How this article was built. Every threshold, bound and formula quoted here was read on the review date in the primary standards document that defines it, not in a secondary summary, and the load-bearing language is quoted rather than paraphrased so its exact measurement conditions travel with it. That matters more than usual in this subject, because the same numbers circulate widely with their conditions stripped off: the 150 ms figure is mouth to ear and the 400 ms figure is network only, and the delay variation bound is a quantile rather than an average. No answer rate, connect rate, call quality benchmark or percentage of tickets figure is quoted anywhere in this article, because no current primary source supports the ones commonly repeated and they would not transfer to your topology, carriers, codecs or sites in any case. The diagnostic sequence, the scope dimensions, the four owner routing model and the escalation evidence list are this article’s own operating guidance and are not requirements of any standard named here. ITU-T Recommendations are voluntary international standards rather than regulations, and conformance to a Y.1541 class is something a network provider offers or does not; the public internet is class 5, which sets no objective at all. Standards are revised and this article’s review date is the date its citations were verified in force. This is general engineering and operations guidance for sales teams, not a network design for any particular deployment, and configuration changes to routers, firewalls, VLANs, QoS policies, NAT or session timers should be made by qualified network personnel with a rollback plan. Kixie publishes this article and sells sales engagement software for business calling and texting.
- ITU-T Recommendation G.114, One-way transmission time, the International Telecommunication Union Telecommunication Standardization Sector, the 05/2003 edition approved 6 May 2003 and in force on the review date together with its Amendment 1 (09/2003) and Amendment 2 (11/2009), read in the published Recommendation text, for the clause stating that “Regardless of the type of application, it is recommended to not exceed a one-way delay of 400 ms for general network planning (i.e., UNI to UNI” as illustrated in ITU-T Rec. Y.1541; for the clause stating that “Although a few applications may be slightly affected by end-to-end (i.e., ‘mouth-to-ear’ in the case of speech) delays of less than 150 ms, if delays can be kept below this figure, most applications, both speech and non-speech, will experience essentially transparent interactivity”; for the statement that delays above 400 ms “are unacceptable for general network planning purposes” while recognising exceptional cases; and for the direction that the E-model of ITU-T Rec. G.107 “should be used to estimate the effect of one-way delay (including all delay sources, i.e., ‘mouth-to-ear’) on speech transmission quality for conversational speech.”
- ITU-T Recommendation Y.1541, Network performance objectives for IP-based services, the International Telecommunication Union Telecommunication Standardization Sector, the 12/2011 edition in force on the review date together with its Amendment 1 (12/2013), read in the published Recommendation text, for the Table 1 QoS class objectives giving class 0 an upper bound on the mean IPTD of 100 ms, an IPDV bound of 50 ms defined as the “Upper bound on the 1 – 10-3 quantile of IPTD minus the minimum IPTD,” and an IPLR bound of 1 x 10-3, and giving class 1 the corresponding 400 ms, 50 ms and 1 x 10-3; for the definition of “U” as unspecified, where “ITU-T establishes no objective for this parameter” and “performance with respect to that parameter may, at times, be arbitrarily poor,” which is the value class 5 carries on every parameter; for the general note that “An evaluation interval of 1 minute is suggested for IPTD, IPDV, and IPLR and, in all cases, the interval must be recorded with the observed value. Any minute observed should meet these objectives”; for Note 4 stating that the class 0 and 1 IPLR objectives are “partly based on studies showing that high quality voice applications and voice codecs will be essentially unaffected by a 10-3 IPLR”; for the Table 2 class guidance naming class 0 as “Real-time, jitter sensitive, high interaction (VoIP, VTC)” and class 1 as “Real-time, jitter sensitive, interactive (VoIP, VTC),” with the note that any listed application “could also be used in class 5 with unspecified performance objectives, as long as the users are willing to accept the level of performance prevalent during their session”; and for the Appendix VII endpoint delay analysis giving 40 ms packet formation, 30 ms average jitter buffer and 10 ms packet loss concealment for a total of 80 ms, the statement that “the total average delay for the user-to-user path is 100 + 80 = 180 ms,” the low delay alternative totalling 50 ms whose “class 0 path IPTD and customer installation delays sum to a 1-way mouth-to-ear transmission time of 150 ms, satisfying the needs of most applications,” and the caution that “a de-jitter buffer’s contribution to mouth-ear delay is based on the average time packets spend in the buffer, not the peak buffer size.”
- RFC 3550, RTP: A Transport Protocol for Real-Time Applications, the Internet Engineering Task Force, dated July 2003, Standards Track and the current Internet Standard for RTP on the review date, read in the published RFC text, for the section 6.4.1 definition of interarrival jitter as “the mean deviation (smoothed absolute value) of the difference D in packet spacing at the receiver compared to the sender for a pair of packets”; for the definition D(i,j) = (Rj – Ri) – (Sj – Si); for the specified formula J(i) = J(i-1) + (|D(i-1,i)| – J(i-1))/16 and the requirement that it “SHOULD be calculated continuously as each data packet i is received”; for the requirement that “The jitter calculation MUST conform to the formula specified here in order to allow profile-independent monitors to make valid interpretations of reports coming from different implementations”; for the description of 1/16 as “the gain parameter” that “gives a good noise reduction ratio”; and for the Appendix A.8 reference implementation of the same estimate.
- RFC 4594, Configuration Guidelines for DiffServ Service Classes, the Internet Engineering Task Force, dated August 2006, Informational, read in the published RFC text, for the section 4.1 Telephony service class recommended “for applications that require real-time, very low delay, very low jitter, and very low packet loss for relatively constant-rate traffic sources (inelastic traffic sources)” and which “SHOULD be used for IP telephony service”; for the statement that “the inelastic types of RTP payloads in this class do not react to loss or significant delay in any substantive way”; for the recommendation that the class “SHOULD use Expedited Forwarding (EF) PHB, as defined in [RFC3246]” with a priority queuing system; for the requirement that “the bandwidth in the core network and the number of simultaneous VoIP sessions that can be supported needs to be engineered and controlled so that there is no congestion for this service” and that “the call admission procedure should have verified that the newly admitted flow will be within the capacity of the Telephony service class forwarding capability in the network”; for the network edge conditioning guidance that “Packet flow marking (DSCP setting) from untrusted sources (end user devices) SHOULD be verified at ingress to DiffServ network”; and for the treatment of marked flows “At peering points (between two DiffServ networks) where SLAs are in place.”
Sources verified and content reviewed by the Kixie Research Team on October 6, 2026. All source links checked on October 6, 2026.
Ready to close more deals with Kixie?
See how Kixie's AI-powered tools can transform your sales and support operations.
Start Free Trial