A name lookup fails and the application reports the failure after roughly ten seconds. The wait is the same whether two DNS servers are configured or six. Adding another server to the list does not shorten the wait, and the fourth one appears never to be used at all.
None of that is a fault. Both numbers — the ten seconds, and the point at which later servers first get a question — are published defaults, and they are documented precisely enough to be read off a timeline. What follows is where the time goes, why the length of the server list stops mattering after the third entry, and the one place where the client and the server disagree about whether the lookup is over.
What to Change First
Before any of the mechanism below, three changes account for most real improvement.
- Put the server that answers fastest first. The first entry is queried alone at
0 s. Every other entry costs at least one full second of waiting before it is consulted. - Stop at three entries on a client. A fourth and later entry is not contacted until the four-second mark. Everything after the third position behaves as one undifferentiated group.
- Check the application's own timeout. If the application gives up at three seconds, the fourth DNS server has not been asked anything yet. Extending the list cannot help a client that has already stopped waiting.
The rest of this explains why those three, and not the usual suggestions.
Where the Ten Seconds Go
Microsoft documents the client-side schedule step by step. With three or more servers configured, the sequence is fixed:
0 s— the first server is queried1 s— no answer, so the second server is queried2 s— still nothing, so the third server is queried4 s— all configured servers are queried simultaneously8 s— all configured servers are queried simultaneously again10 s— the client stops
Five attempts inside ten seconds. The gaps between them are one, one, two and four seconds — non-decreasing, but not by a constant factor — and the last two seconds are not an attempt at all. They are the wait for a reply to the query sent at eight.
The single-server case runs on the same clock. One server is retried at 1, 2, 4 and 8 seconds and abandoned at 10. The number of entries changes who is asked, not how long the client is willing to wait.
One shortcut ends the sequence early. A negative answer — the name definitively does not exist — stops everything immediately. That is why a misspelled hostname fails instantly while a hostname pointing at an unreachable server takes the full ten seconds. The two failures feel unrelated and produce very different waits from the same resolver.
Two Servers Produce a Different Shape
The count changes who is asked and when, inside the same ten seconds. With exactly two servers configured, the documented sequence is:
0 s— the first server1 s— the second server2 s— the second server again4 s— both servers at once8 s— both servers at once10 s— stop
The two-second slot goes to the second server rather than to a third. From the four-second mark onward the two-server and three-server cases are identical: everything configured is queried together, twice, and then the client gives up.
So the list length controls only the first two seconds of behaviour. That is the entire window in which the ordering of entries has any individual effect, and it is why the position of a server matters far more than the number of servers.
Why the Fourth Entry Rarely Matters
Positions one, two and three each get a query of their own, at 0, 1 and 2 seconds. Position four does not. In Microsoft's wording, servers beyond the third are subject to a minimum four-second delay, because the first moment every configured server is queried at once is the four-second mark.
That produces a threshold worth stating plainly: an application whose own timeout is shorter than four seconds will never use the fourth server, or the fifth, or the twentieth. The entries are present in the configuration, they are syntactically valid, and they are never asked anything before the application has already returned an error.
This is the specific reason that lengthening the list is not a redundancy strategy. Four entries and ten entries have identical behaviour for the first four seconds, and identical behaviour after it as well — at 4 s and 8 s all of them are queried together, so positions four through ten are not ranked against each other in any way. Redundancy that arrives later than the caller is willing to wait is not redundancy.
The Cost of a Silent First Entry
There are two ways for a DNS server to be unhelpful, and they cost very different amounts.
A server that refuses or answers negatively is cheap. The response arrives, and if it is a definitive negative the whole sequence ends at once. A server that is silent — powered off, filtered by a firewall, listening on the wrong interface — is expensive, because silence can only be detected by waiting.
The waiting is a fixed quantity, and it can be read straight off the schedule. If the first entry is silent, the second is not queried until 1 s, so every uncached lookup costs at least one extra second. If the first two are silent, the third waits until 2 s, and the floor is two extra seconds per uncached lookup.
That is per lookup, not per session. Where a page or an application resolves several distinct hostnames in sequence, the added delay multiplies by the number of lookups that miss the cache. Ten sequential lookups behind one silent first entry carry at least ten seconds of added waiting, none of which appears as an error anywhere.
This produces the most misleading symptom in the set: everything works, and everything is slow, and no component reports a fault. Each lookup eventually succeeds, so there is nothing in a log to find. Only the position of a dead entry in a list explains it.
A slow server is a third case, and the schedule treats it differently again. Because every configured server is queried again at 4 s and at 8 s, a server that is late rather than absent can still answer inside the window and the lookup succeeds. A lookup that eventually works is not the same problem as one that fails at ten seconds, and the two need opposite responses.
On a Windows DNS Server, the Ceiling Is Three
A Windows DNS server forwarding queries elsewhere runs a different clock, with three defaults that interact.
- RecursionTimeout — 8 seconds. How long the server waits for remote servers while working on a recursive query.
- ForwardingTimeout — 3 seconds. How long it waits at each ordinary forwarder.
- ForwarderTimeout — 5 seconds. The same idea for conditional forwarders, set per zone.
Microsoft's documented schedule for ordinary forwarders is 0 s for the first, 3.5 s for the second — three seconds plus about half a second of overhead — and 7.5 s for the third. A fourth would be queried at 11.5 seconds.
It never is. RecursionTimeout expires at 8 seconds, which is 3.5 seconds earlier. The documentation states the consequence directly: at default settings only three forwarders can be queried, and the server sends SERVFAIL at 11.5 seconds instead of trying a fourth.
The margin on the third forwarder is thinner than it looks. It is queried at 7.5 seconds against a timeout that expires at 8.0 — half a second. A forwarder that is reachable but slightly slow is, in practice, being asked a question it has no time to answer.
Conditional forwarders are worse, because the interval is five seconds rather than three. The first is queried at 0 s and the second at 5.5 s, and a third falls outside the same 8-second expiry. Two is the ceiling, not three. Microsoft is explicit that going beyond these counts requires changing the timeout values, not just adding entries.
The Second and a Half Where the Two Sides Disagree
Put the two published schedules side by side and a figure appears that neither document states, because neither is about the other.
The client stops at 10.0 seconds. A Windows DNS server that has exhausted its forwarders returns SERVFAIL at 11.5 seconds. The difference is 1.5 seconds, and it belongs to nobody.
The subtraction assumes one thing worth stating: that both clocks start at the same moment, which is the case for the first query of a lookup. The client's retries at 1, 2, 4 and 8 seconds are separate messages, and whether any of them restarts the server's own recursion is not something either document addresses. The 1.5 seconds describes the first query, not necessarily the whole exchange.
During that window the client has already reported failure to the application, and the server is still processing the same question. When the answer finally arrives, there is no longer anything waiting for it. Both components are behaving exactly as documented.
Two things follow for anyone reading logs. The server's log will show a query it worked on and answered; the client's log will show a lookup that timed out. They can both be describing the same lookup. A gap between the two records is not evidence of a dropped packet.
And the observable symptom — ten seconds, then failure — is a client-side ceiling, not a measurement of how long the name actually took. Timing the failure measures the client's patience. It says nothing about which forwarder was slow, or whether any of them ever replied.
What Not to Change
Three adjustments come up constantly and do not address this.
- Adding more DNS servers to the client. Covered above: positions beyond the third share one four-second starting line.
- Flushing the resolver cache. The cache is not involved in the schedule above. A lookup that reaches the network at all has already missed the cache, and clearing it changes nothing about the ten seconds.
- Adding a fourth forwarder on the server. It will sit in the configuration and never be queried until RecursionTimeout is raised or ForwardingTimeout is lowered.
The adjustments that do change the outcome are the ones that alter a documented number: reordering so the reliable server is first, shortening the list, changing the timeout values, or extending the application's own patience past four seconds so later servers become reachable at all.
When This Doesn't Apply
The figures above are Windows defaults for a specific path. Several common situations sit outside them.
- Non-Windows clients. Linux resolvers, macOS, Android and iOS each use their own retry schedules. The 1-2-4-8 pattern and the ten-second ceiling are not portable.
- Applications with their own resolvers. Browsers using DNS over HTTPS, and applications that resolve names themselves, do not necessarily follow the system client's schedule at all.
- Any environment where the defaults were changed. All of these values are configurable. Where a group policy or a build script has adjusted them, the published defaults describe nothing.
- A name that definitively does not exist. A negative answer ends the sequence immediately. Nothing here explains a lookup that fails in well under a second.
- Failures that are not name resolution. A lookup that succeeds and is followed by a connection failure is a different problem with a different clock. Confirming which one is happening comes before any of this.
Comments
Post a Comment