An App Store download took 10+ minutes on home Wi-Fi. The same phone on cellular, at the same moment, took about a minute.
Two AdGuard Home defaults were working together: ratelimit: 20 and ratelimit_subnet_len_ipv4: 24. The limit applies per subnet, and a /24 is the whole house. So every device shared one 20-queries-per-second budget on each resolver. Anything over that got dropped. Dropped queries never get written to querylog.json. The logs looked fine the entire time.
Why it looks like a bandwidth problem
A modern app update resolves dozens of hostnames in a burst. CDN shards, OCSP endpoints, telemetry, fallbacks. A per-subnet rate limit punishes that shape.
Downloads don’t fail. It stalls at each resolution, retries after a timeout, and finishes eventually. It looks like a slow link.
Cellular was fast because it skipped my resolvers entirely. That comparison pointed at DNS. It also killed the WAN theory. Same phone, same second, same CDN.
What I ruled out first, and shouldn’t have
Six theories died before I looked at the resolver config.
| Hypothesis | Why it died |
|---|---|
| WAN throughput / CDN edge selection | Same speed to the same edge from a wired host |
| Wrong DNS answers | Answers were correct and correctly blocked |
| QUIC / HTTP-3 blocked on egress | No egress filtering in play |
| Wi-Fi QoS shaping | No rate rules matched the client |
| Local content caching interfering | Not enabled |
| RF congestion | 5 GHz channel utilisation measured 2–4% |
Each one was cheap on its own and expensive as a sequence. The test I reached last should have been first. If a service has a rate limit, measure the limit first. It’s a documented, default-on feature with a number sitting in a config file. One burst test settled it.
The measurement
I fired 60 concurrent queries at each resolver. Then counted the answers that came back.
| Resolver | Before | After |
|---|---|---|
| Resolver A | 29 of 60 dropped (48%) | 0 of 120 dropped |
| Resolver B | 23 of 60 dropped (38%) | 0 of 120 dropped |
| Router (no rate limiting) | 0 of 60 | 0 of 60 |
That router row is the control. It’s what proves the problem is the resolver and not the network path. Half a burst disappearing isn’t subtle. It was invisible because of where it got recorded, which was nowhere.
The change
On 2026-08-11, on both resolvers:
ratelimit: 200
ratelimit_subnet_len_ipv4: 32
ratelimit_subnet_len_ipv6: 128
That 200 raises the ceiling. The subnet-length keys do the work. At /32, and /128 for IPv6, the budget is per client. That’s what I assumed “rate limit” meant. I read the default and didn’t think about it.
Now a phone doing something pathological throttles itself instead of the household. I ran a burst from a second client during a first client’s burst. 0 drops.
Blocking still works. The ad and telemetry domains still return 0.0.0.0. The change cost me protection I thought I had. It was protecting against nobody.
The part worth generalising
A metric that only counts successes can’t show you a drop.
querylog.json records queries that were answered. The rate limiter throws requests away before they get there. The log was complete, and that’s what made it useless.
When a component can silently refuse work, you need the refusal count. If nobody exports that number, you can’t observe the component. You have to provoke it.