Netfilter’s connection tracking table. Every flow the kernel NATs or filters statefully gets an entry keyed by the 5-tuple (protocol, source IP/port, destination IP/port). NAT rules — including everything kube-proxy writes in iptables mode — are applied once, when the entry is created; subsequent packets of the flow are rewritten by looking up the existing entry rather than re-evaluating rules.
Three consequences that keep showing up in incidents:
- It is finite.
nf_conntrack_maxcaps it. Full table →nf_conntrack: table full, dropping packetand connections fail with no application-level error. - UDP entries are guesses. UDP has no handshake, so the kernel invents a flow from the first packet and expires it on a timer (
nf_conntrack_udp_timeout, 30s by default). DNS lives here. - Insertion is racy. Two packets of the same “flow” NAT’d at the same instant can both build an entry; one loses the insert and its packet is dropped silently. Visible only as
conntrack -S | grep insert_failedclimbing.
The third one is why DNS lookups inside pods stall for exactly five seconds — a dropped packet is not an error, it is a timeout.
conntrack -S # per-CPU counters: insert_failed, drop, early_drop
sysctl net.netfilter.nf_conntrack_count net.netfilter.nf_conntrack_max