Where Packets Get Dropped: Tracing the Linux Network Stack with eBPF
Connection timeouts and rising TCP retransmission rates tell you something is wrong. They do not tell you where packets are being dropped. Is a firewall rule rejecting traffic, a route lookup failing, or a receive buffer filling up? Tools such as dropwatch and perf record -e skb:kfree_skb can help identify drop locations, but a location alone does not tell you whether a dropped packet belongs to the affected application.
HUATUO’s dropwatch adds packet context to kernel drop tracing. An eBPF probe attached to tracepoint/skb/kfree_skb captures the IP five-tuple, process information, network device, MAC addresses, and kernel call stack. It also supports tcpdump-style filters compiled into eBPF bytecode at load time. Filtering in the kernel ensures that only matching events reach userspace, keeping the output focused on the traffic under investigation.
This walkthrough covers targeted captures, event analysis, and continuous monitoring with huatuo-bamai, starting with a service experiencing intermittent latency spikes.
Investigating a Latency Spike
In a Kubernetes cluster, the P99 latency of an order service periodically spikes from 50 ms to 2 s. Application logs show sporadic connection reset by peer errors, and netstat -s on the node reveals a steady increase in TCPLostRetransmit.
An investigation might involve capturing traffic with tcpdump, inspecting firewall rules with iptables -L, checking routing with ip route get, and tracing kernel drop locations with perf record -e skb:kfree_skb. Connecting those observations to the affected service requires packet context as well as drop locations.
Start with a 60-second dropwatch capture filtered to TCP port 8080 on eth0:
|
|
The output brings process information, source and destination addresses and ports, and the kernel call stack together in each event. Use these fields to identify relevant traffic, then filter with jq to examine the drop paths associated with the incident.
Filtering Drop Events in the Kernel
On busy hosts, kfree_skb events can be frequent. Some reflect routine cleanup rather than packet loss affecting application traffic, as discussed in the noise-filtering section below. Sending every event to userspace increases processing overhead and makes relevant events harder to find.
Dropwatch evaluates traffic filters in the kernel. Its pure-Go pcap compiler (internal/pcapfilter) compiles tcpdump-style expressions into eBPF bytecode when the program loads. That bytecode runs within the probe, which submits only matching events to userspace through the perf ring buffer.
Supported Filter Primitives
internal/pcapfilter implements a practical subset of the standard tcpdump syntax:
Protocol and Direction
|
|
Boolean Composition
|
|
Unsupported expressions include byte offsets (tcp[tcpflags]), numeric protocol numbers (ip proto 6), and port ranges (portrange). Refer to the usage documentation for the complete list of unsupported features.
Filter Examples
|
|
Traffic filters (--filter) and device filters (--device / --device-excluded) are combined with AND semantics: an event must satisfy both. With --device, events for SKBs that have no net_device are excluded from the output. With --device-excluded, those events pass the device filter. These filters control event reporting, not packet forwarding; keep this distinction in mind when tracing container veth traffic.
Reading a Drop Event
With --output json, dropwatch emits one JSON object per event (NDJSON). The following example is formatted for readability:
|
|
Notable fields:
layers: Parsed packet headers, with a nested object for each protocol layer. Missing layers are omitted. Consumers can check which objects are present to identify the protocols, for example withev.Layers.TCP != nil.stack: The kernel call stack at the drop site, with frames separated by newlines. Use it to identify the code path involved:tcp_v4_rcvpoints to receive processing, whileip_outputpoints to the output path.container_id: Resolved by huatuo-bamai using the memory cgroup CSS address or network namespace cookie. It identifies the associated container, which can then be mapped to its Kubernetes Pod.netdev_linkstatus: Network device link flags that help identify conditions such as loss of carrier or a dormant interface.
Complete field list for each layers entry:
| Layer | Fields |
|---|---|
ether |
src, dst, type, len (802.3 frames only) |
ipv4 |
version, ihl, tos, len, id, flags, frag_offset, ttl, protocol, checksum, src, dst |
ipv6 |
version, traffic_class, flow_label, len, next_header, hop_limit, src, dst |
tcp |
sport, dport, seq, ack, data_offset, flags, window, checksum, urgent, sk_state |
udp |
sport, dport, len, checksum |
icmp |
type, code, checksum, id, seq |
arp |
addr_type, protocol, hw_address_size, prot_address_size, operation, sender_mac, sender_ip, target_mac, target_ip |
Userspace Filtering with jq
Use jq to select events by their JSON fields or reshape the output after kernel-side filtering:
|
|
jq -c writes each JSON object on a single line, preserving the NDJSON format for storage or further processing.
Filtering Routine Cleanup Events
Not all kfree_skb events represent actual data-plane packet drops. The following three categories are filtered by huatuo-bamai under the default configuration:
| Pattern | Stack Signature | Why It Is Not a Drop |
|---|---|---|
TCP CLOSE_WAIT + skb_rbtree_purge |
Frame skb_rbtree_purge/ |
Normal socket closure: the kernel frees in-flight SKBs in CLOSE_WAIT sockets |
| ARP/neighbor table expiry | Frame neigh_invalidate/ |
Neighbor table entry cleanup; does not affect active data flows |
| bnxt NIC TX completion | bnxt_tx_int/ |
Broadcom bnxt driver frees SKBs after DMA transmission completes; normal behavior |
In the huatuo-bamai configuration, noise rules are managed through EventTracing.IssuesList:
|
|
If you need to observe neighbor-table-related drops (for example, when debugging an ARP storm), remove the corresponding rule from IssuesList.
Continuous Monitoring with huatuo-bamai
For ongoing monitoring, huatuo-bamai runs dropwatch as a child process. The --output-storage option sends events over a Unix socket to huatuo-bamai’s processing pipeline for storage in Elasticsearch.
|
|
Once stored, you can:
- Correlate drops with latency: Overlay drop events on application latency charts in Grafana to see whether their timestamps coincide with spikes.
- Find affected containers and interfaces: Group events by
container_id,netdev_name, orlayers.labelto see where drops are concentrated. - Investigate past incidents: Query recorded events and their context without waiting for the problem to recur.
Command-Line Reference
| Parameter | Default | Description |
|---|---|---|
--bpf-path <path> |
required | Path to the eBPF object file |
--filter <expr> |
(none) | tcpdump-style filter expression |
--device <names> |
(none) | Device whitelist, comma-separated (e.g., eth0,eth1) |
--device-excluded <names> |
(none) | Device blacklist; mutually exclusive with --device |
--duration <n> |
0 | Exit after N seconds (0 = run indefinitely) |
--output <json|text> |
text |
Output format; ignored when --output-storage is set |
--output-storage <path> |
(none) | Send events to huatuo-bamai via Unix socket |
--task-id <id> |
(none) | Task ID to associate with this session |
--max-events-per-second <n> |
0 | Global report rate limit; 0 = no limit |
Summary
HUATUO dropwatch combines kernel drop locations with packet headers, process information, and call stacks. Kernel-side filters keep captures focused on the traffic you need to investigate and reduce the volume of events sent to userspace.
With huatuo-bamai, the same tracing workflow supports continuous monitoring and historical analysis. Recorded events help connect application symptoms to specific traffic and kernel code paths, giving you concrete evidence for the next step in an investigation.