🚀 Listed in the CNCF Landscape and eBPF Foundation Emerging Project. If you like this project, give it a star on GitHub! ❤️

Blog

Where Packets Get Dropped: Tracing the Linux Network Stack with eBPF

Connection timeouts and rising TCP retransmission rates tell you something is wrong. They do not tell you where packets are being dropped. Is a firewall rule rejecting traffic, a route lookup failing, or a receive buffer filling up? Tools such as dropwatch and perf record -e skb:kfree_skb can help identify drop locations, but a location alone does not tell you whether a dropped packet belongs to the affected application.

HUATUO’s dropwatch adds packet context to kernel drop tracing. An eBPF probe attached to tracepoint/skb/kfree_skb captures the IP five-tuple, process information, network device, MAC addresses, and kernel call stack. It also supports tcpdump-style filters compiled into eBPF bytecode at load time. Filtering in the kernel ensures that only matching events reach userspace, keeping the output focused on the traffic under investigation.

This walkthrough covers targeted captures, event analysis, and continuous monitoring with huatuo-bamai, starting with a service experiencing intermittent latency spikes.


Investigating a Latency Spike

In a Kubernetes cluster, the P99 latency of an order service periodically spikes from 50 ms to 2 s. Application logs show sporadic connection reset by peer errors, and netstat -s on the node reveals a steady increase in TCPLostRetransmit.

An investigation might involve capturing traffic with tcpdump, inspecting firewall rules with iptables -L, checking routing with ip route get, and tracing kernel drop locations with perf record -e skb:kfree_skb. Connecting those observations to the affected service requires packet context as well as drop locations.

Start with a 60-second dropwatch capture filtered to TCP port 8080 on eth0:

1
2
3
4
5
sudo dropwatch --bpf-path bpf/dropwatch.o \
  --filter "tcp and port 8080" \
  --device eth0 \
  --duration 60 \
  --output json

The output brings process information, source and destination addresses and ports, and the kernel call stack together in each event. Use these fields to identify relevant traffic, then filter with jq to examine the drop paths associated with the incident.


Filtering Drop Events in the Kernel

On busy hosts, kfree_skb events can be frequent. Some reflect routine cleanup rather than packet loss affecting application traffic, as discussed in the noise-filtering section below. Sending every event to userspace increases processing overhead and makes relevant events harder to find.

Dropwatch evaluates traffic filters in the kernel. Its pure-Go pcap compiler (internal/pcapfilter) compiles tcpdump-style expressions into eBPF bytecode when the program loads. That bytecode runs within the probe, which submits only matching events to userspace through the perf ring buffer.

Supported Filter Primitives

internal/pcapfilter implements a practical subset of the standard tcpdump syntax:

Protocol and Direction

1
2
3
4
5
ip   ip6   tcp   udp   icmp   icmp6   arp
ip proto tcp      ip6 proto udp
src host 10.0.0.1    dst host 10.0.0.1
src port 443         dst port 8080
src net 192.168.1.0/24

Boolean Composition

1
2
3
4
tcp and port 443
tcp or udp
not arp
ip and src net 192.168.1.0/24 and tcp dst port 3306

Unsupported expressions include byte offsets (tcp[tcpflags]), numeric protocol numbers (ip proto 6), and port ranges (portrange). Refer to the usage documentation for the complete list of unsupported features.

Filter Examples

1
2
3
4
5
6
7
8
# Monitor only TCP drops on a target port
--filter "tcp and port 443"

# Exclude noise from the metadata service IP
--filter "tcp and not host 169.254.169.254"

# Pinpoint drops from a specific subnet to a database port
--filter "src net 192.168.1.0/24 and tcp dst port 3306"

Traffic filters (--filter) and device filters (--device / --device-excluded) are combined with AND semantics: an event must satisfy both. With --device, events for SKBs that have no net_device are excluded from the output. With --device-excluded, those events pass the device filter. These filters control event reporting, not packet forwarding; keep this distinction in mind when tracing container veth traffic.


Reading a Drop Event

With --output json, dropwatch emits one JSON object per event (NDJSON). The following example is formatted for readability:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
{
  "observed_timestamp": "2026-06-20T08:30:12.123456789Z",
  "comm": "nginx",
  "pid": 4521,
  "container_id": "a1b2c3d4",
  "netdev_name": "eth0",
  "packet_eth_proto": "0x0800",
  "packet_len": 74,
  "layers": {
    "label": "IPv4/TCP",
    "ether": { "src": "aa:bb:cc:dd:ee:ff", "dst": "11:22:33:44:55:66", "type": "IPv4" },
    "ipv4": { "src": "10.0.1.5", "dst": "10.0.2.10", "ttl": 64, "protocol": "TCP" },
    "tcp": { "sport": 54321, "dport": 8080, "flags": "SYN", "seq": 0, "ack": 0, "window": 65535, "sk_state": "SYN_SENT" }
  },
  "stack": "kfree_skb\ntcp_v4_rcv\ntcp_rcv_established\n..."
}

Notable fields:

  • layers: Parsed packet headers, with a nested object for each protocol layer. Missing layers are omitted. Consumers can check which objects are present to identify the protocols, for example with ev.Layers.TCP != nil.
  • stack: The kernel call stack at the drop site, with frames separated by newlines. Use it to identify the code path involved: tcp_v4_rcv points to receive processing, while ip_output points to the output path.
  • container_id: Resolved by huatuo-bamai using the memory cgroup CSS address or network namespace cookie. It identifies the associated container, which can then be mapped to its Kubernetes Pod.
  • netdev_linkstatus: Network device link flags that help identify conditions such as loss of carrier or a dormant interface.

Complete field list for each layers entry:

Layer Fields
ether src, dst, type, len (802.3 frames only)
ipv4 version, ihl, tos, len, id, flags, frag_offset, ttl, protocol, checksum, src, dst
ipv6 version, traffic_class, flow_label, len, next_header, hop_limit, src, dst
tcp sport, dport, seq, ack, data_offset, flags, window, checksum, urgent, sk_state
udp sport, dport, len, checksum
icmp type, code, checksum, id, seq
arp addr_type, protocol, hw_address_size, prot_address_size, operation, sender_mac, sender_ip, target_mac, target_ip

Userspace Filtering with jq

Use jq to select events by their JSON fields or reshape the output after kernel-side filtering:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
# Show only RST packets
sudo dropwatch --bpf-path bpf/dropwatch.o --output json 2>/dev/null \
  | jq 'select(.layers.tcp.flags == "RST")'

# Exclude events whose call stack contains ip_finish_output (typically normal routing output)
sudo dropwatch --output json --duration 10 --bpf-path bpf/dropwatch.o \
  | jq -c 'select(.stack | test("ip_finish_output") | not)'

# Show metadata only, without the call stack (useful for quick drop distribution statistics)
sudo dropwatch --output json --duration 10 --bpf-path bpf/dropwatch.o \
  | jq -c 'del(.stack)'

jq -c writes each JSON object on a single line, preserving the NDJSON format for storage or further processing.


Filtering Routine Cleanup Events

Not all kfree_skb events represent actual data-plane packet drops. The following three categories are filtered by huatuo-bamai under the default configuration:

Pattern Stack Signature Why It Is Not a Drop
TCP CLOSE_WAIT + skb_rbtree_purge Frame skb_rbtree_purge/ Normal socket closure: the kernel frees in-flight SKBs in CLOSE_WAIT sockets
ARP/neighbor table expiry Frame neigh_invalidate/ Neighbor table entry cleanup; does not affect active data flows
bnxt NIC TX completion bnxt_tx_int/ Broadcom bnxt driver frees SKBs after DMA transmission completes; normal behavior

In the huatuo-bamai configuration, noise rules are managed through EventTracing.IssuesList:

1
2
3
4
5
6
[EventTracing]
    IssuesList = [["neigh_invalidate", "neigh_invalidate"], ["bnxt_tx_int", "bnxt_tx_int"]]

[EventTracing.Dropwatch]
    Filter = "tcp"
    MaxEventsPerSecond = 100

If you need to observe neighbor-table-related drops (for example, when debugging an ARP storm), remove the corresponding rule from IssuesList.


Continuous Monitoring with huatuo-bamai

For ongoing monitoring, huatuo-bamai runs dropwatch as a child process. The --output-storage option sends events over a Unix socket to huatuo-bamai’s processing pipeline for storage in Elasticsearch.

1
2
3
4
dropwatch \
  --bpf-path <CoreBpfDir>/dropwatch.o \
  --output-storage /var/run/huatuo/events.sock \
  --filter "tcp"

Once stored, you can:

  • Correlate drops with latency: Overlay drop events on application latency charts in Grafana to see whether their timestamps coincide with spikes.
  • Find affected containers and interfaces: Group events by container_id, netdev_name, or layers.label to see where drops are concentrated.
  • Investigate past incidents: Query recorded events and their context without waiting for the problem to recur.

Command-Line Reference

Parameter Default Description
--bpf-path <path> required Path to the eBPF object file
--filter <expr> (none) tcpdump-style filter expression
--device <names> (none) Device whitelist, comma-separated (e.g., eth0,eth1)
--device-excluded <names> (none) Device blacklist; mutually exclusive with --device
--duration <n> 0 Exit after N seconds (0 = run indefinitely)
--output <json|text> text Output format; ignored when --output-storage is set
--output-storage <path> (none) Send events to huatuo-bamai via Unix socket
--task-id <id> (none) Task ID to associate with this session
--max-events-per-second <n> 0 Global report rate limit; 0 = no limit

Summary

HUATUO dropwatch combines kernel drop locations with packet headers, process information, and call stacks. Kernel-side filters keep captures focused on the traffic you need to investigate and reduce the volume of events sent to userspace.

With huatuo-bamai, the same tracing workflow supports continuous monitoring and historical analysis. Recorded events help connect application symptoms to specific traffic and kernel code paths, giving you concrete evidence for the next step in an investigation.