🚀 Listed in the CNCF Landscape and eBPF Foundation Emerging Project. If you like this project, give it a star on GitHub! ❤️

Blog

Tracing Interrupt Spikes with irqtracing

Background

In production, we often encounter a difficult problem: several containers on the same host experience sudden service latency spikes at the same time. Historical monitoring shows that the host handled a large volume of interrupts during the incident, with sharp increases in cpu.irq and cpu.softirq utilization on multiple CPU cores. Monitoring can identify the affected cores, but it rarely preserves enough detail to reconstruct what was happening.

These incidents tend to be sporadic, lasting only tens of seconds or less. Because the interrupt spike affects only some CPU cores, changes in overall host load and individual container load may be too small to trigger the usual automatic diagnostic capture tools.

We developed irqtracing, a CLI tool, to capture evidence that can help investigate the root causes of these incidents.

Approach

We break the problem into three steps:

  1. Detect abnormal system behavior promptly, recognizing signals such as a sudden increase in interrupt activity as soon as it occurs.

  2. Capture runtime context during the incident, preserving evidence for root cause analysis while the anomaly is still present.

  3. Present the evidence clearly, turning raw data into results that help developers narrow their investigation.

Detecting Interrupt Spikes

Our detection approach is straightforward: sample interrupt utilization on each CPU at fixed intervals, then apply rules to determine whether interrupt activity has spiked.

First, calculate the share of CPU time spent handling hardware and software interrupts between adjacent samples:

1
Interrupt time share = 100% × (irq time delta + softirq time delta) / total CPU time delta

Based on observations from production, we use two detection rules. Their thresholds can be adjusted to suit the workload:

  1. Simultaneous spikes across multiple CPUs: Flag an anomaly when several CPUs show a sharp increase in interrupt utilization at the same time.

  2. Sustained high utilization on one CPU: Flag a CPU when its interrupt utilization remains high across multiple consecutive sampling windows.

The default thresholds are:

Pattern Detection rule Capture target
Spike across multiple CPUs At least 3 CPUs simultaneously increase by at least 20 percentage points, with at least 30% growth relative to the previous window The CPU with the largest increase
Sustained high utilization on one CPU One CPU reaches 80% utilization in 10 consecutive windows The affected CPU

Once these rules identify a target CPU, HuaTuo automatically triggers a capture on that CPU and collects the diagnostic data.

Capturing Runtime Context

HuaTuo already provides several tools for diagnosing interrupt problems. For example, checking whether the scheduler tick fires as expected can reveal whether interrupts have been disabled for too long and help identify why. That mechanism does not fit this case. Here, the problem is a burst of frequent interrupts rather than a prolonged period with interrupts disabled. Each interrupt takes very little time to execute, but repeated interruptions disrupt normal execution and affect system performance.

irqtracing aims to reveal where these interrupts come from, which tasks they affect, and what work they perform.

The approach relies on two tracepoints: irq:softirq_raise and irq:softirq_entry.

  • irq:softirq_raise collects the call stacks that raise softirqs on the selected CPU, identifying their sources. These paths include raising a softirq after the immediate work of a hardware interrupt handler, as well as raising one directly from other kernel code.

  • irq:softirq_entry captures the tasks present when softirq execution begins on the selected CPU, providing the victim view. Affected tasks can include kernel threads and user processes interrupted while running in either kernel mode or user mode.

Here are example call stacks for softirq_raise and softirq_entry:

softirq_raise:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
pid = 1663386
kstack:

        raise_softirq_irqoff+130
        enqueue_to_backlog+630
        netif_rx_internal+59
        __netif_rx+19
        loopback_xmit+230
        dev_hard_start_xmit+167
        __dev_queue_xmit+1253
        ip_finish_output2+818
        ip_output+100
        ip_send_skb+136
        udp_send_skb+509
        udp_sendmsg+2175
        __sock_sendmsg+103
        ____sys_sendmsg+388
        ___sys_sendmsg+658
        __sys_sendmmsg+395
        __x64_sys_sendmmsg+35
        do_syscall_64+301
        entry_SYSCALL_64_after_hwframe+118

ustack:

        __sendmmsg+82
        udp_main+502

softirq_entry:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
pid = 1663386
kstack:

        handle_softirqs+401
        do_softirq+86
        __local_bh_enable_ip+115
        __dev_queue_xmit+2813
        ip_finish_output2+818
        ip_output+100
        ip_send_skb+136
        udp_send_skb+509
        udp_sendmsg+2175
        __sock_sendmsg+103
        ____sys_sendmsg+388
        ___sys_sendmsg+658
        __sys_sendmmsg+395
        __x64_sys_sendmmsg+35
        do_syscall_64+301
        entry_SYSCALL_64_after_hwframe+118

ustack:

        __sendmmsg+82
        udp_main+502

The two call chains suggest the following sequence:

The same task raises a receive softirq inside sendmmsg() and then executes it inline after re-enabling bottom halves (BH). Both samples occur within the same send call made by PID 1663386.

In chronological order:

The same sending task raises and executes a NET_RX softirq

  • Raise identifies the source of the work: PID 1663386 sends UDP packets. The loopback path places the packets on the receive backlog and sets the current CPU’s NET_RX_SOFTIRQ pending bit.

  • Entry marks the start of execution: When the same send path re-enables BH, it finds pending work and immediately begins processing softirqs. The CPU, PID, and user stack therefore remain the same; the system call has not yet returned.

  • Matching user stacks indicate the same calling context: Both samples show udp_main → __sendmmsg. At entry, the current task is still the sending task, but execution is now in interrupt context.

These stacks do not establish that PID 1663386 is the UDP receiver. It happens to be the task in whose context softirq processing begins.

Together, these two tracepoints let us reconstruct much of the activity behind interrupt load and distinguish its sources (source) from the tasks present when softirqs execute (victim).

There is still a practical problem. A call stack can describe an individual event clearly, but we are dealing with a burst of many interrupts in a short period. One stack may be representative and even point toward the root cause, yet a handful of samples cannot characterize everything that happened. Collecting more samples creates another challenge familiar to anyone who has debugged production systems: a terminal flooded with stack traces. Extracting useful clues from that volume of output takes effort, even with AI assistance.

To capture enough evidence, reduce misleading conclusions from sparse samples, and make the result easy to read, we use a flame graph. During a short capture window, we collect as many interrupt events as possible, then aggregate their call stacks into two views, source and victim, within a single flame graph.

Softirq capture and aggregation into source and victim views

Here is an example:

An irqtracing flame graph with separate source and victim roots

In this example:

  1. A total of 3,709 captured stack traces are aggregated.

  2. The source side shows 2,000 occurrences of the same call stack raising NET_RX softirqs.

  3. The victim side records 1,706 NET_RX softirq entry events with the same process, PID 1740795, as the current task.

  4. Two RCU softirq events also occur while that process is running. Their bars are so narrow that they are difficult to see beside the much larger NET_RX count.

This view supports several common investigation needs:

First, it turns individual stack traces into a distribution of activity. A call path that appears repeatedly occupies a wider region. Investigators can start with the dominant softirq type and then inspect its call chains, rather than reading every stack trace separately.

Second, source and victim each have their own root in the same result. If the source side is dominated by a particular interrupt type, such as NET_RX, while a business process appears frequently on the victim side, that suggests a concrete question to investigate: is network softirq activity competing with that process for CPU time? Both views retain softirq types and task labels, making it possible to narrow the investigation from event categories to specific call paths and tasks.

Third, it preserves the distinction between the task that raises a softirq and the task present when it executes. Merging every stack into one tree, as in a conventional flame graph, can draw attention only to the widest functions. Separate views expose both where softirqs are frequently raised and which tasks are often on the CPU when those softirqs execute.

This is an aggregate view for finding investigative leads. It cannot establish that a particular source event interrupted a particular victim. Bar width represents event counts and must not be interpreted as elapsed time.

Trying irqtracing

Combining detection and capture, irqtracing works with HuaTuo as follows:

  1. Apply the detection rules to identify an anomaly and select a target CPU.

  2. Invoke the irqtracing CLI to capture interrupt activity.

  3. In automatic tracing mode, let HuaTuo save the CLI’s capture results.

  4. Alternatively, run the irqtracing CLI independently to produce a flame graph directly.

irqtracing anomaly detection, runtime capture, and result storage

To run a capture manually, build the project, select a CPU, and write the result as an SVG flame graph:

1
2
3
4
5
6
make build
sudo ./_output/bin/irqtracing \
  --bpf-path ./_output/bpf/irqtracing.o \
  --target-cpu 1 \
  --duration 3 \
  --output flamegraph > irqtracing.svg

The current approach has limitations. The rules for deciding whether IRQ utilization constitutes a spike are empirical and leave room for better detection methods. A more rigorous algorithm could improve capture accuracy and could also apply to similar metrics, such as cpuidle and cpusys. Softirqs can also execute in thread context, where their impact on other tasks is primarily scheduling delay. We have not included analysis of that case in the tool’s scope, because doing so would add substantial complexity for limited benefit.

irqtracing is a step toward understanding these incidents and preserving useful evidence for production debugging. We welcome users to try it and suggest improvements.