How to Profile Interrupt Latency and ISR Execution Time on Cortex-M
Your ISR takes “maybe 2 microseconds.” You’re pretty sure. You did the math on the instructions, eyeballed the loop, and called it good. Then a deadline slips. Then another. You add a safety margin, bump a timer period, and move on. Six months later, someone adds a feature inside that ISR, and the whole system falls over at 3 AM on a Saturday.
The problem isn’t the ISR. It’s the “maybe.”
On any interrupt-driven firmware, profiling ISR timing is basic hygiene. And on Cortex-M, the hardware practically begs you to measure it. You just have to know where to look.
This article covers 3 concrete techniques to put real numbers on interrupt latency (how long from the hardware event to your first ISR instruction) and ISR execution time (how long your handler actually runs). By the end, you’ll be able to set up at least one of them on your own hardware, probably in under 10 minutes.
Two Quantities, Not One
People say “interrupt latency” when they mean different things. Let’s be precise.
Interrupt latency is the time from when the hardware peripheral asserts its interrupt request to when the first instruction of your ISR executes. On Cortex-M, this includes the automatic context save (stacking registers onto the stack), which the hardware does without any help from your code.
ISR execution time is the time from that first ISR instruction to the exception return.
INTERRUPT TIMELINE
==================
IRQ Source ──────┐
(HW event) │
▼
┌───────────┐ ┌──────────────────────┐
│ STACKING │ │ ISR EXECUTING │
│ (HW auto) │ │ (your handler code) │
└───────────┘ └──────────────────────┘
|<----------->|<-------------------------->|
Interrupt ISR Execution
Latency Time
|<-------------------------------------------->|
Total Interrupt Response TimeDifferent things inflate each number. Latency grows when higher-priority interrupts are pending, when interrupts are masked (PRIMASK/BASEPRI), when flash wait states stall fetches, or when bus contention holds up the stacking writes. Execution time grows from algorithm complexity, cache misses on M7, or (please don’t) blocking calls inside the handler.
You need to measure both separately. Conflating them hides the real bottleneck.
What Cortex-M Gives You for Free
ARM designed profiling support directly into the core. These aren’t vendor-specific peripherals; they work the same whether you’re on an STM32, an NXP LPC, or a Nordic nRF.
DWT Cycle Counter (Cortex-M3/M4/M7/M33). The Data Watchpoint and Trace unit includes a 32-bit free-running counter (DWT->CYCCNT) that ticks once per core clock cycle. No pins, no extra hardware, no trace probe required.
SysTick Timer (all Cortex-M). Every Cortex-M has a SysTick, a 24-bit down-counter. M0 and M0+ lack DWT, so SysTick is your fallback on those cores. Resolution is 1 cycle if you clock it from the processor clock directly, which is the common default.
ITM/SWO Trace is worth mentioning for completeness. If your debugger supports it, you can stream timestamped trace data out through the SWO pin. But it’s tool-specific, so I won’t lean on it here.
Technique 1: DWT Cycle Counter for ISR Profiling (Software-Only)
This is where you should start on any Cortex-M3 or above. No extra hardware. No special debugger. Just a few lines of C using CMSIS register definitions.
Step 1: Enable the DWT cycle counter.
// Enable trace and debug blocks
CoreDebug->DEMCR |= CoreDebug_DEMCR_TRCENA_Msk;
// Reset the cycle counter
DWT->CYCCNT = 0;
// Enable the cycle counter
DWT->CTRL |= DWT_CTRL_CYCCNTENA_Msk;Run this once during init, before your interrupts are enabled.
Step 2: Capture timestamps in your ISR.
volatile uint32_t isr_entry_cyccnt;
volatile uint32_t isr_exit_cyccnt;
volatile uint32_t trigger_cyccnt; // set at the point you expect the IRQ
void TIM2_IRQHandler(void)
{
isr_entry_cyccnt = DWT->CYCCNT; // first line!
// ... your actual ISR work here ...
// (clear interrupt flag, process data, etc.)
isr_exit_cyccnt = DWT->CYCCNT; // last line before return
}Step 3: Compute the numbers.
uint32_t latency_cycles = isr_entry_cyccnt - trigger_cyccnt;
uint32_t exec_cycles = isr_exit_cyccnt - isr_entry_cyccnt;
// Convert to microseconds
uint32_t latency_us = latency_cycles / (SystemCoreClock / 1000000U);
uint32_t exec_us = exec_cycles / (SystemCoreClock / 1000000U);For interrupt latency, you need a known trigger point. With a timer compare event, capture DWT->CYCCNT right before you write the register that arms the interrupt. For external interrupts, this is trickier; you may need the GPIO technique below instead.
Caveats worth knowing. The counter wraps at 2^32 cycles. At 168 MHz, that’s about 25.5 seconds, so you’re fine for ISR-scale measurements. The DWT->CYCCNT read itself adds roughly 1 to 2 cycles of overhead. Mark your timestamp variables volatile so the compiler doesn’t reorder the reads. On higher optimization levels, consider a __DSB() memory barrier before the final read if you’re concerned about pipeline effects.
Technique 2: Measuring Interrupt Latency with GPIO and a Scope
When you need ground truth, nothing beats an oscilloscope or logic analyzer watching actual pin transitions. This technique is immune to software measurement overhead and captures jitter distribution over millions of events.
Step 1: Pick a GPIO pin. Configure it as push-pull output, high-speed mode.
Step 2: Set the pin HIGH as the absolute first action in your ISR. Set it LOW as the last action before return.
Step 3: Trigger your scope on the interrupt source. For an external interrupt, trigger on the input pin edge. For a timer interrupt, you can route the timer output compare to a pin.
Step 4: Measure.
Ext. IRQ pin ────┐ ┌──────────────────────────
(trigger) └─────────┘
GPIO probe ─────────────────┐ ┌────────
(ISR toggle) └──────────────┘
|<--------->|<------------>|
Interrupt ISR Execution
Latency Time
(+ GPIO (approx.)
overhead)The rising edge of your GPIO relative to the triggering event gives you interrupt latency (plus a small GPIO write overhead, typically 1 to 3 cycles). The pulse width gives you ISR execution time. Run in “infinite persistence” or histogram mode on your scope, and you’ll see the full jitter distribution without writing any statistics code.
Caveats. The GPIO write itself goes through the AHB bus, adding a consistent overhead of a few nanoseconds. It’s repeatable, so you can calibrate it out: toggle the GPIO in a tight loop and measure the minimum pulse width. That’s your floor.
Technique 3: SysTick as a Fallback for M0/M0+
Cortex-M0 and M0+ don’t have DWT, but they do have SysTick, and it counts at 1 cycle resolution if you clock it from the processor clock (check your vendor’s default; most do).
SysTick counts down from its reload value, so the math flips:
void SomeIRQHandler(void)
{
uint32_t entry_tick = SysTick->VAL; // counts down
// ... ISR work ...
uint32_t exit_tick = SysTick->VAL;
// Handle wrap-around (down-counter!)
uint32_t exec_cycles;
if (exit_tick <= entry_tick) {
exec_cycles = entry_tick - exit_tick;
} else {
exec_cycles = entry_tick + (SysTick->LOAD + 1 - exit_tick);
}
}The resolution matches DWT (1 core cycle), but you’re limited to 24 bits, which wraps at 16.7 million cycles; still plenty for ISR measurement. The main downside: if your RTOS is already using SysTick for its tick, you need to be careful not to disturb the reload value. Reading SysTick->VAL is non-destructive, though, so the read itself is safe.
Building a Statistics Buffer, Not a Single Sample
A single measurement tells you almost nothing. ISR timing varies based on cache state, bus contention, and what other interrupts are doing. You need distributions.
Keep the statistics collection inside the ISR cheap:
typedef struct {
uint32_t min;
uint32_t max;
uint64_t sum;
uint32_t count;
} isr_stats_t;
volatile isr_stats_t stats = { .min = UINT32_MAX, .max = 0 };
// Inside the ISR, after computing exec_cycles:
if (exec_cycles < stats.min) stats.min = exec_cycles;
if (exec_cycles > stats.max) stats.max = exec_cycles;
stats.sum += exec_cycles;
stats.count++;This adds maybe 10 to 15 cycles of overhead per interrupt. Negligible for profiling purposes.
Dump the results via UART after your test run, read them out in a debugger memory view, or use semihosting if your setup supports it. Never print from inside the ISR. Collect first, report separately. A printf call inside a handler will destroy the very timing you’re trying to measure.
If you want a histogram, a simple 32-bucket array (each bucket covering, say, 10 cycles) fits in a few hundred bytes of RAM and paints a clear picture of your distribution.
Expected Cortex-M Latency: Your Sanity Check
These are best-case numbers from the ARM technical reference manuals, assuming zero-wait-state memory and no pending higher-priority interrupt:
| Cortex-M Variant | Best-Case Latency (cycles) | Notes |
|---|---|---|
| M0 / M0+ | 16 | No tail-chaining |
| M3 / M4 | 12 | With tail-chaining |
| M7 | 12+ | Varies with cache/TCM config |
| M33 (TrustZone) | 13+ | Secure-to-secure; more across security boundaries |
If your measured latency is close to these numbers, your system is healthy. If it’s 2x or 5x higher, something is inflating it.
When Your Numbers Look Wrong: Cortex-M Interrupt Debugging Pointers
Latency way higher than expected. First suspect: someone is masking interrupts. Search your codebase for __disable_irq(), __set_PRIMASK(1), or __set_BASEPRI(). Critical sections that hold interrupts off for hundreds of cycles are a common culprit. Also check your flash wait-state configuration and whether prefetch/instruction cache is enabled.
Execution time swings wildly between samples. Look for data-dependent branches in the ISR. On Cortex-M7, instruction and data cache misses can add dozens of cycles per miss. Shared resources (like a DMA controller hogging the bus) also cause variability.
Jitter on GPIO measurements. Verify the GPIO clock is enabled and running at full speed. Confirm the pin is in push-pull mode, not open-drain. Check that no DMA transfer is stalling the AHB bus during your measurement window.
These are directions to investigate, not full root-cause analysis. Once you have the numbers, the debugging gets a lot more targeted.
Add DWT Profiling Today
If you’re on Cortex-M3 or above, adding the DWT cycle counter takes about 5 minutes. Three lines in your init, two reads in your ISR, and you’ve gone from “maybe 2 microseconds” to a real number you can make decisions on.
If you’re working on a BLE device where tight ISR timing affects advertising schedules, the Hubble terrestrial SDK API overview covers how the advertising packet structure interacts with your firmware’s timing constraints.
For M0/M0+ projects, SysTick gets you the same resolution with slightly more bookkeeping. And for the times you need to see jitter across thousands of events, a spare GPIO and a $15 logic analyzer will show you things no software measurement can.
That number might confirm your intuition, or it might surprise you. Either way: measure first, optimize second. You’ll save yourself a lot of Saturday mornings.
Hubble Network connects your Cortex-M devices directly to satellite from a single Bluetooth chip—no gateways, no extra radios. See how it works →