How to Implement a Ring Buffer for Embedded Logging and Telemetry

Your ISR fires, writes a log entry halfway into the buffer, and then the main loop starts reading from that same buffer. You get half a message, a corrupted pointer, maybe a hard fault if you’re unlucky. The ugly part: this bug only shows up at -O2, only under load, and never while your debugger is attached.
If that sounds familiar, you’ve outgrown the textbook ring buffer. What you need is a production SPSC (single-producer, single-consumer) ring buffer in bare-metal C that’s ISR-safe without disabling interrupts. And you probably need two variants: one for runtime debug logging (variable-length, lossy, human-readable) and one for sensor telemetry capture (fixed-size structs, lossless, binary). Same core structure, different rules.
Here’s how to build both.
Core Ring Buffer Implementation in C
Every field earns its place.
typedef struct {
uint8_t *buf;
volatile uint32_t head; // ← written ONLY by producer
volatile uint32_t tail; // ← written ONLY by consumer
uint32_t size; // must be power of two
uint32_t mask; // size - 1
} ringbuf_t;Why volatile on head and tail? Because the compiler will otherwise cache them in registers across function calls. At -O2, GCC will optimize away a re-read of tail inside your push function: it can prove that your code path never writes to it. But the consumer, running in a different context, does. volatile forces the reload.
Power-of-two sizing eliminates the modulo operator. On a Cortex-M0 with no hardware divider, index % size compiles to a library call. index & mask compiles to a single AND instruction. The difference matters at ISR rates.
Buffer size = 8 (mask = 0x07)
Index: 0 1 2 3 4 5 6 7
┌───┬───┬───┬───┬───┬───┬───┬───┐
│ │ D │ D │ D │ D │ │ │ │
└───┴───┴───┴───┴───┴───┴───┴───┘
^ ^
tail=1 head=5
head & mask = 5 & 0x07 = 5 ✓
After 4 more pushes: head=9, 9 & 0x07 = 1 → wraps!Notice that head and tail are free-running unsigned integers. They never get reset to zero. You only mask when you index into the buffer. This avoids the ambiguity where head == tail could mean either full or empty.
void ringbuf_init(ringbuf_t *rb, uint8_t *buf, uint32_t size) {
rb->buf = buf;
rb->size = size;
rb->mask = size - 1;
rb->head = 0;
rb->tail = 0;
}
bool ringbuf_is_full(ringbuf_t *rb) {
return (rb->head - rb->tail) == rb->size;
}
bool ringbuf_is_empty(ringbuf_t *rb) {
return rb->head == rb->tail;
}
bool ringbuf_push(ringbuf_t *rb, uint8_t byte) {
if (ringbuf_is_full(rb)) return false;
rb->buf[rb->head & rb->mask] = byte;
rb->head++; // ← only producer writes this
return true;
}
bool ringbuf_pop(ringbuf_t *rb, uint8_t *byte) {
if (ringbuf_is_empty(rb)) return false;
*byte = rb->buf[rb->tail & rb->mask];
rb->tail++; // ← only consumer writes this
return true;
}The key insight: producer only writes head, consumer only writes tail. Each side reads the other’s index but never writes it. That one-way ownership is what makes lock-free SPSC possible.
Making It ISR-Safe: Memory Barriers and Volatile
volatile handles compiler reloads, but it doesn’t handle ordering. The compiler can reorder the store to buf[head & mask] so it lands after the increment of head. The consumer then sees the new head, reads the slot, and gets stale data.
On ARM Cortex-M (single-core, in-order pipeline), the hardware won’t reorder stores. But the compiler still might. You need a compiler barrier between the data write and the index update:
rb->buf[rb->head & rb->mask] = byte;
__asm volatile("" ::: "memory"); // ← compiler barrier
rb->head++;On Cortex-M3/M4/M33, this is sufficient. On multi-core targets (Cortex-A, dual-core RP2040), you’d need a hardware barrier: __DMB() after the data store. If you’re running on the Hubble device SDK or similar BLE stacks on Nordic or TI parts, you’re single-core, so compiler barriers are all you need.
A note on honesty: If you’re disabling interrupts around every push/pop, you don’t have a lock-free buffer. You have a mutex. That’s sometimes the right call (multi-producer scenarios, mixed ISR priorities), but know what you’re paying for: worst-case interrupt latency grows by the duration of your critical section.
Logging vs. Telemetry: Two Buffers, Different Rules
Most tutorials give you the generic ring buffer and wave goodbye. But the requirements for a debug log buffer and a telemetry capture buffer are different enough that using the wrong pattern causes real problems.
| Concern | Logging | Telemetry |
|---|---|---|
| Record size | Variable-length | Fixed-size struct |
| Encoding | ASCII / formatted | Binary / packed |
| Overflow policy | Overwrite oldest | Reject newest, count drops |
| Consumer | UART drain, debug probe | DMA, radio TX, flash write |
| Latency tolerance | High (human reader) | Low (real-time) |
| Integrity need | Nice to have | Critical (CRC/checksum) |
Buffer full?
│
┌────┴────┐
│ YES │ NO
▼ ▼
┌─────────────┐ Push normally
│ Use case? │
├──────┬──────┤
▼ ▼
LOG TELEMETRY
│ │
▼ ▼
Advance Reject write,
tail increment
(lose dropped_count
oldest)Logging: Byte Buffer, Overwrite-Oldest
For logging, you operate at the byte level. Each log entry gets a length prefix (1 or 2 bytes) followed by the ASCII payload. When the buffer is full, you advance the tail to make room, discarding the oldest entry.
void log_write(ringbuf_t *rb, const uint8_t *data, uint8_t len) {
// Make room if needed by advancing tail
while ((rb->size - (rb->head - rb->tail)) < (uint32_t)(len + 1)) {
uint8_t old_len = rb->buf[rb->tail & rb->mask];
rb->tail += old_len + 1; // skip past oldest entry
}
rb->buf[rb->head & rb->mask] = len; // length prefix
rb->head++;
for (uint8_t i = 0; i < len; i++) {
rb->buf[rb->head & rb->mask] = data[i];
rb->head++;
}
}Old logs are the least valuable. Dropping them to keep the freshest messages is the right tradeoff for debug output.
Telemetry: Struct Slots, Reject-Newest
For telemetry, each slot holds a fixed-size record. You’re buffering sensor readings, GPS fixes, or packed BLE advertising payloads before they get transmitted. Silently overwriting a telemetry record is dangerous; your system might depend on that data. Reject the write and count the drop instead.
typedef struct {
uint32_t timestamp;
int16_t accel_x, accel_y, accel_z;
uint16_t crc;
} telemetry_record_t;
typedef struct {
telemetry_record_t *buf;
volatile uint32_t head;
volatile uint32_t tail;
uint32_t size;
uint32_t mask;
volatile uint32_t dropped_count; // ← monitor this
} telem_ringbuf_t;
bool telem_push(telem_ringbuf_t *rb, const telemetry_record_t *rec) {
if ((rb->head - rb->tail) == rb->size) {
rb->dropped_count++;
return false; // ← reject, don't overwrite
}
rb->buf[rb->head & rb->mask] = *rec;
__asm volatile("" ::: "memory");
rb->head++;
return true;
}That dropped_count isn’t just bookkeeping. You can include it in your next telemetry packet or log line to catch buffer sizing issues before they become field failures.
Don’t mix these policies. Overwriting telemetry silently loses data your application may depend on. Rejecting log messages wastes buffer space on stale entries nobody cares about.
Practical Hardening
A few things separate a “works on the bench” buffer from one that survives production.
High-water mark tracking. Add a uint32_t hwm field. After every push, compare (head - tail) to hwm and update if larger. This tells you how close your buffer came to full in the field. If your 256-byte log buffer peaks at 248, you need a bigger buffer or a faster drain.
Drain hooks. Register a callback or set a flag when occupancy exceeds a threshold (say 75%). The main loop checks this flag and kicks off a UART drain or flash page write. This turns your buffer into an active flow-control mechanism rather than passive storage.
Compile-time size enforcement:
_Static_assert((LOG_BUF_SIZE & (LOG_BUF_SIZE - 1)) == 0,
"Buffer size must be power of two");Host testability. The entire ring buffer module should compile on x86 with no MCU headers. Write your tests with a simple assertion harness or your favorite framework. Fill it, drain it, overflow it, check every edge. If you’ve structured the code without MCU dependencies, you can exercise it thoroughly before it ever touches silicon.
Pitfalls That’ll Bite You at 2 AM
- Using
%instead of& mask. On Cortex-M0 and AVR, division is a function call. At 10 kHz ISR rates, you’ll feel it. - Forgetting
volatileon head/tail. Works fine at-O0. Breaks silently at-O2when the compiler caches the index in a register. The worst kind of optimization-dependent bug, because your debug build never reproduces it. - Multiple producers at different ISR priorities. If your ADC ISR and your timer ISR both push to the same buffer, you’ve broken the single-producer assumption. Either merge them into one producer or protect with a critical section.
- Allocating the buffer on the stack. The function returns, the stack frame gets reused, and your buffer pointer now aims at someone else’s local variables.
- Confusing full and empty. With free-running indices,
head == tailmeans empty andhead - tail == sizemeans full. No ambiguity, no wasted slot.
Picking Your Pattern
The core data structure is identical: a power-of-two buffer, a head, a tail, a mask. Logging and telemetry diverge at the policy layer. Start with the byte-level SPSC buffer for logging. Get it solid. Run it on your host machine with adversarial test cases. Then layer the fixed-size telemetry variant on the same foundation.
If you’re building BLE devices that ship telemetry through Hubble’s network, the terrestrial SDK reference shows how advertising packets get structured. Your telemetry ring buffer sits right between the sensor read and that advertising payload. Size it based on your worst-case transmit interval, add the dropped counter, and you’ll have data you can trust.
Hubble Network connects your embedded devices directly from a ring buffer to the cloud—no gateways, no relay infrastructure. See how it works →