SDR(Software Defined Radio)

 

 

 

Memory and Buffering

An SDR's streaming problem is not just a memory-capacity problem. It has a chain of buffers, and every one of them makes the same trade in a different place. I spent a long time tuning the wrong buffer, which seems to be the usual way to learn this. Each section below takes one stage of that chain. The order follows a sample from the data converter to the application, because that is the order the failures propagate in.

Executive Summary

The table is a lookup. Equivalent sample time is the buffer size in bytes divided by the payload rate. The examples use 800 MB/s with decimal MB; actual waiting time depends on occupancy and service.

Stage

What it absorbs

Depth and sizing

Who sets it

What it cannot do

FPGA FIFO

Clock-domain crossing when designed as an asynchronous FIFO, and short downstream stalls.

Image-specific. For illustration, 64 KiB holds 81.92 microseconds at 800 MB/s.

The FPGA image; changing depth may require a custom build and available resources.

Apply backpressure to the antenna. When it fills, samples are lost.

On-board DRAM

Finite capture or playback, and stalls when the FPGA explicitly routes the stream through DRAM.

1.25 seconds for 1 GB, or about 1.34 seconds for 1 GiB, at 800 MB/s.

The FPGA image, through a block such as RFNoC Replay.

Automatically protect every stream. The image must expose and connect the memory; controller and access behaviour affect latency.

Transport buffer

Host scheduling jitter, switch queueing and interrupt delay.

Transport-specific. An 8 MB usable payload buffer represents 10 ms at 800 MB/s.

Driver parameters, and the operating system limits behind them.

Fix an average rate deficit. It only delays the overflow.

Application ring

Variation in the processing time per block.

Application-defined, subject to memory, queue limits and the permitted processing latency.

The application, and it is the one stage fully under your control.

Repair a sustained processing deficit. Occupied bytes add waiting time; unused capacity alone does not.

DMA descriptor ring

The gap between the device writing and the driver reclaiming.

Posted descriptors times usable bytes per descriptor, when each descriptor owns one fixed-size buffer.

The driver and hardware; supported settings may be applied during setup or through runtime configuration.

Access arbitrary application pointers. DMA mappings and backing storage must remain valid until completion.

Why is an SDR a buffering problem before it is a memory problem?

You can have plenty of free RAM and still lose samples. The converter keeps producing data while the host processes blocks at irregular intervals, and only the buffers on that path can absorb the difference.

Start with the payload rate: R = fs * C * B. Here fs is the complex sample rate per channel, C is the channel count, and B is bytes per complex sample. Two channels at 100 MS/s with 16-bit I and 16-bit Q produce 800 MB/s. MB means one million bytes here. Transport headers and line coding require additional link capacity.

A buffer has both a capacity and an occupancy. Capacity says how much it could hold. Occupancy says how much data is waiting now. For reception, free capacity determines how long the consumer can stop. For transmission, queued playable samples determine how long the producer can stop.

At 800 MB/s, a 1 ms pause adds 0.8 MB to an RX queue. But that calculation works only if the data can reach that queue. Free application memory cannot protect a full FPGA FIFO when the transport itself has stopped.

A temporary pause is different from a sustained rate deficit. If samples arrive at 800 MB/s and processing removes only 700 MB/s, occupancy grows by 100 MB each second. Another 100 MB delays failure by about one second; it does not make continuous operation possible. After a temporary pause, the consumer also needs spare throughput to remove the accumulated backlog.

  • Calculate the rate at each stage. Sample format changes can change the byte rate without changing the sample rate.
  • Free space must be reachable. Buffers protect the stages connected to them, not every stage in the system.

What does the FIFO inside the FPGA actually do?

The FPGA processes samples on a clock, but its next stage may accept data in bursts. A FIFO lets those two stages operate at different instants without immediately losing the sample sequence.

A synchronous FIFO uses one clock domain. An asynchronous FIFO bridges different clock domains using a design that safely transfers control information between them. The FIFO preserves ordering while its read and write sides advance independently. It does not permit an unlimited difference between their average rates.

Flow control can pause an upstream digital block when the next FIFO becomes full. However, a continuously operating ADC cannot postpone the arrival of the RF signal. If the entire receive path stops accepting data, some finite storage eventually fills. The implementation must then drop data, stop capture or report another defined failure.

For example, a hypothetical 64 KiB FIFO holds 65,536 bytes. At 800 MB/s, that represents only 81.92 microseconds when empty. If it is already half full, its remaining protection against a complete downstream stall is about 40.96 microseconds. These values illustrate sizing; they are not specifications for a particular FPGA image.

On-board DRAM can provide much more capacity, but the image must connect that memory into the intended data path. Merely having DRAM on the board does not make every stream use it. RFNoC Replay provides explicit recording and playback through external memory on supported configurations. Its usable address space and port arrangement depend on the image.

FIFO depth therefore belongs to the hardware design. A custom image may change it within resource limits, while a normal host parameter usually cannot.

  • FIFO capacity covers short mismatches. Clock-domain crossing and burst absorption do not solve sustained overload.
  • DRAM needs an explicit data path. Verify the loaded FPGA configuration before counting its capacity as streaming protection.

Why is a circular buffer the right structure for a sample stream?

A sample stream has no natural end, but its storage must have a fixed size. A circular buffer reuses that storage as the consumer finishes each block, avoiding repeated allocation during streaming.

The producer writes into available slots and publishes its progress. The consumer reads completed slots and returns them for reuse. At the end of the allocation, each index wraps to the beginning. The implementation must distinguish full from empty, for example by reserving one slot or tracking a separate count.

The diagram below follows one block through a ring shared by a receive thread and a DSP thread. The arrows represent ownership changes. Only a free block can return to the producer; a published block remains unavailable for reuse until consumption finishes.

FREE WRITING READY READING claimpublishacquire Release only after the consumer finishes using the block.

Figure 1. Reusing a block requires an ownership transfer, even when the samples are never copied.

  • WRITING belongs to the producer; READING belongs to the consumer.
  • READY means the payload and associated metadata are complete and visible.

Correct ordering matters on multicore hosts. The consumer must not see the published index before the payload is visible. Use the synchronization primitives appropriate to the implementation, such as release/acquire operations or locks. A volatile variable alone is not a complete thread-synchronization protocol.

The same block should carry its first-sample timestamp, sample count and discontinuity status. Otherwise a ring can preserve the bytes while losing their timing meaning. Multiple readers also complicate reuse: a slow recorder must not retain a block forever and prevent the real-time consumer from receiving new data.

  • Index updates transfer ownership. Publish after writing and release after reading, with the required memory ordering.
  • Keep timing metadata with the payload. Buffer reuse must not separate samples from their timestamps.

How does DMA move samples without the CPU touching them?

DMA removes the CPU from the byte-by-byte transfer, but software still prepares memory and manages ownership. The important question is when the device may access a buffer and when the CPU may reuse it.

The driver supplies descriptors that identify DMA-accessible memory and transfer lengths. Hardware processes those descriptors and reports completion. The driver can then deliver received data or reclaim transmitted buffers. Descriptor storage and sample storage serve different purposes; a ring of descriptors is not itself a ring full of IQ samples.

A device uses DMA addresses, which need not equal CPU virtual or physical addresses. An IOMMU can translate between address spaces. The driver must establish a valid mapping and keep the underlying storage alive until the transfer completes. User pages may require pinning; driver-allocated DMA memory follows the relevant allocation API.

Cache coherency and memory ordering are separate issues. Some mappings need explicit synchronization when ownership changes. Even coherent descriptor memory needs ordering so hardware does not see a valid descriptor before its address and length are ready. The operating system's DMA API defines those rules.

Completion processing can still become the limiting stage. If software does not reclaim buffers and post replacements promptly, the device can exhaust its available descriptors. The machine may have gigabytes of unused RAM while the particular DMA queue has no usable destination.

Zero-copy adds an application-side constraint: the consumer must finish before returning the buffer. Keeping a pointer after release risks reading samples from a later transfer. DMA avoids a CPU copy, but the transfer still consumes interconnect bandwidth and accesses memory.

  • DMA needs valid mappings and lifetimes. A normal application pointer is not a device address.
  • Completion does not imply unlimited reuse. Driver and application ownership must agree before the next transfer.

What does the buffer chain look like on each interface?

A buffer parameter has meaning only within its transport. Increasing a USB transfer count changes a different queue from increasing an Ethernet socket buffer, even when both settings appear in the same SDR driver.

The receive paths below show representative stages rather than every internal queue. Each row ends at the application ring, but the host reaches that ring through different hardware and software. TX data travels in the opposite direction through corresponding transmit queues.

USBEthernetPCIe FPGA / device buffersUSB controller / transfersDriver / libraryApplication ring FPGA / packetizerNetwork / NIC RX ringSocket / libraryApplication ring FPGA / DMA engineHost DMA buffersDriver / libraryApplication ring Arrows show sample flow. Backpressure and error reporting depend on the specific implementation.

Figure 2. A larger application ring cannot repair losses in an earlier transport queue.

  • USB uses pending transfers to keep the controller supplied with work.
  • Ethernet adds network and NIC queues before the socket path.
  • PCIe can DMA into host buffers, which still need timely reclamation.

UHD's USB transport uses num_recv_frames for simultaneous receive transfers and recv_frame_size for their size. For UDP sockets, recv_buff_size requests socket capacity subject to operating-system limits. The documented UDP performance notes do not give frame count the same tuning role as USB.

Packet size affects batching too. For a fixed sample rate, a larger sample payload takes longer to accumulate. Smaller packets reduce that waiting interval but increase packet-processing work. The path MTU and driver constraints limit the available choices.

A DPDK path differs from the Ethernet row above by bypassing the ordinary kernel socket path. Its packet buffers come from managed pools, and polling changes how work is serviced. It still needs available buffers and sufficient processing capacity. Always identify the active transport implementation before applying a parameter from another setup.

  • Tune the queue that is filling. Similar parameter names do not establish identical behaviour across transports.
  • Check the effective configuration. Requested sizes can be limited or adjusted by the driver and operating system.

Where does memory bandwidth actually run out?

The link rate is only the first part of the host's data movement. Samples may be copied, converted and read several times before decoding finishes, so memory traffic can substantially exceed the incoming payload rate.

Consider the same 800 MB/s stream in complex int16 format. Each complex sample occupies four bytes. Converting it to complex float32 produces eight bytes per sample, so the output becomes 1.6 GB/s. A stage sized using only the wire format will underestimate its host-side storage requirement.

A simple traffic accounting example starts with an 800 MB/s DMA write. One full copy adds an 800 MB/s read and an 800 MB/s write. Format conversion then reads 800 MB/s and writes 1.6 GB/s. These operations total 4.8 GB/s of logical data movement before later DSP reads.

This sum is not a measured DRAM bandwidth requirement. Caches may satisfy some accesses, while write allocation and other effects can add traffic. It is a way to identify expensive passes through the data. Measure the real workload before comparing it with the memory system's sustainable bandwidth.

NUMA placement can add another transfer. A NIC attached to one CPU socket may deliver buffers that a worker on another socket consumes. Keeping the device, memory allocation and processing thread appropriately located can reduce remote traffic. More worker threads do not necessarily help once they compete for the same memory channels.

Large batches reduce per-call overhead, but also increase working-set size and waiting time. Compare throughput, processing latency and memory traffic together when choosing a block size.

  • Count reads and writes, not just input bytes. Copies and sample conversion can dominate data movement.
  • Measure the active working set. Cache reuse and NUMA placement determine how logical traffic reaches physical memory.

What do overflow and underrun really tell you?

An error marker tells you that a streaming condition failed, but it may not identify the original cause. The useful next step is to connect that marker to queue occupancy, timestamps and transport statistics.

On RX, overflow means some stage could not retain all incoming data. UHD's ERROR_CODE_OVERFLOW also covers receive sequence errors; out_of_sequence distinguishes that case. So the code alone does not prove that the FPGA FIFO was the first buffer to fill.

On TX, an underrun means a send buffer emptied while data was needed. UHD reports EVENT_CODE_UNDERFLOW asynchronously. A late timestamp is different: EVENT_CODE_TIME_ERROR reports a packet whose requested time was already past. Queue capacity cannot make an expired deadline valid.

The table below separates observations from possible explanations. These are diagnostic starting points, not one-to-one fault assignments. Record several indicators around the same event before changing a buffer setting.

Observation

What to investigate

Application RX occupancy rises steadily

Sustained processing deficit, blocked storage output or a consumer that retains buffers.

RX gaps with low application occupancy

Earlier device, NIC or socket loss; insufficient posted transfers; transport interruptions.

TX underruns after occasional long processing intervals

Producer scheduling, processing-time variation and available queued sample time.

TX time errors despite queued data

Timestamp calculation, scheduling lead and samples arriving after the requested start.

Preserve the discontinuity in the output. If valid timestamps show a missing interval, do not concatenate samples as though they were contiguous. A recorder can mark the gap; a modem may need to discard affected processing state and reacquire alignment. The recovery depends on the application and streaming mode.

  • An error code is evidence, not a complete diagnosis. Correlate device, transport and application observations.
  • Keep missing time visible. Silent concatenation can turn a transport failure into misleading DSP results.

How deep should each buffer be?

The useful size comes from a delay budget, not from the amount of RAM available. Decide how long a stage may stop, and how much waiting time the application can tolerate.

For a complete RX consumer stall lasting J seconds, reserve at least R * J bytes of free space at the affected queue. Add allowance for bursts and measurement uncertainty. If the queue already contains Q bytes, its total capacity must cover Q plus that reserve.

At 800 MB/s, a 2 ms stall needs 1.6 MB of free capacity. Suppose processing subsequently runs at 1,000 MB/s while input continues at 800 MB/s. The net drain rate is 200 MB/s, so removing that backlog takes 8 ms. Another stall before recovery requires additional reserve.

Capacity is not automatically latency. An empty 100 MB ring does not delay each sample by 125 ms. Actual queued data creates waiting time; a fixed batching policy can add further delay. Under steady flow, Q / R expresses occupancy as equivalent sample time. With a stalled consumer, actual waiting can be longer.

TX sizing uses the complementary quantity: samples already available for playback. Prebuffering can protect against producer pauses, but timed transmission still has to begin at the scheduled instant. Large queues cannot justify sending old data after its useful time.

Measure occupancy high-water marks, service gaps and error counts under representative load. Repeat with all channels, sample conversion and recording enabled. Tune one stage at a time and confirm that its effective size changed. If the measured stall allowance exceeds the latency budget, reduce the stall or processing cost instead of only increasing capacity.

  • Size RX reserve and TX inventory separately. Free receive space and queued transmit data protect opposite failures.
  • Budget recovery time as well as stall time. A system needs spare throughput to recover before the next interruption.

Reference

The list below is where the device figures and the driver parameter behaviour come from. The vendor documentation is the authority for anything specific to a board or a driver. The ShareTechnote pages carry the transport, timing and converter context this page assumes.