SDR(Software Defined Radio)

 

 

 

Where the Processing Runs

A buffer buys us some time, but it does not process any samples. Something still has to keep up with the incoming stream. I used to think this was mainly a software question, with hardware somewhere in the background. But the hardware limit comes first. One number gives us a first estimate of whether the work belongs on a CPU, GPU or FPGA. Let's look at that number and see what it means for each one.

Executive Summary

Following table compares the three options. Before reading it, keep throughput and latency separate. Latency is the time from a sample's arrival until its result is available.

Target

What it is good at

What determines latency

Cost of changing it

Real example

CPU

Complex logic, branches and functions that change often.

Algorithm, block size, memory access and scheduling. Measure the complete path and its variation.

Software rebuild and validation; duration depends on the project and change.

VOLK provides SIMD kernels; srsRAN Project implements a software 5G CU/DU.

GPU

Applying the same operation to many items in parallel.

Batch accumulation, transfers, launch, execution and completion. No universal latency applies.

Kernel rebuild plus integration, numerical validation and timing measurements.

NVIDIA Aerial runs 5G layer 1 on CUDA.

FPGA

Dedicated streaming pipelines with timing controlled by the implemented design.

Pipeline depth, block accumulation, memory access and stalls. Determinism requires bounded behaviour.

Hardware changes require synthesis, implementation and validation; exposed parameters may change at runtime.

RFNoC blocks on supported USRPs. An O-RU functional role does not mandate FPGA implementation.

The deciding number

A first estimate: divide core clock rate by the aggregate sample rate assigned to that core.

At 200 Msps, a 3 GHz core has only 15 cycles per sample.

A higher core clock or lower assigned sample rate increases the ratio; algorithm and memory optimizations reduce the work needed.

Combine this estimate with measured workload, memory traffic and deadline constraints.

How much time do you actually have per sample?

At 200 MS/s, another sample arrives every 5 ns. That looks impossible for software until you separate the arrival interval from the processing schedule: software normally handles blocks, with several samples in progress together.

A 3 GHz core provides 3 billion clock cycles per second. Dividing by 200 million complex samples gives 15 cycles per complex sample for one stream. Four equal streams sharing that core reduce the nominal budget to 3.75 cycles per sample. This assumes the core spends all its time on those streams.

Cycles are not instructions or arithmetic operations. A CPU can issue multiple instructions, and a SIMD instruction can operate on several values. Conversely, a cache miss or dependency can consume time without completing useful arithmetic. The ratio is a first capacity estimate, not a rule that every sample must finish within 15 cycles.

Now consider a block of 4,096 samples at 200 MS/s. A new block arrives every 20.48 microseconds. A single worker must sustain that average service interval, with enough spare capacity to recover from interruptions. Its result may also have a separate deadline that is shorter or longer than one block interval.

The sample definition matters too. A complex multiply includes several real operations, while decoding work often scales with code blocks and iterations. Channel count, bandwidth, modulation and traffic load can therefore change the workload differently. A sample-rate figure alone cannot predict the entire modem's compute demand.

  • State the denominator. Distinguish real samples, complex samples, channels and processing blocks.
  • Use cycles per sample as an initial estimate. Measure service time and deadline margin for the complete workload.

What makes a CPU fast enough at all?

A CPU rarely processes a high-rate stream as one isolated sample at a time. It becomes useful when the implementation exposes parallel arithmetic, reuses cached data and avoids unnecessary work between blocks.

SIMD applies one instruction to several values held in a vector register. A 256-bit register can hold eight 32-bit floating-point values, or four interleaved complex values. That does not mean one instruction completes four complex multiplies. The operation still needs the appropriate arithmetic, rearrangement and loads.

Libraries such as VOLK provide optimized kernels for supported SIMD architectures. They help with operations such as vector multiplication and dot products, but they do not eliminate the cost of moving data. Check which implementation is actually selected on the target machine, and benchmark realistic block lengths and alignment.

Several cores can process independent channels or blocks. Dependencies limit that freedom: a later stage may need the previous stage's result, and stateful filters must preserve their history. Splitting work also introduces queues and synchronization. More threads can reduce performance when they compete for memory bandwidth or frequently exchange small blocks.

Keep the hot processing path predictable. Allocate reusable buffers before streaming, avoid blocking file writes in the receive worker, and separate detailed logging from deadline-sensitive work. CPU affinity can reduce migration, but it does not remove interrupts, cache contention or every source of scheduling delay.

The useful performance number is sustained processing time with the actual sample format and channel count. Include format conversion and queue handling. A fast arithmetic kernel does not prove that the full receive-to-result path meets its deadline.

  • Vectorize useful work. SIMD width helps only when the algorithm and memory layout can use it.
  • Parallelism has a coordination cost. Measure queues and memory traffic as well as CPU utilization.

When does a GPU help, and when does it hurt?

A GPU can finish a large parallel calculation quickly and still deliver the result too late. The decision depends on the entire offload path, including when enough input becomes available and where the output is needed.

FFT batches, matrix operations and parallel decoding workloads can expose enough independent work to use a GPU effectively. Small operations with frequent host decisions may not. For a discrete GPU, include data transfer, launch and completion handling alongside kernel execution. Integrated or direct-I/O arrangements change these costs, so measure the actual topology.

A simple non-overlapped comparison is T_offload = T_input + T_launch + T_kernel + T_output + T_completion. Suppose those terms total 80 microseconds while the CPU takes 60 microseconds. A 10-microsecond GPU kernel has not accelerated the complete operation. These are illustrative values, not hardware measurements.

Keeping successive stages on the GPU can remove repeated transfers of intermediate results. Larger batches can also amortize launch overhead, but waiting for a batch adds latency. A batch of 4,096 samples at 30.72 MS/s takes about 133.3 microseconds to accumulate from an initially empty stream.

Transfer and compute can overlap for independent work when the hardware and software support it. CUDA streams and suitable host memory are part of that setup. Overlap improves sustained throughput; it does not remove the dependency between an individual block's input transfer and its computation.

NVIDIA Aerial provides a GPU-accelerated radio processing example. It should not be interpreted as proof that any SDR block benefits from offload. Benchmark both a sustained loaded stream and small, deadline-sensitive work before choosing the partition.

  • Time the complete offload. Kernel duration alone excludes several costs visible to the radio application.
  • Reuse device-resident data. Avoid moving intermediate results back to the CPU without a processing reason.

What does an FPGA give that neither of the others can?

An FPGA lets the designer build a processing path that follows the sample stream directly. Its advantage comes from dedicated pipelines and interfaces, but the resulting timing still depends on the implemented design.

A pipeline divides an operation across registered stages. Suppose a design runs at 250 MHz, accepts one sample each clock and has 20 cycles of latency. It can accept 250 MS/s while each sample spends 80 ns in the pipeline. These are two different properties: acceptance interval and result latency.

If the same design accepts a sample only every four clocks, its maximum input rate becomes 62.5 MS/s. Replicating lanes can raise aggregate throughput, provided the memory and interfaces supply them. A high clock frequency alone does not establish a one-sample-per-cycle implementation.

Filtering, frequency translation and decimation fit naturally near the converter. Reducing the rate there also reduces later transport and host work. However, a decimation filter has its own group delay, and a block FFT must collect input before producing its result. FPGA processing is not universally a nanosecond operation.

Fixed-point precision is another design decision. Word lengths, scaling, rounding and saturation affect both resource use and signal quality. Compare the implemented output with a trusted reference across the expected input range, including strong signals and accumulated filter gain.

RFNoC allows processing blocks to be connected within supported USRP FPGA images. Changing the hardware graph can require rebuilding the image, while exposed parameters may be adjustable at runtime. Deterministic timing requires bounded stalls, controlled arbitration and successful timing closure; the FPGA label alone does not guarantee it.

  • Check both initiation interval and latency. A deep pipeline can sustain a high sample rate.
  • Verify numerical behaviour and timing. Resource fit does not prove signal fidelity or deadline compliance.

Why is latency a different question from throughput?

A processor can keep up with every arriving block while delivering each result after the radio needs it. Throughput asks how much work finishes per second; latency asks how long one result takes.

For an end-to-end deadline, include acquisition, queueing, transport, processing and delivery. Define the start and finish points before measuring. Timing only the function call can exclude the period when samples were waiting in a driver or application queue.

The diagram below uses an illustrative 500-microsecond deadline. Processing occupies 180 microseconds, but the complete path takes 380 microseconds. The remaining 120 microseconds covers variation; it is not additional compute time that can be spent without considering the other stages.

From first input sample to result available at its consumer Acquisition: 100 us Queue / I/O: 80 us Processing: 180 us Margin: 120 us 0100180360380500 us Delivery: 20 us Deadline

Figure 1. Processing time consumes only part of the end-to-end deadline.

  • The queue interval includes waiting and data movement before processing in this example.
  • Delivery finishes at 380 microseconds; the deadline remains fixed at 500 microseconds.

In a pipeline, different blocks can occupy different stages simultaneously. The slowest sustained stage limits throughput, while a block's latency includes all dependent stages. Adding a buffer can prevent loss during a short interruption, but occupied buffers also add waiting time.

Measure the distribution, not just the mean. Report percentiles, observed maximum, measurement duration and deadline-miss count under representative load. A measured maximum is evidence from that test, not a proof of the worst possible delay. Radio scheduling offsets and transport windows determine the actual deadline; a slot duration is not automatically the available compute budget.

  • Define the timing endpoints. Sample arrival, kernel completion and usable output are different events.
  • Reserve margin for variation. Average throughput does not establish deadline reliability.

How do you partition a real system?

The choice is usually not one processor for the whole radio. A useful partition keeps high-rate work near its input, places parallel work where it can run efficiently, and leaves room for control decisions.

The diagram below shows one possible receive chain. The FPGA reduces the sample rate before transport. The CPU performs control and scheduling, while a GPU is optional for a sufficiently large parallel workload. This is an example architecture, not a required split for every SDR.

ADC samples FPGA: filter / decimate200 MS/s to 25 MS/s CPU: processing / controlSchedule work and results Optional GPU stagesKeep intermediates here TransportOffloadResults At 4 bytes per complex sample: 800 MB/s before decimation, 100 MB/s after.

Figure 2. Early rate reduction lowers the work and data movement required after the FPGA.

  • The example assumes one complex stream and unchanged four-byte sample representation.
  • The GPU boundary adds transfer and scheduling costs that must fit the deadline.

Decimation by eight changes 200 MS/s into 25 MS/s. At four bytes per complex sample, payload falls from 800 MB/s to 100 MB/s. This is useful only if the filter preserves the wanted bandwidth and suppresses aliasing. Reducing the rate is a signal-processing decision, not a free performance setting.

At each boundary, specify sample format, scaling, block size, timestamp meaning and buffer ownership. Include discontinuity reporting. A fast accelerator is difficult to integrate if its result cannot be associated with the correct channel or radio time.

Keep frequently changing control logic on a programmable processor when that fits the timing requirement. Move stable high-rate operations only when the saved compute and transport costs justify integration and verification. For an O-RAN functional split, the interface defines responsibilities and exchanged data; it does not mandate a particular semiconductor technology.

  • Measure boundary costs before moving a block. Data format conversion and ownership transfer are part of the partition.
  • Preserve the signal and timing contract. Faster processing is useful only when the result remains correct and arrives in time.

What actually goes wrong on a real machine?

A short benchmark can pass while a live radio misses deadlines. Real operation adds traffic variation, competing work, temperature changes and transport interruptions, so the first failure may occur outside the arithmetic kernel.

Measure the chain in stages: input delivery, queue wait, processing and output delivery. Preserve radio timestamps for sample timing and use an appropriate host clock for elapsed host time. Do not subtract timestamps from unrelated clock domains without establishing their relationship.

The table below connects common observations to useful checks. None identifies a unique cause by itself. Correlate processing traces with queue occupancy, transport errors and the workload active at the time.

Observation

What to check

One worker stays busy while total CPU usage is low

Serial dependencies, a single saturated core and work distribution.

Processing slows as channel count increases

Memory traffic, cache working set, NUMA placement and thread contention.

GPU kernels are fast but results are late

Batch wait, transfers, launch queues and host synchronization.

FPGA output stalls intermittently

Downstream backpressure, FIFO occupancy, shared memory and arbitration.

Performance falls during a long run

Clock frequency, temperature, power limits and accumulating backlog.

Warm up the system before measuring steady state, but also test startup if its deadline matters. Include all intended channels, modulation modes and background tasks. Decoder iteration counts and traffic load can change work even when the sample rate stays constant.

Validate signal quality after optimization. SIMD changes, reduced precision and different arithmetic ordering can affect results. Compare against reference vectors and the application's error limits. Then repeat sustained timing measurements; correct output and timely output are separate acceptance conditions.

  • Profile the complete chain under load. Low average utilization can hide a saturated stage or long scheduling gaps.
  • Verify performance and numerical correctness together. A speedup must preserve the radio function it accelerates.

Reference

Following references cover the platform behaviour discussed here. For framework-specific details, use the project's own documentation. The linked ShareTechnote pages explain the stages that deliver samples to the processing chain.

  • VOLK : Vector Optimized Library of Kernels - hand-written SIMD kernels with run-time selection of the best implementation.
  • GNU Radio Manual : GNU Radio - the block scheduler, buffer handling and how VOLK is used inside it.
  • NVIDIA Aerial SDK : NVIDIA - GPU-accelerated radio processing; actual timing depends on the deployed pipeline.
  • RFNoC : Ettus Research - inserting user processing blocks into the USRP FPGA chain.
  • srsRAN Project : Software Radio Systems - the software 5G CU/DU. The separate srsRAN 4G project covers the earlier LTE stack.
  • CUDA C++ Best Practices Guide : Host/device work partitioning, transfer overhead, device-resident intermediates and transfer/compute overlap. Relevant sections were consulted, not the complete guide.
  • srsRAN Project Documentation : The 5G CU/DU implementation and its software processing context.
  • ShareTechnote - NVIDIA/Aerial : The GPU accelerated RAN platform referenced in the GPU section.
  • ShareTechnote - Memory and Buffering : The buffers this processing consumes, and the bandwidth and NUMA constraints around them.
  • ShareTechnote - Digital Front End : The decimation that decides how large the per-sample budget is.
  • ShareTechnote - Adaptors/Connectors : Host transports and the Open RAN 7.2 split.