SDR(Software Defined Radio)

 

 

 

Digital Front End

A converter runs at a rate the designer chose. A host asks for a rate the application needs. The digital front end is what reconciles those two numbers, and it also picks the channel out of the band. For years I treated it as a black box that produced whatever rate I typed. It is not a black box, and the reason some rates work well and others do not is entirely inside it. Each section below takes one stage of that block.

Executive Summary

The table is a lookup. The chain runs top to bottom on receive, and the same stages run in reverse on transmit.

Stage

What it does

Why this structure

What it costs

NCO and complex mixer

Shifts the wanted channel down to 0 Hz, so later filters can be low-pass.

A complex multiply is the cheapest exact frequency shift available in digital form.

Phase truncation spurs, at roughly 6 dB below carrier per bit of retained phase.

CIC decimator

Takes the large integer rate reduction, often by 16 to several hundred.

Integrators and combs only, so it needs no multipliers at the highest rate in the chain.

A sinc shaped passband droop, and large internal word growth.

Compensation FIR

Applies the inverse of the CIC droop, and usually decimates by two as well.

It runs after the CIC, so it works at a low rate where multipliers are affordable.

Coefficients that depend on the CIC configuration, so the pair has to be designed together.

Halfband chain

Each stage decimates by exactly two and sharpens the band edge.

Close to half the coefficients are zero, so a halfband costs about half of a general FIR.

Only powers of two, so the chain alone cannot reach an arbitrary rate.

Resampler

Reaches rates the integer chain cannot, by a rational or an arbitrary ratio.

Polyphase computes only the samples that are kept, rather than filtering then discarding.

Real cost in logic or in host CPU, which is why some requested rates are far more expensive.

What sits between the converter and the transport?

The converter delivers the whole band at a fixed rate. The application wants one channel at a rate of its own choosing. Something has to bridge those two, and that something is the digital front end.

It has two jobs. The first is selection, which means moving the wanted channel to 0 Hz and removing everything else. The second is rate reduction, which means producing fewer samples per second than the converter does.

Those two jobs are linked rather than independent. Reducing the rate is only legal once the band has been narrowed, because anything left outside the new Nyquist limit folds back in. So every decimation stage carries a filter, and the filter is the expensive part.

The arithmetic that motivates the whole block is worth stating once. An X310 class converter runs at 200 MS/s. A narrowband application may want 200 kS/s. That is a rate reduction of a thousand, and doing it in one step would need an enormous filter.

So the reduction is split into stages, and each stage uses the cheapest structure that works at its rate. Early stages run fast and must be simple. Later stages run slowly and can afford to be precise. That principle explains the whole chain, and every section below is one consequence of it.

One property connects this page to the rest of the section. The rate leaving this block is the rate the transport has to carry. The Adaptors/Connectors page uses that number in all of its arithmetic. Decimation here is what makes a 1 GbE radio possible at all.

Figure 1 draws the receive chain with the rates marked. The reduction is spread across four stages, and each one is a different structure for the reason just given.

Fast and simple first, slow and precise last ADC 200 MS/s NCO and mixer shift to 0 Hz CIC no multipliers compensation FIR undoes the droop halfband chain divide by 2 each transport 200 kS/s sample rate falls by a factor of 1000 across the block Every rate reduction needs a filter first, or the discarded band folds back in. Cost per stage falls with the rate, so precision is bought at the slow end of the chain.

Figure 1. The chain exists because one filter cannot do a thousandfold reduction affordably. Splitting it lets the fast stages be crude and the slow stages be exact, which is the whole design argument.

  • Two jobs, tightly linked : Select the channel and reduce the rate. The second is only legal after the first, because anything left outside folds back in.
  • The reduction can be enormous : A 200 MS/s converter feeding a 200 kS/s application is a factor of a thousand. No single filter does that cheaply.
  • Cost per sample follows the rate : Early stages run fast and must be simple, and later stages run slowly and can afford precision.
  • Every decimation carries a filter : The filter is the expensive part, and the decimation itself is free.
  • This block sets the transport load : The rate leaving here is the rate the host interface carries. Decimation is what makes a 1 GbE radio possible.
  • The transmit path is the same chain reversed : Interpolation replaces decimation, and the ordering of the stages turns around.

How does the NCO move a channel to baseband?

Selection has to happen before filtering, because a low-pass filter only keeps what sits near 0 Hz. So the wanted channel is moved there first. The operation is a multiplication, and the thing doing the multiplying is a numerically controlled oscillator.

The idea is the digital version of a mixer. Multiplying a signal by a complex exponential shifts its whole spectrum. The shift equals the exponential's frequency. Choosing that frequency to be minus the channel centre puts the channel at 0 Hz.

One property makes the digital version better than its analogue counterpart. A complex multiply shifts the spectrum in one direction only, so no image is produced. The analogue mixer of the RF Front End page creates an image because its two paths are never perfectly matched. Here the two paths are numbers, and numbers match exactly.

The oscillator is built from a phase accumulator. A constant is added to a register on every sample. The register overflows naturally at the end of a cycle. The accumulated value is the phase, and the frequency is set by the size of the constant.

Turning that phase into a sine and cosine is where the design choices appear. A lookup table is fast and needs memory. A CORDIC rotates a vector using only shifts and additions, and it needs no table at all. CORDIC costs one iteration per bit of accuracy, so the two approaches trade memory against logic.

Either way the phase is truncated before use, and the truncation matters. Discarding low phase bits introduces a small periodic error, and a periodic error appears as a discrete spur rather than as noise. The usual rule of thumb puts the worst spur about 6 dB below the carrier for every phase bit retained.

That rule has a practical consequence when hunting for weak signals. Spurs from the NCO sit at frequencies related to the tuning offset rather than to anything real. Changing the requested centre frequency by a small amount moves them, which is the same test the ADC and DAC page suggests for converter artefacts.

  • Shift first, then filter : A low-pass filter keeps only what sits near 0 Hz. So the channel moves there before anything is removed.
  • A complex multiply shifts one way : It produces no image, which is the advantage the digital mixer holds over the analogue one.
  • A phase accumulator is the whole oscillator : Add a constant every sample and let the register overflow, and the constant sets the frequency.
  • Lookup table or CORDIC : A table spends memory and a CORDIC spends logic, at roughly one iteration per bit of accuracy.
  • Phase truncation makes spurs : Discarding phase bits creates a periodic error, and the worst spur sits about 6 dB down per retained bit.
  • Retune slightly to identify them : NCO spurs follow the tuning offset rather than the air. A small frequency change moves them and not real signals.

Why is the first decimation stage almost always a CIC?

The first stage runs at the converter rate, which is the fastest anything in the chain will ever run. Multipliers are expensive at that rate, and a general filter needs a great many of them. The cascaded integrator comb structure exists to avoid them entirely.

Its construction is remarkably plain. Several integrators run at the input rate. The stream is then decimated, and the same number of comb sections run at the output rate. An integrator is an addition and a comb is a subtraction, so the whole filter contains no multiplication at all.

That is the entire reason it is used. A structure with no multipliers runs at the converter rate in modest logic. It can also decimate by a large factor in one step. Rate reductions of 16 to several hundred are routine.

The price is the shape of its response. A CIC has a sinc response raised to the power of the number of stages, so its passband is not flat. It droops, and the droop worsens toward the band edge.

The numbers make the trade concrete. With four stages, using only 20 percent of the output band costs about 0.6 dB at the edge. Using 50 percent of it costs about 3.6 dB. So designs deliberately use a fraction of the available output bandwidth, and accept extra decimation to get it.

A second cost is invisible from outside and matters to anyone building one. The integrators accumulate without bound, so the internal word grows. The required growth is the number of stages multiplied by the log base two of the decimation factor. Four stages decimating by 64 therefore add 24 bits, so a 16-bit input needs 40-bit internal registers.

The response has one convenient property that partly justifies the droop. Its nulls land exactly on the frequencies that would fold onto 0 Hz after decimation. So a CIC attenuates most strongly at precisely the places where aliasing would do the most damage.

  • It contains no multipliers : Integrators add and combs subtract, which is why the structure can run at the converter rate in modest logic.
  • It takes the large factor in one step : Decimation by 16 to several hundred is routine. That is what makes the rest of the chain affordable.
  • The passband droops : The response is a sinc raised to the number of stages, so the band edge is attenuated relative to the centre.
  • Use a fraction of the output band : With four stages, 20 percent of the band costs 0.6 dB and 50 percent costs 3.6 dB.
  • Internal words grow a great deal : Four stages decimating by 64 add 24 bits, so a 16-bit input needs 40-bit registers inside.
  • The nulls land where they are needed : They sit on the frequencies that would fold onto 0 Hz, so the worst aliasing is attenuated most.

What fixes the CIC droop?

The droop is predictable, and anything predictable can be inverted. So the stage after the CIC applies the opposite shape, and the pair together comes out flat. That stage is the compensation FIR, and it is designed against a specific CIC rather than in general.

Position is what makes it affordable. The compensation filter runs after the decimation, so it operates at a much lower rate than the CIC did. Multipliers that were unaffordable at 200 MS/s become ordinary at a few MS/s, which is the same principle the framing section stated.

Its response rises toward the band edge by exactly the amount the CIC fell. The CIC shape depends on the number of stages and on the decimation factor. So the coefficients depend on both. Change the decimation and the compensation has to be recomputed, which is why the two are designed as a pair.

After compensation the remaining reduction is done by halfband filters. A halfband decimates by exactly two, and it has a property that makes it unusually cheap. Close to half of its coefficients are exactly zero, so the multiply count is roughly half that of a general filter of the same length.

Cascading halfbands is therefore the standard way to cover the last part of the reduction. Each stage halves the rate and sharpens the transition. Several in a row cost far less than one long filter. The last stage in the chain usually carries the final shaping, since it runs slowest of all.

One limitation follows directly, and the next section deals with it. A halfband chain only reaches rates related by powers of two. Any other rate needs a different mechanism.

Figure 2 draws the droop and its correction. The CIC response falls toward the band edge, the compensation filter rises by the same amount, and the product is flat across the usable band.

The droop is predictable, so the next stage inverts it frequency, toward the band edge gain, dB 0 dB CIC alone, sinc to the Nth power compensation FIR, the inverse the two together edge of the usable band The compensation coefficients depend on the CIC stage count and decimation factor, so the two are designed as a pair. It runs after the decimation, where multipliers are affordable.

Figure 2. Changing the decimation factor changes the droop, so a compensation filter is never generic. This is why a radio exposes a list of supported rates rather than an arbitrary one.

  • Predictable distortion can be inverted : The compensation FIR rises by exactly the amount the CIC fell, and the pair together is flat.
  • It runs after the decimation : Multipliers that were unaffordable at the converter rate become ordinary a hundred times slower.
  • The pair is designed together : Coefficients depend on the CIC stage count and decimation, so changing one forces a redesign of the other.
  • A halfband costs about half a FIR : Close to half its coefficients are exactly zero, so cascading them is cheaper than one long filter.
  • Each halfband divides by exactly two : Cascading covers the last part of the reduction while sharpening the band edge at each step.
  • Powers of two are the limitation : A halfband chain cannot reach an arbitrary rate, which is what the next section has to solve.

How do you get a sample rate that is not a simple ratio?

The chain so far reaches the converter rate divided by an integer, and often only by convenient integers. Applications do not respect that. A cellular standard wants 30.72 MS/s, an audio path wants 48 kS/s, and neither divides neatly into a converter clock.

The general answer is rational resampling. Upsample by an integer, filter, then downsample by another integer, and the ratio of the two reaches any rational rate. Getting 48 kS/s from 44.1 kS/s means upsampling by 160 and downsampling by 147.

Done literally that is enormously wasteful. The intermediate rate is 160 times the input, and almost every sample computed at that rate is then discarded. Nobody implements it that way.

The polyphase decomposition removes the waste. It reorganises the filter, so only the samples actually kept are ever computed. The result is arithmetically identical and costs a small fraction as much, which is why polyphase appears in every serious resampler.

Some ratios are not rational at all, or the integers are impractically large. Then an arbitrary resampler is used. It interpolates between existing samples to estimate a value at a new instant. The Farrow structure is the common implementation, and it fits a low-order polynomial through nearby samples.

Arbitrary resampling is more expensive and less exact than a rational design. The polynomial is an approximation, so it adds a small error that depends on the fractional delay. That is a real cost, and it is the reason a radio prefers to give you a rate its integer chain can reach.

This explains a behaviour that otherwise looks arbitrary. A radio has a master clock, and the rates reachable by integer decimation from it are cheap. An X310 runs from a master clock of 200 MHz or 184.32 MHz. A rate that divides one of those cleanly costs nothing extra. Requesting something else forces resampling, either in the FPGA or on the host CPU.

The practical advice is short. Ask what master clock the radio is using, then choose a rate that divides it. A B200 or B210 allows the master clock itself to be changed, which makes far more rates reachable without any resampling at all.

  • Rational resampling reaches any ratio : Upsample by one integer and downsample by another, so 44.1 to 48 kS/s is 160 up and 147 down.
  • The literal implementation is wasteful : It computes 160 times as many samples as it keeps, so nobody builds it that way.
  • Polyphase computes only what is kept : The result is arithmetically identical and costs a small fraction, which is why every real resampler uses it.
  • Arbitrary ratios need interpolation : A Farrow structure fits a low-order polynomial between samples, and the approximation adds a small error.
  • The master clock decides what is cheap : An X310 runs at 200 MHz or 184.32 MHz, and rates dividing those need no resampling.
  • Some radios let you move the master clock : A B200 or B210 does. That turns an awkward rate into an integer one and removes the cost.

Where does fixed point arithmetic bite?

Everything in this block is fixed point, because floating point in an FPGA costs far more logic than it is worth here. Fixed point is exact and it is unforgiving. Two failure modes come from it, and they look completely different on a spectrum display.

The first is word growth. Adding two numbers can need one more bit, and a filter adds many numbers. The CIC section gave the sharpest case, where four stages decimating by 64 add 24 bits to the word.

Somewhere those extra bits have to be discarded, since the transport carries 16-bit samples. Discarding the low bits adds quantisation noise. Discarding the high bits risks losing the signal entirely. Where the cut is made is a design decision called scaling, and it is the main thing a fixed point designer spends time on.

The second failure mode is what happens when a value exceeds the range. Two behaviours are possible and they are not equally bad. Wraparound turns a large positive number into a large negative one, which is a catastrophic discontinuity. Saturation clips at the maximum, which is merely distortion.

The difference on a display is unmistakable once you know it. Saturation raises the noise floor and adds harmonics, roughly like an overdriven amplifier. Wraparound produces broadband hash across the entire band, because a discontinuity contains every frequency.

Well designed hardware saturates rather than wraps, and it is worth confirming rather than assuming. Sudden full-band noise on strong signals usually means something is wrapping. The fix is a scaling change rather than a filter change.

One habit avoids most of this on the host side. Keep some headroom in the digital chain rather than filling the word, exactly as the transmit path leaves headroom for peaks. A signal sitting at full scale in the FPGA has nowhere to go when a filter overshoots, and filters do overshoot on transients.

  • Filters grow the word length : Every addition can need another bit, and a four-stage CIC decimating by 64 adds 24 of them.
  • Scaling decides where bits are discarded : Cutting the low bits adds noise and cutting the high bits loses the signal. That choice is the core design decision.
  • Wraparound and saturation are not equivalent : One turns a large positive into a large negative, and the other simply clips at the limit.
  • The two look different on a display : Saturation adds harmonics and raises the floor, and wraparound produces hash across the whole band.
  • Full-band hash means wrapping : A discontinuity contains every frequency, so sudden broadband noise on strong signals points at scaling.
  • Leave digital headroom too : Filters overshoot on transients, so a signal already at full scale inside the chain has nowhere to go.

What does the transmit side do differently?

The transmit chain is the receive chain reversed, and the reversal is not merely cosmetic. Interpolation replaces decimation, and the stages run in the opposite order. What changes is which artefacts have to be suppressed, and where.

Interpolation means inserting samples between existing ones. Inserting zeros raises the rate. It also creates copies of the spectrum at multiples of the old rate. A filter after the insertion removes those copies, and the filter is what actually performs the interpolation.

So the transmit filter has a different job from its receive counterpart. On receive the filter prevents unwanted energy from folding in. On transmit the filter prevents unwanted energy from being emitted. The second is stricter, because the consequence leaves the building.

The cheap stages sit at the fast end here too, which reverses their position. A halfband chain runs first at the low rate, and the CIC interpolator runs last at the converter rate. The compensation filter therefore comes before the CIC rather than after it, since it has to pre-distort the signal ahead of the droop.

The droop itself works the same way and gets corrected the same way. A CIC interpolator imposes the same sinc shape, so the compensation filter applies the inverse in advance. The ADC and DAC page adds a second droop from the converter hold, and a well built transmitter corrects both together.

Headroom becomes a harder constraint on transmit. Every interpolation filter overshoots on transitions. So the peak level after filtering exceeds the peak before it. A signal at full scale entering the chain will clip inside it, and the clipping happens where no meter is watching.

Figure 3 draws both directions together, so the mirror is visible in one picture.

Receive : rate falls left to right NCO and mixer CIC decimate compensation FIR halfband chain host Transmit : the same stages, reversed halfband chain compensation FIR CIC interpolate NCO and mixer DAC Compensation moves ahead of the CIC, because the droop now has to be pre-distorted. On receive a filter stops energy folding in. On transmit it stops energy being emitted, which is the stricter job. Interpolation filters overshoot, so the peak grows inside the chain where no meter is watching.

Figure 3. The mirror is exact except for one thing. Compensation moves to the far side of the CIC, because a droop that has not happened yet has to be pre-distorted rather than corrected.

  • Interpolation creates spectral copies : Inserting zeros raises the rate and replicates the spectrum, and the following filter is what removes the copies.
  • The filter does the interpolating : Zero insertion only changes the rate, so all the actual work happens in the filter after it.
  • Transmit filtering is the stricter job : On receive an imperfect filter costs you accuracy, and on transmit it emits energy other people receive.
  • Compensation moves ahead of the CIC : The droop has not happened yet, so the signal is pre-distorted rather than corrected afterwards.
  • Two droops need correcting on transmit : The CIC contributes one and the converter hold contributes another, and a good design handles both.
  • Peaks grow inside the chain : Interpolation filters overshoot on transitions, so a full-scale input clips at a stage nobody is monitoring.

What does the digital front end look like on real hardware?

The structures above are generic, and each board implements a particular subset of them. What differs is how much is fixed and how much you can configure. That difference decides which sample rates and which channel shapes are actually available.

An AD9361 puts the whole chain inside the transceiver. Its converters feed configurable decimation filters. Those feed a 128-tap programmable FIR before the 12-bit output leaves the chip. The programmable stage is the interesting one, because it lets the channel filter be shaped by the user rather than by the vendor.

That programmability is why a Pluto or a B210 can be configured for a wide range of bandwidths. It also explains a common surprise. The 128 taps have to cover the whole channel shaping requirement, so a very sharp filter at a wide bandwidth is simply not available.

A USRP with an FPGA motherboard puts the chain in the FPGA instead. The X310 does its NCO tuning and its decimation there, which is why the chain can be modified. RFNoC exists precisely to let a user insert their own processing into that path rather than only choosing parameters for it.

The trade against the AD9361 approach is worth naming. An FPGA chain is far more flexible and it consumes FPGA resources and development effort. A transceiver chip chain costs nothing to use and stops exactly where the vendor stopped it.

An RTL-SDR sits at the far end of that scale. The RTL2832U carries a fixed decimation chain designed for digital television, and a user cannot reconfigure it. That is why only certain sample rates work reliably, and why 2.4 MS/s is the rate most guides recommend.

One pattern runs across all three, and it is a useful thing to check before buying. The more of the digital front end that is fixed in silicon, the fewer sample rates and channel shapes are available. Flexibility in this block is bought with FPGA logic. The boards offering it are the ones carrying an FPGA large enough to spare.

  • An AD9361 holds the chain on the chip : Configurable decimators feed a 128-tap programmable FIR, so the channel filter is shaped by the user.
  • 128 taps is a real budget : Sharp filtering at wide bandwidth exceeds it. That is the usual surprise when a filter refuses to meet a mask.
  • An FPGA chain can be modified : An X310 runs its NCO and decimation in the FPGA. RFNoC exists so a user can insert processing there.
  • Flexibility costs logic and effort : A transceiver chip chain is free to use and stops where the vendor stopped, and an FPGA chain does neither.
  • An RTL-SDR chain is fixed : The RTL2832U decimator was designed for television, which is why only some rates work and 2.4 MS/s is recommended.
  • Fixed silicon means fewer rates : The proportion of this block that is configurable predicts how many sample rates and channel shapes a board offers.

Reference

The list below is where the device behaviour comes from. The vendor documentation is the authority for anything specific to a part, and the ShareTechnote pages carry the stages on either side of this block.

  • AD9361 Data Sheet : Analog Devices - the decimation filter chain and the 128-tap programmable FIR ahead of the 12-bit output.
  • AD9361 Reference Wiki : Analog Devices - filter configuration, achievable bandwidths and the constraints on the programmable stage.
  • UHD and USRP Manual : Ettus Research - master clock rates, supported decimation factors and where resampling happens.
  • X300/X310 : Ettus Research - the 200 MHz and 184.32 MHz master clock options behind the reachable sample rates.
  • RFNoC : Ettus Research - inserting user processing into the FPGA signal chain rather than only configuring it.
  • ShareTechnote - ADC and DAC : The converter rate this block reduces, and the processing gain that decimation delivers.
  • ShareTechnote - RF Front End : The analogue mixer this block replaces in the digital domain, and its image problem.
  • ShareTechnote - Adaptors/Connectors : The transport that carries whatever rate leaves this block.