The physical layer makes no decisions. It executes what the scheduler asked for, and reports what it received. That sounds like the simple part of the stack. It consumes most of the machine, and it is the layer that decides whether the whole system meets its deadline. Each section below takes one part of it.
- Executive Summary
- What does the PHY have to do in each direction?
- How does a receiver find the cell in the first place?
- What does the uplink receive chain actually compute?
- Why is the decoder the expensive part?
- What is different about the transmit chain?
- How is the PHY parallelised?
- Where does numerical precision actually matter?
- Reference
Executive Summary
The table is a lookup. The cost column is a rough share of the physical layer budget, not an absolute figure.
Stage |
What it does |
Cost |
Is the cost predictable? |
How it parallelises |
|---|---|---|---|---|
Transform and extraction |
FFT per symbol, then pull out the resource elements that belong to each device. |
Small, and it grows with bandwidth. |
Completely. It depends only on the configuration. |
Per symbol, or per antenna. |
Channel estimation |
Compare received reference symbols against the known ones, then interpolate. |
Moderate, and it grows with antenna count. |
Yes, for a given configuration. |
Per device, and per antenna pair. |
Equalisation and demodulation |
Undo the channel, then turn symbols into soft bit values. |
Moderate, and it rises sharply with MIMO layers. |
Yes, apart from the layer count. |
Per resource element, which is very fine grained. |
Decoding |
LDPC for data, and Polar for control. Iterative, with early termination. |
The largest single share, often most of the uplink budget. |
No. A poor channel needs more iterations. |
Per code block, which is the natural unit. |
The whole transmit chain |
Encode, rate match, scramble, modulate, map, precode and transform. |
Smaller than the receive chain. |
Completely. Nothing in it is iterative. |
Per code block, then per symbol. |
What does the PHY have to do in each direction?
The two directions look symmetric on a block diagram. They are not symmetric in software. One of them knows what it is looking for, and the other has to work it out.
The transmit side is a translation. A transport block arrives from the MAC with a full instruction attached. The modulation, the resource allocation and the coding rate are all specified. Nothing has to be estimated.
The receive side has to recover a signal that passed through an unknown channel. It knows what was scheduled, so it knows where to look. It does not know what the channel did to the signal on the way.
That asymmetry explains most of the cost difference. Estimating the channel, inverting it, and decoding through the residual errors is far more work than generating a clean signal. In practice the receive chain dominates the budget.
One further difference matters for the real-time loop. Transmit cost is fixed by the configuration, so it can be budgeted exactly. Receive cost depends on the channel, so it can only be bounded.
The real-time loop page describes the consequence. A stack sized for average receive cost misses deadlines exactly when the channel degrades. That is the moment the system is under most pressure anyway.
There is a third job that belongs to neither direction. The physical layer measures things and reports them upward. Signal strength, timing advance, and the success or failure of each decode all become inputs to the scheduler.
The two directions are not symmetric : Transmit translates a full instruction, and receive recovers a signal that passed through an unknown channel.Transmit estimates nothing : Modulation, allocation and coding rate all arrive with the transport block from the MAC.Receive dominates the budget : Estimating a channel, inverting it and decoding through the residue is much more work than generating a clean signal.Transmit cost can be budgeted, receive cost only bounded : One depends on configuration, and the other depends on the channel.The worst case arrives at the worst moment : A stack sized for average receive cost misses deadlines exactly as the channel degrades.Measurement is the third job : Signal strength, timing advance and decode outcomes all flow upward as scheduler inputs.
How does a receiver find the cell in the first place?
Every other operation in the physical layer knows where to look. Cell search does not. It is the only genuinely blind operation in the stack, and its structure is different from everything else.
The device begins with almost nothing. It knows a frequency range to try. It does not know the exact carrier, the timing, the cell identity, or whether a cell is present at all.
The primary synchronization signal solves the first problem. It is one of a small number of known sequences, so the device correlates the received samples against each candidate. A correlation peak gives symbol timing and part of the cell identity together.
Correlation is used because it works without synchronization, which is the same argument the RF Front End page makes about preamble design. A matched filter finds a known sequence in noise without needing to know when it starts.
The secondary signal completes the identity and resolves the frame boundary. Once both are found the device knows which cell it has found and where the frame begins. Only then can it read anything.
What it reads first is the broadcast channel, which carries the master information block. That is a very small message, protected heavily, and it points at where the rest of the system information lives. The device follows that pointer to system information block one.
From a software point of view the interesting property is the cost profile. Cell search is expensive and it happens rarely. Every other operation is cheaper and happens every slot.
So cell search is normally not on the real-time path at all. It runs once during acquisition, often on a separate thread, and hands over to the steady-state loop when it succeeds. Treating it as part of the slot budget is a design error that costs a great deal of headroom.
Cell search is the only blind operation : Everything else in the physical layer knows where to look, because the scheduler told it.Correlation works without synchronization : A matched filter finds a known sequence in noise without needing to know when it starts.The primary signal gives timing and identity together : One correlation peak resolves symbol timing and part of the cell identity at the same time.Then the device follows a chain : Secondary signal, broadcast channel, master information block, and a pointer to system information.The cost profile is the opposite of everything else : Expensive and rare, against cheap and every slot.Keep it off the real-time path : Running acquisition inside the slot budget wastes headroom that the steady-state loop needs.
What does the uplink receive chain actually compute?
This is where most of the machine goes, so it is worth following in order. Each stage feeds the next, and each one has a different cost behaviour. Knowing which is which is what makes optimisation productive rather than random.
The chain starts with the transform. The cyclic prefix is removed and an FFT converts each symbol to the frequency domain. After that the signal is a grid of resource elements rather than a stream of samples.
Extraction comes next and it is bookkeeping rather than arithmetic. The scheduler allocated specific resource elements to specific devices. The physical layer pulls out the ones belonging to each transmission it expects.
Channel estimation is the first real computation. Known reference symbols were transmitted at known positions, so comparing what arrived against what was sent gives the channel at those points. Interpolating between them gives the channel everywhere else.
The quality of that interpolation matters more than most people expect. A channel that varies quickly in frequency needs denser references. A channel that varies quickly in time, because the device is moving, needs them spread differently.
Equalisation then inverts the estimated channel. With one antenna that is a division per resource element. With several antennas and several layers it becomes a small matrix operation per resource element, and the cost rises quickly with the layer count.
Demodulation converts each equalised symbol into soft bit values. The word soft is doing important work there. A hard decision would discard the confidence information that the decoder needs, and the loss is worth several decibels.
The remaining stages are descrambling, rate dematching and decoding. Descrambling and dematching are cheap rearrangements. Decoding is the expensive one, and it gets its own section.
Figure 1 shows the chain with the cost behaviour marked on each stage.
Figure 1. The chain is easy to budget until the last box. That single stage is why uplink cost has to be bounded rather than calculated, and why a busy cell on a poor channel is the case that breaks a deadline.
The transform turns samples into a grid : After the cyclic prefix and the FFT, the signal is resource elements rather than a sample stream.Extraction is bookkeeping : The scheduler already decided which elements belong to which device, so this stage only collects them.Estimation compares known against received : Reference symbols at known positions give the channel there, and interpolation gives it everywhere else.Equalisation cost rises with layers : One antenna is a division per element, and several layers make it a small matrix operation per element.Soft values are worth several decibels : A hard decision discards the confidence information the decoder depends on.Only the decoder is unpredictable : Every stage before it can be budgeted once bandwidth and antenna count are fixed.
Why is the decoder the expensive part?
The decoder usually consumes more of the physical layer budget than everything else together. The reason is structural rather than a matter of implementation quality. It is the only stage that repeats itself.
NR uses LDPC codes for data and Polar codes for control. LDPC decoding works by passing confidence values around a graph. Each pass improves the estimate, and the decoder repeats until the result is consistent.
How many passes are needed depends entirely on the channel. A strong signal converges in two or three iterations. A weak one may need fifteen or twenty, and sometimes it never converges at all.
Early termination is what makes this affordable. After each iteration the decoder checks whether the codeword is now valid. If it is, the decoder stops immediately rather than running its configured maximum.
So the average cost is far below the worst case, and the worst case is what the deadline must survive. A stack that budgets for the average works beautifully in a laboratory with a strong signal. It fails on the day a device moves to the cell edge.
The maximum iteration count is therefore a real design parameter rather than a detail. Raising it improves the error rate slightly and raises the worst-case cost proportionally. Lowering it does the opposite, and it is the usual lever when a deadline is being missed.
Code block segmentation is what makes the decoder tractable at all. A large transport block is split into code blocks of bounded size, each with its own check sequence. That bound is what keeps the decoder's working set small and its cost predictable per block.
Segmentation also creates the parallelism the next section relies on. Each code block decodes independently of every other one. That is the single most useful property the physical layer has.
Figure 2 shows the iteration count against signal quality, with the early termination effect visible.
Figure 2. Early termination is why the average cost is affordable and the worst case is not. Lowering the configured maximum is the standard response to a missed deadline, and it costs a little error rate at the cell edge.
It is the only stage that repeats : LDPC decoding passes confidence values around a graph until the result becomes consistent.The channel decides the iteration count : A strong signal converges in two or three passes, and a weak one may need twenty.Early termination makes it affordable : The decoder checks validity after every pass and stops as soon as the codeword is correct.Budget the maximum, not the average : A stack sized for the average works in a laboratory and fails when a device reaches the cell edge.Maximum iterations is a real design parameter : Lowering it is the standard response to a missed deadline, at a small cost in error rate.Segmentation bounds the problem : Splitting a transport block into code blocks keeps the working set small and the per-block cost predictable.Segmentation also creates the parallelism : Each code block decodes independently, which is the most useful property the physical layer has.
What is different about the transmit chain?
The transmit chain is the receive chain reversed, with one property that changes everything about how it is engineered. Nothing in it is iterative. Every stage costs exactly what it costs.
The sequence is straightforward. A transport block gets a check sequence, is segmented into code blocks, and each block is encoded. Rate matching selects which coded bits are actually sent, which is what turns a fixed-rate code into any rate the scheduler asked for.
Rate matching deserves a note because it does more than it appears to. The circular buffer it selects from is also what makes incremental redundancy work. A retransmission sends different coded bits rather than the same ones, which is why HARQ combining gains so much.
After that come scrambling, modulation, layer mapping and precoding. Then the resource elements are placed in the grid, an inverse transform produces time-domain samples, and the cyclic prefix is added.
Predictability is the property worth exploiting. Because the cost is fixed by configuration, the transmit chain can be scheduled tightly against its deadline. There is no need for the safety margin that the receive path requires.
It also means transmit work can be done early. A slot can be generated as soon as the scheduler has decided it, rather than at the last moment. That flexibility is what the lead in the real-time loop uses.
One asymmetry catches people during optimisation. Effort spent making the transmit chain faster rarely helps, because it was not the stage missing the deadline. The receive chain is where the variance lives, so that is where the headroom has to come from.
The exception is a cell with many devices and little traffic. Then the downlink control channel dominates, since every scheduled device needs its own control message. That cost scales with device count rather than with throughput, which is the same distinction the real-time loop page draws.
Nothing in the transmit chain iterates : Every stage costs exactly what the configuration says, so the whole path is predictable.Rate matching creates the rate : Selecting which coded bits are sent turns one fixed-rate code into whatever rate the scheduler chose.The same mechanism gives incremental redundancy : A retransmission sends different coded bits, which is where the HARQ combining gain comes from.Predictable cost allows a tight schedule : The transmit path needs none of the safety margin the receive path requires.Transmit work can be done early : A slot can be built as soon as it is decided, which is what the pipeline lead exploits.Optimising transmit rarely helps : The variance lives on the receive side, so that is where headroom has to be found.Many idle devices are the exception : Control channel cost scales with the number of scheduled devices rather than with throughput.
How is the PHY parallelised?
The physical layer is the only part of the stack that parallelises easily. Several natural divisions exist, and choosing the right one decides how well it scales. The wrong choice produces threads that spend their time waiting for each other.
The finest division is per resource element. Equalisation and demodulation genuinely are independent at that level. It is far too fine for threads, and it is exactly the right granularity for vector instructions.
The coarsest division is per device. Each transmission is independent of the others in the slot. It works, and it balances badly, because devices have very different amounts of data.
The code block is the division that works best. It is independent, it is bounded in size, and there are usually many of them per slot. That combination gives a work queue with many similar tasks, which is what a thread pool needs.
So the standard arrangement is a pool of workers pulling code blocks from a queue. The number of workers matches the cores available for the real-time path. Nothing is assigned in advance, so a slow block does not stall an idle worker.
The join is where the design gets interesting. All code blocks of one transport block must finish before the result is known. So the slowest block sets the completion time, and one unusually difficult block delays a whole transmission.
That is the argument for capping the iteration count rather than letting a hard block run. A block that will not converge is better abandoned than allowed to delay everything behind it. The retransmission that follows costs less than a missed slot.
Figure 3 shows the pool, the join, and the deadline the join has to meet.
Figure 3. The join turns the slowest code block into the cost of the whole transport block. That is why an iteration cap protects the deadline better than any amount of worker tuning.
Per resource element is for vectors : Equalisation and demodulation are independent at that level, and it is far too fine for threads.Per device balances badly : Transmissions are independent, and they differ enormously in size, so workers finish at very different times.The code block is the right unit : Independent, bounded in size, and numerous, which is exactly what a thread pool needs.Pull, do not assign : Workers taking blocks from a queue avoids the idle worker that pre-assignment always produces.The join makes the slowest block the cost : All blocks of a transport block must finish before any result exists.Cap the iterations to protect the join : Abandoning a block that will not converge costs a retransmission, and letting it run costs a slot.
Where does numerical precision actually matter?
A physical layer can be built in floating point or in fixed point. The choice is usually made once for the whole implementation, and it is worth understanding what it actually costs. The answer is not uniform across the chain.
Floating point is the easier choice on a general-purpose processor. Modern cores have wide vector units for it, so the throughput is there. Software stacks running on CPUs almost always use it, and the processing page explains why the vector width is what matters.
Fixed point is the necessary choice in an FPGA, where a floating point unit costs far more logic than it is worth. So the same algorithm exists in two forms, and porting between them is real work rather than a compiler flag.
Within the chain, precision requirements differ sharply. Channel estimation needs reasonable dynamic range, because it deals with the difference between strong and weak subcarriers. Equalisation needs it too, since inverting a small number produces a large one.
Soft bit values are the interesting case, and they need less precision than people expect. The decoder cares about the sign and the rough magnitude of each value. Quantising them to a handful of bits costs very little performance and saves a great deal of memory bandwidth.
That saving matters more than it sounds. Soft values are the largest data structure in the receive chain, and they are written once and read repeatedly by an iterative decoder. Halving their width halves the traffic the decoder generates.
One failure mode is worth recognising because it is silent. Fixed point arithmetic that saturates produces a gradual performance loss, and arithmetic that wraps produces sudden nonsense. The Digital Front End page describes the same distinction on the sample path.
The practical guidance is short. Use floating point on a CPU unless something forces otherwise. Quantise the soft values deliberately rather than by accident. Then check the result against a known signal, because a precision bug looks exactly like a bad radio.
The choice is made once for the whole chain : Floating point on a CPU, fixed point in an FPGA, and porting between them is real work.Estimation and equalisation want dynamic range : Both deal with large ratios, and inverting a small number produces a large one.Soft values need surprisingly little : The decoder uses the sign and a rough magnitude, so a few bits each costs very little performance.Narrow soft values save memory bandwidth : They are the largest structure in the chain, and an iterative decoder reads them repeatedly.Saturation degrades and wrapping breaks : One loses performance gradually and the other produces sudden nonsense, exactly as on the sample path.Verify against a known signal : A precision bug looks identical to a bad radio, and only a controlled input separates them.
Reference
The list below is where the channel structure and coding definitions come from. The specifications are the authority, and the ShareTechnote pages carry the hardware and the layers around this one.
- 3GPP TS 38.211 : NR physical channels and modulation. Reference signals, synchronization signals and the resource grid.
- 3GPP TS 38.212 : NR multiplexing and channel coding. LDPC for data, Polar for control, segmentation and rate matching.
- 3GPP TS 38.214 : NR physical layer procedures for data. Transport block sizes and the modulation and coding tables.
- 5G FAPI : Small Cell Forum - the per-slot requests and indications that drive this layer from the MAC.
- ShareTechnote - The Real-Time Loop : The deadline this layer has to meet, and what happens when it does not.
- ShareTechnote - MAC and the Scheduler : What decides the allocations this layer executes.
- ShareTechnote - Where the Processing Runs : SIMD, GPU and FPGA choices for the stages described here.
- ShareTechnote - Digital Front End : The transform and rate conversion that sit below this layer.