A cellular base station is the most demanding software an SDR can run. It is not the hardest because the algorithms are exotic. It is hardest because four difficult requirements arrive together, and none of them can be relaxed to help another. This page maps the stack, and the rest of the series takes one part of it each. ShareTechnote already explains what each protocol layer does, so this series covers how the software is built instead.
- Executive Summary
- What does a RAN stack actually have to contain?
- Why is this the hardest software an SDR can run?
- Where are the boundaries, and why are they there?
- What runs in which process and which thread?
- How do the open implementations differ?
- What does the SDR boundary look like from the stack side?
- What is the smallest stack that actually works?
- Reference
Executive Summary
The table is a lookup. The time budget column is the one that separates this software from ordinary software.
Layer |
What it decides |
Time budget |
State it holds |
Where it usually runs |
|---|---|---|---|---|
PHY, layer 1 |
Nothing. It executes what the MAC asked for, and reports what it received. |
A fraction of one slot, so hundreds of microseconds. |
Almost none between slots, apart from HARQ soft buffers. |
A worker thread pool, a GPU, or an FPGA for the lowest part. |
MAC and scheduler |
Which UE transmits, on which resources, with which modulation, in each slot. |
Tens of microseconds, once per slot, for all UEs together. |
Per-UE channel quality, buffers and HARQ processes. |
One thread, usually, because the resource grid is shared state. |
RLC |
Whether a lost packet is retransmitted, and how packets are segmented. |
Milliseconds, driven by timers rather than by slots. |
Transmission windows and reordering buffers, per bearer. |
The same process as MAC, or a separate distributed unit thread. |
PDCP and SDAP |
Ciphering, integrity, header compression and which flow maps to which bearer. |
Milliseconds, and it is throughput bound rather than deadline bound. |
Security keys, sequence numbers and compression context. |
The central unit user plane, once the stack is split. |
RRC |
Whether a UE connects, what it is configured with, and when it moves. |
Tens of milliseconds, governed by protocol timers. |
The full context of every UE, which is the largest state in the system. |
The central unit control plane, on ordinary threads. |
What does a RAN stack actually have to contain?
The layer list is short and the implication is not. A base station needs a physical layer, a MAC with a scheduler, RLC, PDCP, SDAP and RRC. Every one of them has to be present before a single byte of user data moves.
That is the first thing that makes this different from most software. There is no useful subset. A stack with an excellent physical layer and no RRC does not carry a reduced service, since it carries nothing at all.
The layers also differ from each other far more than the diagram suggests. The physical layer is arithmetic under a hard deadline. The scheduler is a decision engine under a slightly softer one. RRC is an asynchronous state machine with no deadline and enormous state.
Those three characters need different code, different threading and different testing. Treating them as one program with a uniform structure is the most common architectural mistake, and it shows up as either a slow physical layer or an unmaintainable control plane.
One division runs through everything and is worth fixing early. The user plane moves bytes and cares about throughput and latency. The control plane makes decisions and cares about correctness and state. They share almost nothing beyond a UE identifier.
This series takes the layers in that order. The real-time loop comes first, because it constrains everything above it. Then the physical layer, the scheduler, and the upper layers. The last page covers bring-up and debugging, which is where most of the time actually goes.
What this series does not do is explain the protocols. ShareTechnote covers NR and LTE procedures in depth already, and repeating that here would help nobody. The subject here is the software that has to implement them.
There is no useful subset : PHY, MAC, RLC, PDCP, SDAP and RRC all have to exist before one byte of user data moves.The layers have different characters : Arithmetic under a hard deadline, a decision engine, and an asynchronous state machine are three different programs.Uniform structure is the usual mistake : Writing all three the same way produces a slow physical layer or an unmaintainable control plane.User plane and control plane share almost nothing : One moves bytes and one makes decisions, and they meet only at a UE identifier.The real-time loop constrains everything : It comes first in this series because every layer above inherits its deadlines.This series is about software, not protocol : The procedures are covered elsewhere on the site. The subject here is how they get implemented.
Why is this the hardest software an SDR can run?
Plenty of software is complicated. What makes a RAN stack unusual is that four hard requirements apply at once, and satisfying one does not relax any other. Most engineering problems allow a trade, and this one allows very few.
The first requirement is hard real time. A slot boundary arrives every 500 microseconds at 30 kHz subcarrier spacing, whether or not the software is ready. Missing it does not slow the system down. It produces a transmission that never happened, which the far end observes as a fault.
The second is exact conformance to a wire format. Every field, every bit position and every encoding rule is specified. A device on the other end was built by someone else from the same document. There is no negotiation and no version tolerance to fall back on.
The third is interoperability with equipment you cannot inspect. A commercial handset is a closed system, and when it refuses to attach it does not explain why. Debugging therefore proceeds by inference from what the other side did next, which the bring-up page covers in detail.
The fourth is stateful concurrency at scale. Every connected device has a context spanning all the layers, and those contexts are created, modified and destroyed while the real-time loop keeps running. A lock held too long in the control plane becomes a missed slot in the physical layer.
Combining the first and the fourth is what makes the architecture difficult. Ordinary concurrent software solves contention with locks and accepts occasional delay. Here an occasional delay is a dropped transmission, so the usual answer is not available.
The result is a characteristic style. Allocation happens before the loop starts. Queues are lock-free where the real-time path touches them, and the control plane is kept off it deliberately. Those choices appear in every serious implementation, and they are a response to this combination rather than to taste.
Four requirements arrive together : Hard real time, exact wire format, interoperability with closed equipment, and stateful concurrency across many devices.A missed deadline is a fault, not a slowdown : The slot boundary arrives regardless, so late work becomes a transmission that never happened.There is no version tolerance : The other end was built from the same specification by someone else. Nothing about the format is negotiable.The other end does not explain itself : A handset that refuses to attach gives no reason. Debugging works by inference from its next action.Real time and shared state conflict directly : A lock held in the control plane becomes a missed slot in the physical layer.The style follows from the constraint : Preallocation, lock-free queues and a control plane kept off the fast path are responses rather than preferences.
Where are the boundaries, and why are they there?
A stack could be one program. Real deployments cut it into pieces at four specific places, and none of the cuts is arbitrary. Each one sits where the traffic character changes, which is what makes a standardised interface worth defining.
The lowest cut is between the radio unit and the distributed unit. That is the 7.2x split, carried over eCPRI, and the Adaptors/Connectors page covers its rates and timing. Below it the data is frequency-domain samples, and above it the data is transport blocks.
The next cut is between MAC and PHY, and it is the one specific to this series. The Small Cell Forum defines FAPI for it, with a P5 interface for configuration and a P7 interface for the per-slot exchange. P5 sets up a cell, and P7 carries the requests and indications that flow every slot.
That division matters because the two sides have different natures. P5 traffic is rare and complex, and P7 traffic is simple and relentless. Separating them lets the per-slot path stay narrow enough to be fast.
nFAPI extends the same interface across a network, which is what allows the MAC to be virtualised away from the PHY. It is the same message set, wrapped for transport, and it exists so a MAC and a PHY from different vendors can work together.
The third cut is F1, between the distributed unit and the central unit. RLC and everything below stay in the distributed unit near the radio. PDCP and RRC move to the central unit, which can then serve several distributed units and can sit far away.
The fourth is E1, which splits that central unit again into a control plane part and a user plane part. The reason is scale rather than latency, since user plane traffic grows with throughput and control plane load grows with the number of devices. Splitting them lets each scale on its own.
Figure 1 places all four cuts on one diagram, with the traffic character on each side.
Figure 1. The cuts are not evenly spaced, because they follow changes in what the data is rather than the layer diagram. Reading them from the bottom shows the data becoming steadily more abstract and steadily less urgent.
Cuts follow traffic character : Each boundary sits where the data changes nature, which is what makes a standardised interface worth having.7.2x separates samples from transport blocks : Below it the fronthaul carries frequency-domain IQ, and above it the stack carries coded data.FAPI splits configuration from the slot loop : P5 sets a cell up and P7 carries the per-slot requests and indications.The two FAPI halves have opposite natures : P5 is rare and complex, and P7 is simple and relentless, so separating them keeps the fast path narrow.F1 moves the latency-tolerant layers away : RLC and below stay near the radio, and PDCP with RRC can sit in another building.E1 splits by scaling curve : User plane load grows with throughput and control plane load grows with device count.
What runs in which process and which thread?
The layer diagram says nothing about concurrency, and concurrency is where the design actually lives. Three groups of work exist, and each one wants a different arrangement. Getting the grouping wrong costs more performance than any algorithm choice.
The physical layer is the easy case, because its work divides naturally. Each transport block is independent of every other one in the same slot, and each code block inside it is independent again. So a worker pool with one task per code block scales well. The PHY page covers why that unit is the right one.
The scheduler is the opposite case. It allocates a single shared resource grid, so every decision depends on the decisions already made in that slot. Splitting it across threads means locking the grid, and the lock would be held for most of the work.
So the scheduler is usually one thread, and that is a deliberate design rather than a limitation. It runs once per slot and finishes in tens of microseconds. Making it concurrent would cost more in coordination than it could win in parallelism.
The upper layers are ordinary concurrent software and should be written that way. RRC handles procedures that take tens of milliseconds, and PDCP moves bytes without a slot deadline. Neither belongs on the real-time path, and putting them there is a common way to create unexplained late slots.
One rule ties the three together. Work with a slot deadline runs on isolated cores with a real-time policy, and everything else runs normally. The boundary between them has to be a queue rather than a function call. A call transmits the caller's timing problems to the callee.
Figure 2 shows the arrangement, with the queues marked. The queues are the interesting part, since they are where the two timing worlds meet.
Figure 2. The green line is the important part of the picture. Everything to its left is written under a deadline, everything to its right is written normally, and the queue is what stops the two disciplines contaminating each other.
The PHY divides naturally : Transport blocks and the code blocks inside them are independent, so a worker pool scales with the core count.The scheduler does not divide : It allocates one shared resource grid, so every decision depends on the ones already made this slot.One scheduler thread is a design, not a limit : Coordination would cost more than the parallelism could return, for work that finishes in tens of microseconds.Upper layers are ordinary software : RRC and PDCP have no slot deadline, and putting them on the real-time path creates unexplained late slots.Separate the two worlds with a queue : A direct call hands the caller's timing problem to the callee. A queue absorbs it instead.Grouping beats algorithms : Getting the concurrency arrangement wrong costs more than any individual algorithm choice recovers.
How do the open implementations differ?
Two open stacks are commonly used with SDR hardware, and comparing their feature lists is not very informative. Comparing their architectures is, because each one made a different structural choice and lives with the consequences. The choice is visible in how you deploy them.
OpenAirInterface is the oldest and the most research-oriented. It covers a very wide range of features and configurations, and it exposes most of them. That breadth is its value for experiments. It also means a working configuration is something you arrive at rather than assume.
srsRAN was built later with a cleaner separation between layers, and it is generally easier to bring up. Its distinctive feature for SDR work is a virtual radio option, which replaces the RF hardware with a network transport between processes. That turns a two-radio experiment into two processes on one machine.
Between them these two settle a question that used to be open. A full stack runs on general-purpose cores, in software, with no dedicated silicon in the signal path. So the remaining differences between implementations are engineering effort rather than feasibility.
One structural point separates both of them from a textbook design. None of them is a straightforward layered program with each layer calling the one below. The real-time path is arranged around the slot loop. The layer structure is how the code is organised, rather than how it executes.
That distinction is worth understanding before reading any of the source. Following a packet downward through the layers is not how the program runs. The scheduler is working several slots ahead of the physical layer. The real-time loop page explains that arrangement.
For learning, the virtual radio option is the single most useful feature in any of them. It removes the radio, the antennas, the timing and the interference from the experiment at once. The software behaviour stays intact. Almost every question about the stack can be answered without an SDR connected at all.
Compare architectures, not feature lists : Each stack made a different structural choice, and that choice is what you actually live with.OpenAirInterface trades ease for breadth : Very wide coverage and most of it exposed, so a working configuration is arrived at rather than assumed.srsRAN separates layers more cleanly : It is generally easier to bring up, and its virtual radio option removes the hardware from an experiment.The performance question is settled : A full stack runs on general-purpose cores in software, so the remaining gap between implementations is engineering effort.None of them is a layered program : The code is organised in layers and executes around the slot loop, which are different things.The virtual radio is the best learning tool : It removes radio, antennas, timing and interference at once while leaving the software behaviour unchanged.
What does the SDR boundary look like from the stack side?
Every page in the Building Blk group describes this boundary from the radio's side. From the stack's side it looks much simpler and it carries one additional obligation. The stack does not ask for samples. It asks for the samples belonging to a particular moment.
The interface is a pair of timestamped streams. The stack reads received samples tagged with their capture time. It writes transmit samples tagged with the time they should be emitted. Everything else in this series depends on those tags being correct.
That is why the Syncrhonization page matters more here than anywhere else on the site. A stack derives its entire notion of time from the sample counter, so an incorrect timestamp does not cause a timing error. It causes the stack to transmit in the wrong slot, which looks like a protocol failure.
The stack also needs the transmit stream to run ahead of real time. Samples for a future slot have to reach the radio before that slot arrives, so the interface always carries a lead. Sizing that lead is the subject of the real-time loop page.
Flow control comes from the radio rather than from a clock, and this is the detail most first attempts get wrong. The stack does not sleep until the next slot is due. It blocks reading samples, and the samples arrive at exactly the rate the converter produces them.
That inversion is worth stating plainly. The radio paces the software. A loop built on a timer drifts against it, however carefully the timer is set. The buffering page explains what the drift does to the buffer chain.
Once the split moves to 7.2x the boundary changes shape but not character. The stack now exchanges frequency-domain samples with an O-RU inside a delivery window, rather than time-domain samples with a converter. It is still a timestamped exchange with a lead, and the deadline is stricter.
The stack asks for a moment, not for samples : Both streams are timestamped, and everything above depends on those tags being right.A wrong timestamp is a protocol failure : The stack derives all of its time from the sample counter, so it transmits in the wrong slot rather than late.The transmit stream always leads : Samples for a future slot must reach the radio before it arrives. The interface carries a deliberate lead.The radio paces the software : The loop blocks on receiving samples rather than sleeping on a timer, and a timer-based loop drifts.Flow control inversion catches first attempts : Waiting for the next slot with a sleep looks correct. It is wrong in a way that grows slowly.7.2x changes shape, not character : Frequency-domain samples inside a delivery window, still timestamped, still leading, with a stricter deadline.
What is the smallest stack that actually works?
The specification is enormous and most of it is optional. A useful question is therefore what the minimum is, because that is what a first bring-up should target. The answer is smaller than the document count suggests and larger than most people hope.
On the downlink the cell has to transmit continuously. That means the synchronization signal block, the master information block on the broadcast channel, and system information block one. Without those three a device does not know the cell exists.
Then random access has to work in both directions. The cell must detect a preamble, send a random access response with a timing advance and an uplink grant, receive the third message, and complete contention resolution. Each of those four steps has its own failure mode, which the bring-up page covers.
Above that the minimum is a working RRC setup, a security procedure, and one data radio bearer. The security step is not optional in the way it is sometimes assumed to be, since a modern device will not proceed without it.
Underneath all of it the scheduler must handle at least one downlink and one uplink allocation per slot, with HARQ retransmissions working. HARQ is the part most often postponed and it cannot be, because the first real channel will need it immediately.
What can be left out is genuinely large. Carrier aggregation, mobility procedures, beam management and dual connectivity are all absent from a first working system. So is nearly all of the measurement configuration. So is almost every optional feature flag in the specification.
The useful conclusion is about ordering rather than scope. Build the continuous downlink first, then random access, then the connection, then data. Each stage is observable from the outside. A device tells you honestly whether the previous one worked, before it attempts the next.
The downlink has to run continuously first : Synchronization signals, the master information block, and system information block one. Without them no device knows the cell exists.Random access is four separate steps : Preamble detection, the response with a timing advance, the third message, and contention resolution.Security is not optional : A modern device will not proceed past the connection setup without it, whatever a minimal design might prefer.HARQ cannot be postponed : It is often left for later, and the first real radio channel needs it immediately.What is omitted is enormous : Carrier aggregation, mobility, beam management, dual connectivity and most measurement configuration are all absent from a first system.Build in the order the device checks : Each stage is externally observable. A device attempts the next only after the previous one worked.
Reference
The list below is where the interface definitions and the implementation details come from. The project documentation is the authority for anything specific to a stack, and the ShareTechnote pages carry the hardware chain underneath.
- 5G FAPI : Small Cell Forum - the P5 configuration interface and the P7 per-slot interface between MAC and PHY.
- 5G nFAPI : Small Cell Forum - the same message set wrapped for network transport, for a virtualised MAC and PHY split.
- srsRAN Project : Software Radio Systems - stack architecture, the virtual radio option and the deployment configurations.
- OpenAirInterface : OpenAirInterface Software Alliance - the reference research stack for 4G and 5G.
- O-RAN Specifications : O-RAN Alliance - the 7.2x split, and the F1 and E1 interfaces referenced in the boundary section.
- ShareTechnote - The Real-Time Loop : The slot loop that every layer above inherits its deadlines from.
- ShareTechnote - Adaptors/Connectors : The fronthaul and host transports underneath the stack.
- ShareTechnote - Syncrhonization : The sample timestamps the whole stack derives its time from.
- ShareTechnote - Where the Processing Runs : The CPU, GPU and FPGA partitioning the stack is deployed onto.