Everything below this point runs against a slot deadline. Everything above it runs against timers measured in milliseconds. That change of pace is the most useful thing to know about the upper layers, because it changes how the code should be written. These layers hold almost all of the state in the system, and that is where their difficulty lives.
- Executive Summary
- What sits above MAC, and why is it separate?
- What does RLC add that HARQ does not?
- Why does reordering exist, and what does it cost?
- What does PDCP actually do?
- How is the control plane structured?
- Why is state management the hard part?
- What does the CU and DU split do to all of this?
- Reference
Executive Summary
The table is a lookup. Notice that no layer here has a slot deadline, which is the property that separates them from everything below.
Layer |
What it adds |
What drives it |
State per device |
Where it lands after the split |
|---|---|---|---|---|
RLC |
Segmentation to fit a grant, and retransmission of what HARQ still lost. |
Timers, and status reports from the far end. |
Transmission and reception windows, per bearer. |
The distributed unit, near the radio. |
PDCP |
Ciphering, integrity protection, header compression and duplicate removal. |
Arriving packets, and a reordering timer. |
Security keys, sequence numbers and compression context. |
The central unit user plane. |
SDAP |
Mapping quality of service flows onto radio bearers. |
Configuration from the core network. |
A mapping table, which is small. |
The central unit user plane. |
RRC |
Connection setup, security, configuration of every layer below, and mobility. |
Protocol timers and messages, both of which take tens of milliseconds. |
The full device context, which is the largest structure in the system. |
The central unit control plane. |
The common property |
None of them touches a resource grid or a symbol. |
Milliseconds, never microseconds. |
All of it survives across slots, unlike the physical layer. |
All of it can be moved away from the radio site. |
What sits above MAC, and why is it separate?
Four layers sit above the MAC, and each exists because the layer below it cannot do that job. Understanding why each one was added explains the design better than any list of its functions.
RLC exists because HARQ is fast and imperfect. It catches the errors that survive, and it also reshapes data to fit whatever size grant the scheduler produced. Neither job belongs in a layer working to a slot deadline.
PDCP exists because security and header compression need a stable, ordered stream. Ciphering depends on a sequence number that both ends agree on. That agreement cannot be maintained by a layer that hands out variable fragments every slot.
SDAP exists because 5G separated quality of service flows from radio bearers. Something has to map one onto the other, and that mapping is configuration rather than processing.
RRC exists because everything below it needs configuring, and the configuration itself has to be negotiated over the air. It is the only layer that changes how the other layers behave.
The property they share is the useful one. None of them touches a symbol or a resource block. They handle packets, sequence numbers and timers, which are ordinary software objects.
So these layers should be written as ordinary software. Normal threading, normal allocation, and normal error handling are all appropriate. The overview page argues that trying to write them in the style of the physical layer produces code that is hard to maintain and no faster.
The queue between MAC and RLC is where the two worlds meet. Below it, everything is measured in microseconds. Above it, a millisecond of delay is invisible.
Each layer exists because the one below cannot : Knowing why each was added explains the design better than listing what it does.RLC catches what HARQ misses : It also reshapes data to fit whatever grant size the scheduler produced, which no deadline-bound layer should do.PDCP needs a stable ordered stream : Ciphering depends on a sequence number both ends agree on, which variable per-slot fragments cannot provide.RRC configures everything below it : It is the only layer that changes how the other layers behave.None of them touches a symbol : They handle packets, sequence numbers and timers, which are ordinary software objects.Write them as ordinary software : Normal threading and allocation are appropriate, and physical layer style here buys nothing.
What does RLC add that HARQ does not?
Two retransmission mechanisms in one stack looks redundant. It is not, and the reason is a deliberate division of labour. One is fast and approximate, and the other is slow and certain.
HARQ operates on the physical layer timescale. A failure is detected within a few slots and corrected almost immediately. It is cheap, and it works because the acknowledgement is a single bit carried on a channel that already exists.
That single bit is also its weakness. An acknowledgement can itself be received incorrectly, and an acknowledgement misread as a positive one loses the data permanently. The probability is small and it is not zero.
So a residual error rate survives HARQ. It is far below the ten percent first-transmission target from the scheduler page, and it is still far above what a data connection can tolerate. Something has to catch it.
RLC acknowledged mode is what catches it. It numbers every packet, tracks which numbers have arrived, and asks for the missing ones by number. Its status reports are full messages rather than single bits, so they are protected properly.
The result is a two-level scheme where each level does what it is good at. HARQ removes almost all errors quickly and cheaply. RLC removes the remainder slowly and reliably. Doing either job with the other mechanism would be far more expensive.
RLC has two other modes, and choosing between them is a real design decision. Unacknowledged mode numbers packets and never retransmits, which suits voice where a late packet is useless anyway. Transparent mode does nothing at all, and it is used for broadcast messages that have no device to acknowledge them.
Segmentation is the other half of the layer, and it is the part that interacts with the scheduler. A grant arrives with a size the scheduler chose, and that size is whatever fitted in the grid. RLC cuts its packets to fit exactly, which is why the layer has to be asked for data rather than pushing it.
The two mechanisms divide the work : HARQ is fast and approximate, and RLC is slow and certain, and each is expensive at the other's job.A single-bit acknowledgement can be misread : A negative read as a positive loses the data permanently, at a small but non-zero rate.A residual error survives HARQ : It is far below the first-transmission target and still far above what a data connection tolerates.RLC asks for missing packets by number : Its status reports are full messages rather than single bits, so they are properly protected.Three modes, three different jobs : Acknowledged for data, unacknowledged for voice where a late packet is useless, and transparent for broadcast.Segmentation makes RLC a pull layer : The grant size comes from whatever fitted in the grid, so RLC is asked for exactly that many bytes.
Why does reordering exist, and what does it cost?
Packets sent in order do not arrive in order, and the reason is the retransmission machinery itself. A packet that needed three attempts arrives after packets that needed one. So the receiver sees a sequence with gaps that fill in later.
Delivering that to the application directly would be a mistake. Higher layers behave badly with out-of-order data, and a transport protocol above will interpret it as congestion. So the receiving side holds packets back until the gaps fill.
That holding is reordering, and it is implemented with a window and a timer. Packets arriving inside the window are stored. When the gap fills, everything up to the next gap is delivered at once.
The timer is what stops it waiting forever. If a gap has not filled when the timer expires, the receiver gives up on it and delivers what it has. Setting that timer is a genuine trade rather than a tuning detail.
Too short and the receiver abandons packets that were about to arrive, which wastes the retransmission that was already paid for. Too long and every lost packet stalls the stream for the full duration of the timer.
That stall is head-of-line blocking, and it is the real cost of reordering. One missing packet holds up every packet behind it, even though those arrived perfectly. A single loss therefore produces a latency spike affecting a whole burst of traffic.
This is why latency measured on a good connection can still show occasional large spikes. The link is not slow. One packet needed several attempts, and everything behind it waited.
Reordering also happens at PDCP, and for a different reason. When a bearer is split across two paths, packets arrive from both and have to be merged in order. The mechanism is the same and the cause is topology rather than retransmission.
Figure 1 shows a single loss producing a stall, with the timer bounding it.
Figure 1. The link was never slow. One packet needed several attempts, and in-order delivery made every packet behind it wait, which is what a latency histogram shows as a long tail.
Retransmission causes reordering : A packet needing three attempts arrives after packets that needed one, so gaps appear and fill later.Out-of-order delivery misleads the layers above : A transport protocol reads it as congestion, so the receiver holds packets back instead.A window and a timer implement it : Packets are stored until the gap fills, then everything up to the next gap is delivered at once.The timer is a real trade : Too short discards retransmissions already paid for, and too long stalls the stream on every loss.Head-of-line blocking is the cost : One missing packet holds up every packet behind it, although those arrived perfectly.Latency spikes on a good link mean this : The connection is not slow, and one packet needed several attempts.PDCP reorders for a different reason : A split bearer delivers packets from two paths, and merging them in order needs the same mechanism.
What does PDCP actually do?
PDCP looks like a collection of unrelated functions, and there is a thread connecting them. Every one of them needs a reliable, ordered, numbered stream, which is exactly what the layer below now provides.
Ciphering is the most obvious. Each packet is encrypted using a key and a counter derived from the sequence number. Both ends compute the same counter from the same number, which is why the numbering has to be agreed rather than local.
Integrity protection is separate from ciphering and is often confused with it. Ciphering hides the content, and integrity protection proves it was not modified. Signalling messages get both, since an altered configuration message is more dangerous than a read one.
Header compression is the function that saves the most and is thought about least. A voice packet may carry forty bytes of headers around twenty bytes of audio. Since those headers barely change between packets, sending only the differences reduces the overhead dramatically.
The compressor keeps a context at each end describing the last header seen. That context is state that must survive, and it makes header compression fragile in a way ciphering is not. Losing synchronisation between the two contexts breaks the flow until it is rebuilt.
Duplication is the newest function and the simplest to describe. The same packet is sent over two paths, and the receiver discards whichever copy arrives second. It costs exactly twice the resources and buys a large reduction in loss.
That trade is only worth making for traffic where reliability matters more than capacity. Duplication is therefore configured per bearer rather than per device, and turning it on for ordinary traffic simply halves the cell's capacity.
One implementation note matters more than it appears to. All of these functions are per packet and per bearer, so their cost scales with throughput rather than with device count. PDCP is the layer that becomes expensive on a fast connection, which is a different scaling law from RRC.
Everything here needs an ordered numbered stream : That is the thread connecting functions that otherwise look unrelated.The sequence number drives the cipher : Both ends derive the same counter from it, so the numbering has to be agreed rather than local.Ciphering and integrity are different : One hides content and the other proves it was not modified, and signalling needs both.Header compression saves the most : Forty bytes of header around twenty bytes of audio becomes a few bytes of difference.Compression context is fragile state : Losing synchronisation between the two ends breaks the flow until the context is rebuilt.Duplication costs exactly double : Configure it per bearer, because enabling it for ordinary traffic halves the cell's capacity.PDCP scales with throughput : Its cost is per packet, which is a different scaling law from the control plane above it.
How is the control plane structured?
RRC is the least like the rest of the stack. It has no throughput requirement and no deadline in microseconds. What it has instead is an enormous number of states, transitions and timers, all of which have to be correct.
The device is in one of three states. Idle means the network has no context for it. Connected means a full context exists and data can flow. Inactive is between them, where a context is kept while the connection is released, so reconnecting is fast.
Every procedure is a message exchange with a timer guarding it. The network sends a message, starts a timer, and waits. If the reply arrives the procedure continues, and if the timer expires the procedure fails and something has to be cleaned up.
That structure repeats everywhere, and it is the reason RRC code is dominated by state machines rather than by algorithms. There is very little computation in this layer. There is a great deal of remembering what was asked and what has not yet been answered.
The messages are defined in ASN.1 and encoded with packed encoding rules. That choice has real engineering consequences. The encoder and decoder are generated from the specification rather than written, and the generated code is large.
Packed encoding is used because it is compact, and compactness matters on a control channel with poor coverage. It is also unusual enough that hand-writing a decoder is a serious mistake. Using a generator is the only maintainable approach.
One property makes this layer harder to test than any other. A procedure can fail at any step, and the correct recovery differs for each. Testing the successful path exercises a small fraction of the code, and the failure paths are where devices in the field actually go.
The asynchronous nature adds the last difficulty. A device can send a message at any moment, including while the network is in the middle of configuring it. Handling a message that arrives in an unexpected state is most of what separates a robust implementation from a demonstration.
No deadline, and enormous state : RRC has no throughput requirement, and it has states, transitions and timers that all have to be correct.Three states, with inactive in between : Keeping a context while releasing the connection is what makes reconnection fast.Every procedure is a message and a timer : Send, start the timer, and wait. Expiry means failure and something to clean up.The code is state machines, not algorithms : There is very little computation, and a great deal of remembering what has not been answered.Generate the ASN.1 codec, never write it : Packed encoding is compact and unusual, and a hand-written decoder is a long-term mistake.The failure paths are where devices go : Testing the successful path exercises a small fraction of the code.Messages arrive in unexpected states : Handling that is most of what separates a robust implementation from a demonstration.
Why is state management the hard part?
The physical layer holds almost no state between slots. The upper layers hold all of it, for every device, for as long as each one stays connected. That difference is what makes the two halves of the stack fail in completely different ways.
A device context is not one object. It is a set of related structures spread across every layer, tied together by identifiers that differ at each level. The scheduler knows a device by one identifier, RLC by a bearer, and RRC by another identifier entirely.
Creating that set is straightforward, since it happens once in a known order. Destroying it is where implementations go wrong, because release can be triggered from several directions at once.
A device can be released because it asked, because a timer expired, because the core network requested it, or because the radio link failed. Those can overlap. A release already in progress can receive a second release request for the same device.
The classic failure is a context that is partly destroyed. The scheduler no longer knows the device and RRC still does, or the reverse. What follows is either a device that cannot reconnect or a slow leak that exhausts the cell over hours.
Leaks in this layer have a distinctive signature worth recognising. A cell works perfectly for a day and then stops accepting new connections. Nothing crashed, and the identifier space or the memory pool is simply full of contexts nobody released.
The defence is to make ownership explicit and singular. One component owns the lifecycle, and every other component holds a reference that it cannot outlive. Release then becomes one operation rather than a sequence that can be interrupted halfway.
The second defence is a counter and a limit. Count active contexts, expose the number, and check it against what the other layers believe. A discrepancy between two layers' counts is the earliest visible symptom of the failure described above.
The two halves fail differently : The physical layer holds almost nothing between slots, and the upper layers hold everything for the whole connection.A context is spread across layers : Different identifiers at each level tie together structures that are created and destroyed as one.Destruction is where it goes wrong : Creation happens once in a known order, and release can be triggered from four directions at once.Partial destruction is the classic bug : One layer forgets the device while another remembers it, producing either a reconnect failure or a leak.Leaks show up a day later : A cell stops accepting new connections with nothing crashed, because the identifier space is full.Make ownership singular : One component owns the lifecycle, and release becomes one operation rather than an interruptible sequence.Count contexts and compare across layers : A discrepancy between two layers' counts is the earliest visible symptom.
What does the CU and DU split do to all of this?
Everything above has described one program. Real deployments cut it in two, and the cut runs through the middle of the layers on this page. Knowing where it falls explains several otherwise puzzling design choices.
The distributed unit keeps the physical layer, the MAC and RLC. Those are the layers with tight timing, so they stay near the radio. The central unit takes PDCP, SDAP and RRC, which tolerate delay.
F1 is the interface between them, and it is a real network protocol rather than a function call. So an RLC instance and its PDCP instance may sit in different buildings, joined by a link with latency of its own.
That latency is why the split lands where it does. RLC retransmission is timer driven and tolerates a few milliseconds. HARQ is not, and moving MAC across a network link would break the slot loop immediately.
The benefit is that one central unit serves many distributed units. Device context and mobility state live in one place, so moving a device between cells becomes an internal operation rather than a full handover. That is the main reason the split exists.
E1 splits the central unit again, into a control plane part and a user plane part. The reason is scaling, as the overview page describes. User plane load follows throughput and control plane load follows the device count.
For an SDR experiment none of this is necessary, and every implementation still lets you run it all in one process. It is worth knowing anyway, because the code is structured around the split even when it is not deployed that way.
That structure explains an oddity in the source. Layers that appear adjacent communicate through message definitions rather than direct calls, and the messages look heavier than the job requires. They are network messages that happen to be delivered locally.
The cut runs through this page's layers : RLC stays with the radio, and PDCP with RRC moves away.Timing decides where it falls : RLC is timer driven and tolerates milliseconds, and moving MAC across a link would break the slot loop.F1 is a network protocol : An RLC instance and its PDCP instance may sit in different buildings, joined by a link with its own latency.One central unit serves many distributed units : Mobility between cells becomes internal rather than a full handover, which is the main reason for the split.E1 splits again by scaling curve : User plane load follows throughput, and control plane load follows the number of devices.The code is split even when the deployment is not : Adjacent layers exchange message definitions rather than calls, because those are network messages delivered locally.
Reference
The list below is where the protocol definitions come from. The specifications are the authority for the procedures and formats. The ShareTechnote pages carry the layers below and the deployment split.
- 3GPP TS 38.322 : NR RLC protocol. The three modes, segmentation, status reporting and the reassembly timer.
- 3GPP TS 38.323 : NR PDCP protocol. Ciphering, integrity protection, header compression, duplication and reordering.
- 3GPP TS 38.331 : NR RRC protocol. The state model, the procedures and the ASN.1 message definitions.
- 3GPP TS 38.473 : F1 application protocol, between the distributed unit and the central unit.
- ShareTechnote - MAC and the Scheduler : The layer below, and the grant sizes RLC has to segment for.
- ShareTechnote - RAN Software Stack : The boundaries and the process model these layers sit inside.
- ShareTechnote - Bring-up and Debugging : How the failures described here appear from the outside.
- ShareTechnote - Open RAN Architecture : The deployment split these layers are divided by.