Every other page in this series inherits its constraints from this one. A RAN stack is not a program that processes data as it arrives. It is a program driven by a clock it does not own, working on several moments at the same time. I found the stack impossible to read until I understood that, and obvious afterwards. Each section below takes one part of the loop.
- Executive Summary
- What actually drives the loop?
- How long is a slot, and what has to happen inside one?
- Why does the stack run ahead of the radio?
- What are k0, k1 and k2, and whose problem are they?
- What happens when a slot is late?
- How is the loop actually built?
- Where does the time go inside one slot?
- Reference
Executive Summary
The table is a lookup. The slot is the unit of everything, and its length is the only real budget the stack has.
Subcarrier spacing |
Slot length |
Slots per second |
Where it is used |
What it does to the software |
|---|---|---|---|---|
15 kHz |
1000 microseconds |
1000 |
Low bands, and the same numerology LTE used. |
Comfortable. A general-purpose machine has real headroom here. |
30 kHz |
500 microseconds |
2000 |
The common mid-band choice, and the default for most deployments. |
Workable, and it is where careful engineering starts to matter. |
60 kHz |
250 microseconds |
4000 |
Wider mid-band carriers and some low latency configurations. |
Demanding. Operating system jitter becomes a first-order problem. |
120 kHz |
125 microseconds |
8000 |
Millimetre wave bands. |
Hardware acceleration is normal rather than optional. |
The rule behind them |
A slot is always 14 symbols, so doubling the spacing halves the slot. |
Doubles each step. |
The choice belongs to the band, not to the implementer. |
Every deadline on this page scales with it. |
What actually drives the loop?
The first surprise in a RAN stack is that nothing calls the system clock. There is no timer firing every 500 microseconds. Time comes from the radio, and specifically from counting samples.
The converter produces samples at a fixed rate. Counting them therefore measures time exactly, because the count and the elapsed time are the same fact expressed differently. A slot is simply a fixed number of samples.
So the loop reads samples and blocks when none are available. The radio releases them at the converter rate. That blocking read is what paces the entire stack, and it is more accurate than any timer the operating system offers.
The consequence is worth stating plainly. The stack's clock is the radio's clock. If the radio's reference drifts, the stack drifts with it, and no internal correction is possible or desirable.
This is the point where the Syncrhonization page stops being background reading. A shared reference means two base stations count at the same rate. A shared time mark means they agree on which slot number they are in.
Slot numbering is the second half of the mechanism. The count gives a position, and the frame structure turns that position into a system frame number and a slot index. Every scheduling decision is expressed in those terms rather than in seconds.
One design consequence follows immediately. A stack should never take a timestamp from the operating system for anything that matters. Two clocks in one program eventually disagree, and the resulting bug appears hours later as a slow drift.
Nothing calls the system clock : There is no periodic timer. Time is derived by counting samples at a known rate.A slot is a number of samples : Counting samples measures time exactly, because the two are the same fact stated differently.The blocking read is the pacer : The loop waits on samples rather than on a timer. That is more accurate than anything the operating system offers.The stack inherits the radio's clock : If the reference drifts the stack drifts, and correcting that internally would be wrong.Everything is expressed in frame and slot : Scheduling decisions name a system frame number and a slot index, never a time in seconds.Never mix in a second clock : An operating system timestamp on the critical path produces a drift bug that appears hours later.
How long is a slot, and what has to happen inside one?
A slot always contains 14 OFDM symbols. Its duration therefore follows the subcarrier spacing directly. Doubling the spacing halves the symbol, which halves the slot.
At 30 kHz spacing a slot lasts 500 microseconds. That is the figure used throughout this series, because it is the common mid-band configuration. It means two thousand complete cycles of work every second.
Listing what has to happen in one of those cycles makes the problem concrete. The uplink samples for the current slot have to be received and processed. The scheduler has to decide what happens in a future slot. The downlink samples for another future slot have to be generated and handed to the radio.
Notice that those three activities concern three different slots. That is the single most important structural fact about the stack. The loop is not processing one moment at a time.
The uplink work is the largest and the least predictable. It includes the transform, channel estimation, equalisation, demodulation and decoding. The decoder is iterative, so its cost depends on the channel rather than on the configuration.
The downlink work is smaller and almost perfectly predictable. Encoding and modulation cost the same regardless of the channel. That asymmetry matters for budgeting, because only one side of the loop has a variable cost.
The scheduler sits between them and is small in time and large in consequence. It runs once, it touches every active device, and it decides what both of the other two activities will do. The scheduler page covers what it actually computes.
Fourteen symbols, always : The slot length follows the subcarrier spacing directly, so doubling the spacing halves the slot.500 microseconds is the working figure : That is 30 kHz spacing, the common mid-band choice, and two thousand complete cycles every second.Three activities, three different slots : Receiving one slot, scheduling another, and generating a third all happen in the same cycle.Uplink cost is variable : The decoder is iterative, so its cost depends on the channel rather than on the configuration.Downlink cost is predictable : Encoding and modulation cost the same whatever the channel does, which makes them easy to budget.The scheduler is small and decisive : It runs once per slot and determines what both of the other activities will do.
Why does the stack run ahead of the radio?
A transmitted slot has to exist before it is transmitted. That sounds trivial. It forces the entire stack into a pipeline, and it is the reason the code cannot be read as a sequence.
Work backwards from the antenna. Samples must reach the radio before the slot boundary they belong to. Generating them takes time. Deciding what they contain takes more time, and that decision needs uplink feedback that arrived even earlier.
So each stage runs a fixed number of slots ahead of the one below it. The scheduler works on slot n plus four while the physical layer builds slot n plus two and the radio emits slot n. Those offsets are configuration, and they are the most important configuration in the system.
Setting the lead too small is the obvious failure. A stage that has not finished when the radio needs its output produces an empty transmission. That is the late slot condition, and the next section covers what it looks like.
Setting the lead too large is the failure people do not anticipate. Every slot of lead is a slot of latency added to every packet. It also forces the scheduler to decide earlier, using channel information that is older, which makes its decisions worse.
So the lead is tuned rather than maximised. It is chosen as the smallest value that never produces a late slot on the worst-case machine, plus a small margin. The processing page explains why the worst case rather than the average is the number that matters.
One property of this arrangement causes real confusion when reading a stack for the first time. A variable named for the current slot usually refers to a slot that has not happened yet. Which slot a piece of code is working on is a question worth answering before reading anything else.
Figure 1 shows the three timelines together, with the lead marked.
Figure 1. The lead is tuned to the worst case rather than the average, because a single late slot is a real transmission failure. Every slot of lead added to fix that is a slot of latency charged to every packet.
A slot must exist before it is transmitted : That single requirement forces the whole stack into a pipeline rather than a sequence.Each stage leads the one below it : The scheduler works several slots ahead of the physical layer, which works ahead of the radio.Too small a lead produces empty transmissions : A stage that misses its handover deadline leaves the radio with nothing to send.Too large a lead costs latency and accuracy : Every slot of lead delays every packet, and it forces decisions on older channel information.Tune it to the worst case : The smallest value that never produces a late slot on the slowest machine, plus a small margin.Ask which slot the code is on : A variable named for the current slot usually refers to one that has not happened. That is the main obstacle to reading a stack.
What are k0, k1 and k2, and whose problem are they?
The previous section covered the lead inside the base station. A second set of offsets governs the exchange with the device. They appear constantly in configuration and in logs, so they are worth understanding as a group rather than individually.
All three are slot offsets carried in downlink control information. Each one says how many slots to wait before doing something. Their names come from the specification and their purpose is identical.
k0 is the offset from the control message to the downlink data it schedules. A device reads the grant in one slot and receives the data k0 slots later. Setting it to zero means the data arrives in the same slot as its grant.
k1 is the offset from downlink data to the acknowledgement of it. The device decodes the data, decides whether it succeeded, then reports on the uplink k1 slots after receiving it. That reply is what drives the HARQ machinery in the scheduler.
k2 is the offset from the control message to the uplink transmission it grants. The device reads the grant and transmits k2 slots later. It is the uplink counterpart of k0.
The reason none of these can be zero on the uplink side is the device's own processing time. The specification defines a minimum, expressed in symbols, that depends on the subcarrier spacing and on the capability the device declared. Scheduling tighter than that produces no reply at all.
That constraint belongs to the scheduler rather than to the physical layer. It has to pick values that respect the slowest device currently connected. Picking them too generously wastes latency on every transaction, which is the same trade as the internal lead.
One practical note explains a whole class of bring-up failures. A device that receives data and never acknowledges it is often being scheduled inside its own processing time. The symptom looks like a decoding failure and the cause is a timing configuration.
Figure 2 places the three offsets on one timeline.
Figure 2. All three offsets are the scheduler's choice within limits the device sets. Generous values are safe and slow, and aggressive values are fast until one device on the cell cannot keep up.
All three are slot offsets in the grant : Each says how many slots to wait, and all are carried in downlink control information.k0 is grant to downlink data : Zero means the data arrives in the same slot as the grant that scheduled it.k1 is data to acknowledgement : The reply it produces is what drives the HARQ machinery inside the scheduler.k2 is grant to uplink transmission : It is the uplink counterpart of k0, and it has the same structure.The device sets the lower bound : A declared minimum processing time, in symbols, depending on the subcarrier spacing and the capability.Too tight looks like a decoding failure : A device scheduled inside its processing time does not answer. The cause is configuration rather than radio.
What happens when a slot is late?
Ordinary software degrades when it runs out of time. This software does not degrade. A deadline missed by one microsecond has the same effect as one missed by a hundred milliseconds.
On the transmit side the radio needs samples for slot n before slot n begins. Samples that arrive afterwards are useless, because the moment they described has passed. The radio transmits whatever was in its buffer instead, which is usually nothing.
A device observes that as a slot where the expected transmission did not appear. If the slot carried a grant, the device never learns it was scheduled. If it carried data, the device reports a failure and the scheduler retransmits.
On the receive side a late slot has a different shape. The samples are still there, since they were captured correctly. What is missing is the decoded result at the moment the acknowledgement was due.
The scheduler then has to answer without the information it wanted. It normally treats a missing result as a failure, which triggers a retransmission of data that may have arrived perfectly well. So a processing overrun costs throughput even when the radio link is ideal.
Every serious stack counts these events, and the counter is the first thing to look at when performance disappoints. A cell that carries traffic but never reaches its expected rate is often losing slots rather than losing packets.
The failure has one property that makes it hard to catch in testing. It is load dependent, so it appears when a second device connects or when the channel degrades. A stack that ran perfectly with one device at high signal strength has not been tested for this at all.
The correct response is rarely to optimise the slow function. It is to check the four causes from the processing page first: scheduler preemption, frequency scaling, thermal throttling and memory placement. Those account for most late slots on a machine that is otherwise fast enough.
There is no graceful degradation : A deadline missed by a microsecond has the same effect as one missed by a hundred milliseconds.A late transmit slot is silence : The moment has passed, so the radio sends whatever was in its buffer, which is usually nothing.A late receive slot is a false failure : The samples were captured correctly. The decoded result is not ready when the answer is due.Overruns cost throughput on a perfect link : A missing result is treated as a failure, so data that arrived correctly gets retransmitted anyway.Check the counter before the code : A cell that carries traffic but misses its expected rate is often losing slots rather than packets.It is load dependent : One device at high signal strength does not exercise it, so a clean single-device test proves very little.Look at the machine before the algorithm : Preemption, frequency scaling, throttling and memory placement account for most late slots.
How is the loop actually built?
The structure is simpler than the constraints suggest. One thread owns the timing, and everything else hangs off it. Getting that thread right is most of the work.
The timing thread does one thing repeatedly. It reads a slot worth of samples from the radio, which blocks until they exist. It converts the sample count into a slot number. Then it hands work out and returns to reading.
Handing work out is the part that must not block. The timing thread posts tasks to the physical layer workers and returns immediately. If it waits for them, the wait comes out of the next slot's budget, and one slow slot becomes a permanent deficit.
The workers process the slot in parallel and signal completion. A separate join point collects their results before the deadline that needs them. That deadline is earlier than the end of the slot, because the result feeds the scheduler.
The scheduler runs on the timing thread itself in most designs. It is short, it needs a consistent view of all state, and moving it to another thread would need locking. The overview page explains why that trade rarely pays.
Everything above RLC runs elsewhere entirely, connected by queues. Those queues are the boundary between hard timing and ordinary software. A packet crossing them can wait a millisecond without anyone noticing, which is exactly the point.
One mistake is common enough to name, because it looks correct in every code review. Sleeping until the next slot is due, using a system timer, appears to do the same job as blocking on samples. It drifts against the radio, slowly, and the drift only becomes visible after minutes of operation.
Figure 3 shows the arrangement, with the blocking point and the non-blocking handovers marked.
Figure 3. The single blocking point is what keeps the loop aligned to the radio. Every other handover is a post rather than a wait, because waiting spends the next slot's budget on this slot's problem.
One thread owns the timing : It reads samples, converts the count to a slot number, dispatches work, and returns to reading.It blocks in exactly one place : On the sample read. Every other handover is a post that returns immediately.Waiting turns one slow slot into a deficit : Time spent waiting comes out of the next slot's budget, so the loss accumulates.The scheduler shares that thread : It is short and needs a consistent view of state, so moving it would only add locking.Queues separate the two timing worlds : Above RLC a millisecond of delay is invisible, and that is precisely why the queue is there.Never sleep on a system timer : It passes review, matches the intended behaviour, and drifts against the radio over minutes.
Where does the time go inside one slot?
The budget is 500 microseconds at 30 kHz spacing. Knowing how it divides is what turns a vague performance concern into a measurement. The division is uneven and it is not what most people guess.
Uplink processing takes the largest share. Within it the decoder dominates, and the decoder's cost varies with the channel. A poor channel needs more iterations, so the worst case appears exactly when the cell is busiest.
Downlink generation takes a smaller and steadier share. Encoding and modulation cost what they cost. That predictability is why the downlink is rarely the reason a stack misses a deadline.
The scheduler takes a small share in time and touches everything. Tens of microseconds is typical for a cell with a modest number of devices. It grows with the device count rather than with the throughput, which is a different scaling law from the rest.
The remainder is the margin, and it is not spare capacity. It absorbs the jitter that the operating system, the memory system and the transport introduce. A design with no margin works on an idle machine and fails on a real one.
Two scaling behaviours are worth separating, because they mislead people in opposite directions. Physical layer cost scales with bandwidth and with the number of active devices. Scheduler and control plane cost scales with the number of connected devices, whether or not they are sending anything.
So a cell with two hundred idle devices and one heavy user is a control plane problem. A cell with two heavy users is a physical layer problem. Testing one and deploying the other is a common and expensive mistake.
The practical instruction is the same as on the processing page. Measure the per-slot time and record the maximum, not the mean. The mean will look comfortable in a system that is already dropping slots.
Uplink processing takes the largest share : The decoder dominates it, and its cost rises with a poor channel rather than staying fixed.The worst case arrives when the cell is busy : More devices and worse channels both push decoder iterations up at the same time.Downlink is steady : Encoding and modulation cost the same every slot, so the transmit path rarely causes a missed deadline.The scheduler is cheap and grows differently : Tens of microseconds, scaling with the number of connected devices rather than with throughput.The margin is not spare : It absorbs operating system, memory and transport jitter, so a design without it fails outside the laboratory.Two scaling laws, two different problems : Many idle devices stress the control plane, and few busy ones stress the physical layer.Record the maximum per-slot time : The mean looks comfortable in a system that is already losing slots.
Reference
The list below is where the timing definitions come from. The specifications are the authority for the slot structure and the scheduling offsets, and the ShareTechnote pages carry the layers around this loop.
- 3GPP TS 38.211 : NR physical channels and modulation. Frame structure, numerology and the 14 symbol slot.
- 3GPP TS 38.213 : NR physical layer procedures for control. The k1 acknowledgement timing and the device processing time requirements.
- 3GPP TS 38.214 : NR physical layer procedures for data. The k0 and k2 offsets, and the scheduling constraints around them.
- 5G FAPI : Small Cell Forum - the P7 per-slot request and indication exchange that carries this timing between MAC and PHY.
- ShareTechnote - RAN Software Stack : The layers and boundaries this loop drives.
- ShareTechnote - Physical Layer Software : What fills the largest share of the budget described here.
- ShareTechnote - MAC and the Scheduler : What chooses the k values and consumes the acknowledgements.
- ShareTechnote - Syncrhonization : The sample counter this loop derives all of its time from.
- ShareTechnote - Where the Processing Runs : The machine level causes of late slots.