SDR(Software Defined Radio)

 

 

 

Bring-up and Debugging

Most of the time spent on a RAN stack is not spent writing it. It is spent working out why a device will not attach, with almost no information about what the device objected to. This page is about that process. It is the last page of the series because it uses everything in the earlier ones, and it is the one I would have wanted first.

Executive Summary

The table is a lookup. Work down it in order, because a device only attempts each step after the previous one succeeded.

Step

What proves it worked

The usual cause when it fails

Where to look

The radio streams

Samples arrive with no overflow reports, before any stack is started.

Transport, buffering or clocking, none of which is a stack problem.

The radio driver, and a plain capture tool.

The cell transmits

A second receiver sees the synchronization block at the expected frequency.

Frequency configuration, transmit gain, or the cell never started.

A spectrum view from any other receiver.

A device finds the cell

The cell appears in the device's manual network scan.

Wrong band or channel number, wrong network identity, or system information not decodable.

The device's own scan, then the broadcast configuration.

Random access completes

All four steps appear in the trace, in order.

Uplink gain, preamble configuration, timing advance, or a duplex pattern mismatch.

The MAC trace, and the uplink signal level.

The connection is established

Setup completes, security runs, and a data bearer exists.

Core network reachability, credentials, or a configuration the device rejected.

The control plane trace, on both the radio and core interfaces.

Data flows at a sensible rate

Throughput is near what the configuration should give.

Link adaptation, scheduling, or late slots from the machine itself.

Per-slot timing counters, and per-device grant statistics.

Why is a RAN so hard to debug?

Ordinary debugging assumes you can observe the failing system. Here the interesting half of the system is a device you did not build and cannot instrument. That single fact shapes everything about the process.

A commercial handset will not tell you why it refused. It has a reason, and the reason is not exposed. What it does instead is stop, or retry, or go back to scanning, and those three responses are all the information available.

So the method becomes inference from behaviour rather than inspection of state. A device that retries once and gives up is saying something different from one that never responds at all. Learning to read those differences is most of the skill.

The second difficulty is that the medium is not observable by ordinary means. A print statement cannot show what was actually transmitted. Between the software and the device sit the converter, the amplifier, the antenna and the air, and each of them can be the problem.

The third is that there are three parties rather than two. The base station, the device and the core network all have to agree, and a failure in any of them looks similar from the others. A device that attaches and immediately detaches is usually a core network problem reported by the radio.

The fourth is timing, which the real-time loop page covers. A stack that is functionally perfect and slightly late produces failures that look like protocol errors. Nothing in the message trace hints at the real cause.

Those four together explain why a systematic order matters so much here. Guessing is unusually expensive, because most guesses cannot be checked quickly. The remedy is to make each step independently verifiable, which is what the next section is about.

  • The other half is not instrumentable : A commercial device has a reason for refusing, and it does not expose it.
  • Behaviour is the only signal : Retrying once, retrying forever and never responding are three different messages about what went wrong.
  • A print statement cannot see the air : The converter, the amplifier, the antenna and the channel all sit between the software and the device.
  • Three parties have to agree : A device that attaches and immediately detaches is usually a core network problem reported by the radio.
  • Late looks like wrong : A functionally perfect stack that misses deadlines produces failures the message trace cannot explain.
  • Guessing is unusually expensive : Most guesses cannot be checked quickly, which is why a fixed order pays for itself.

What is the bring-up order?

The order is not a matter of preference. A device performs these steps in sequence, and it attempts each one only after the previous succeeded. So the order is imposed by the device rather than chosen by you.

That is the most useful property in the whole process. It means a failure localises itself. Whatever the last successful step was, the fault is in the step immediately after it.

Start below the stack entirely. Confirm that the radio streams samples cleanly with no stack running at all. The buffering page covers what to check, and doing this first removes a whole class of cause from every later question.

Then confirm the cell transmits. A second receiver, or a spectrum analyser, should show the synchronization block at the expected frequency and at a sensible level. This step needs no device and it catches most configuration errors.

Only then involve a device, and use its manual network scan first. A cell that appears in the scan proves that synchronization, the broadcast channel and the first system information block all decoded correctly. That is a great deal of the downlink verified in one observation.

Random access is the next rung and it is where most time is spent. It has four steps, each of which fails differently. The next section but one covers what each failure means.

After that come connection setup, security, and the bearer. These are control plane problems, and they usually involve the core network rather than the radio. A device that reaches this point has already proved the radio works.

Data last, and slowly first. A connection that passes a ping but achieves poor throughput is a different problem from one that passes nothing, and mixing the two wastes a great deal of time.

Figure 1 shows the ladder with what each rung proves.

Each rung is attempted only after the one below it succeeded 1. Radio streams, no stack running removes transport and clocking from every later question 2. Cell transmits, seen on a second receiver needs no device, and catches most configuration errors 3. Device lists the cell in a manual scan proves sync, broadcast channel and system information all decode 4. Random access completes, all four steps where most bring-up time is actually spent 5. Connection, security and a bearer usually the core network rather than the radio 6. Data at a sensible rate passing a ping and being slow are different problems The fault is always in the rung above the last one that worked.

Figure 1. The device imposes this order, which is why a failure localises itself. Skipping a rung to save time reliably costs more than it saves, because the skipped step becomes a candidate cause for everything above it.

  • The device imposes the order : Each step is attempted only after the previous succeeded, so you cannot choose a different sequence.
  • A failure localises itself : The fault is in the rung immediately above the last one that worked.
  • Start below the stack : Confirm clean sample streaming with nothing running, which removes transport and clocking from every later question.
  • A manual network scan proves a lot : A cell appearing in it verifies synchronization, the broadcast channel and system information together.
  • Random access is where the time goes : It has four steps and each fails differently, which is why it deserves its own signature table.
  • Separate no data from slow data : They are different problems, and treating them as one wastes a great deal of time.

Where do you actually look?

Three observation points exist, and each one sees a different part of the problem. Using the wrong one is the most common reason a session produces no progress. They are worth learning as a set.

Logs are the first and the least useful in isolation. A stack at a high log level produces an enormous volume, and most of it describes normal operation. Their value is in the error and warning lines, and in the counters rather than the narrative.

Protocol traces are the second and they are where most answers are. A stack can write its exchanges in a capture format, which opens directly in a protocol analyser with proper dissectors. What arrives is a decoded, timestamped message sequence rather than a wall of text.

That format matters more than it sounds. A message trace shows the exchange from both sides, so the point where a device stopped responding is visible immediately. It also decodes the encoded messages, which is not something a log line can do.

The third is the IQ capture, and it answers questions the other two cannot. When the trace shows nothing received, the question becomes whether nothing was sent or whether it was sent wrongly. Only the samples can answer that.

This is where the recording page becomes directly useful. A capture taken at the moment of failure can be replayed offline as many times as needed, against an instrumented receiver, without the device present.

Replay is the technique that turns an intermittent problem into a reproducible one. The samples do not change between runs, so a failure that happened once becomes a failure that happens every time. That alone justifies the disk space.

One rule about which tool to use makes the choice easy. If the message trace shows a message, the problem is above the physical layer, so use the trace. If it shows nothing where something was expected, drop to the samples.

Figure 2 maps the three onto the stack.

Three observation points, three different questions RRC and PDCP RLC and MAC PHY radio and the air protocol trace decoded, timestamped, both sides of the exchange most answers are here IQ capture was it sent at all, and was it sent correctly logs read the counters and the error lines, not the narrative If the trace shows a message, work above the physical layer. If it shows nothing, drop to the samples.

Figure 2. The trace answers what was exchanged and the capture answers what was transmitted. Reaching for logs first is the usual mistake, because they describe intent rather than result.

  • Logs describe intent, not result : Their value is in the counters and the error lines rather than in the narrative volume.
  • The protocol trace holds most answers : A decoded, timestamped exchange shows immediately where a device stopped responding.
  • Only samples answer whether it was sent : When the trace shows nothing, the question moves below the physical layer.
  • Replay makes intermittent problems reproducible : Recorded samples do not change between runs, so a failure that happened once now happens every time.
  • Capture at the moment of failure : The recording can then be analysed offline, repeatedly, with no device present.
  • One rule chooses the tool : A message in the trace means work above the physical layer, and nothing in the trace means drop to the samples.

What do the common failure signatures mean?

Bring-up failures repeat. The same handful of symptoms account for most of the time lost, and each one points at a small set of causes. Having the mapping written down turns a long session into a short one.

The table below is that mapping. Each row is a symptom that can be observed without guessing, followed by what usually produces it. Work down the list from the top, because the rows follow the ladder of the earlier section.

Symptom

Most likely causes

First thing to check

Nothing visible on a spectrum view

Transmit gain at zero, wrong frequency, the cell never started, or the stack exited quietly.

That the process is still running, then the configured frequency against what you are watching.

Signal present, device does not list the cell

Wrong band or channel number, network identity the device rejects, or system information that fails to decode.

Decode your own broadcast with a second receiver before blaming the device.

Device lists the cell, no preamble ever detected

Uplink gain too low, preamble configuration mismatch, or a duplex pattern that leaves no uplink opportunity.

The uplink noise floor first, then the duplex pattern against the preamble occasions.

Preamble detected, the device retries the same one

The response was not transmitted, was transmitted late, or carried a timing advance the device rejected.

Whether the response left at all, and whether it left inside its window.

Response sent, no third message received

The uplink grant is unusable, the timing advance is wrong, or the offset to the uplink slot is tighter than the device allows.

The k2 value against the device processing time, from the real-time loop page.

Connection sets up, then releases within seconds

Core network unreachable, authentication failure, or a configuration the device could not apply.

The core network interface trace, since the radio has already proved itself.

Attached, no user data passes

No session established, user plane tunnel misrouted, or the flow mapped to no bearer.

Whether packets reach the user plane function at all, in either direction.

Data passes, throughput far below expectation

Conservative link adaptation, scheduling starvation, or late slots on the host machine.

The late slot counter, then grants and throughput per device separately.

One pattern runs through the whole table and is worth extracting. Almost every row is resolved by checking your own side before investigating the device. The device is the harder thing to inspect, so it should be the last suspect rather than the first.

A second pattern concerns the two rows in the middle. Both random access failures are commonly timing rather than radio, which is exactly the confusion the real-time loop page predicts. A device that is scheduled inside its own processing time behaves identically to one that could not decode.

  • The same handful of symptoms recur : Having the mapping written down converts a long debugging session into a short one.
  • Decode your own broadcast first : A second receiver reading your system information settles most cell detection failures immediately.
  • No preamble usually means uplink or duplex : Check the uplink noise floor, then whether the duplex pattern leaves any preamble opportunity.
  • Random access failures are often timing : A device scheduled inside its processing time looks exactly like one that could not decode.
  • A quick release points at the core : Reaching connection setup already proved the radio, so the remaining suspects are elsewhere.
  • Check your own side first : The device is the hardest thing to inspect, so it should be the last suspect rather than the first.

How do you test without a real device?

Debugging against a commercial handset is slow, because the handset explains nothing. Several ways exist to replace it, and each removes a different set of variables. Using them in the right order saves a great deal of time.

The most valuable is the virtual radio, which the overview page mentions. It replaces the RF hardware with a transport between two processes, so a base station and a device stack talk to each other with no radio involved.

What that removes is everything physical. No antennas, no gain settings, no interference, no clock drift and no timing between separate boxes. Whatever fails afterwards is unambiguously a software problem.

The second is a device simulator that you control. Unlike a handset it can be instrumented, made to log its reasoning, and modified to test a specific case. It is also the only way to create a hundred devices without buying a hundred handsets.

Both share one limitation that has to be stated clearly. A simulator built from the same understanding as the base station will agree with the base station about anything both misread. Two implementations sharing an author share their misconceptions.

That is why a commercial device remains necessary at some point. It was built independently from the same specification, so it disagrees exactly where a reading was wrong. Interoperability testing is the only thing that finds those cases.

Between the two sits channel emulation, which adds impairments deliberately. Delay, fading, noise and frequency offset can be applied to a virtual link. That tests the receiver against conditions a laboratory bench never produces.

A sensible progression follows from this. Virtual radio first for functional correctness, then channel emulation for robustness, then a real device over the air for conformance. Each step adds one class of variable rather than all of them at once.

  • The virtual radio is the most valuable tool : It replaces the hardware with a transport between processes, so anything that fails afterwards is software.
  • It removes every physical variable : No antennas, gain, interference, clock drift or inter-box timing remain in the experiment.
  • A simulator can be instrumented : It logs its own reasoning and can be modified, which a commercial handset never does.
  • Two implementations by one author share misconceptions : A simulator agrees with the base station wherever both misread the specification.
  • Only an independent device finds those : It was built from the same document by someone else, so it disagrees exactly where you were wrong.
  • Add variables one class at a time : Virtual radio for correctness, channel emulation for robustness, then a real device over the air.

Why does one working device prove so little?

Getting one device to attach and pass traffic feels like the end of the work. It is closer to the beginning. Several important mechanisms are simply not exercised by a single device on a strong signal.

The scheduler is the clearest case. With one device there is no allocation decision to make, because everything goes to that device. Fairness, starvation and priority handling are all untested, and those are most of what a scheduler does.

Contention resolution is untested for the same reason. Random access includes a mechanism for two devices choosing the same preamble, and that mechanism never runs when only one device exists. It is also difficult to trigger deliberately.

Link adaptation is barely exercised on a strong signal. The outer loop from the scheduler page settles at a fixed offset and stays there. Its correction behaviour, and the retransmission path it depends on, only appear when the channel degrades.

The identifier and context handling of the upper layer page also stays hidden. Allocation, reuse and release of device identities cannot fail visibly when only one identity is ever in use.

Timing has the same property, and it is the one that catches people out after deployment. Per-slot cost grows with the number of active devices. A stack with comfortable margin at one device can be missing deadlines at ten.

Time itself is the last untested variable. Leaks, counter wraps and drift all need hours to appear. A demonstration lasting ten minutes proves nothing about a system that has to run for a week.

So the useful next steps after one device are specific. Add a second device, degrade the signal deliberately, and leave it running overnight. Those three tests exercise most of what a single short demonstration leaves untouched.

  • One device means no scheduling decision : Fairness, starvation and priority are untested, and they are most of what a scheduler does.
  • Contention resolution never runs : The mechanism for two devices choosing the same preamble needs two devices to exist.
  • A strong signal hides link adaptation : The outer loop settles and stays there, so its correction behaviour is never exercised.
  • Identifier handling stays hidden : Allocation, reuse and release cannot fail visibly while only one identity is ever in use.
  • Timing margin shrinks with device count : A stack that is comfortable at one device can be missing deadlines at ten.
  • Leaks need hours : A ten minute demonstration proves nothing about a system expected to run for a week.
  • Three tests cover most of the gap : Add a second device, degrade the signal deliberately, and leave it running overnight.

What discipline actually finds bugs here?

The techniques above are tools. What determines whether a session makes progress is method, and the method that works here is stricter than in most software. The reason is that observation is expensive.

Change one thing between runs. That advice appears everywhere and it matters more here than usual, because a run takes minutes and produces ambiguous evidence. Changing two settings means the next result cannot distinguish between them.

Keep the configuration in version control alongside the code. A RAN configuration has hundreds of parameters, and the difference between a working and a broken system is often one of them. Reconstructing yesterday's configuration from memory is not possible.

Record the samples during any failure worth understanding. The recording page explains what metadata to keep, and the value here is that a recording turns a live problem into an offline one. Offline problems can be examined as many times as needed.

Reproduce before fixing. A fault seen once and then repaired leaves you unable to tell whether the repair worked or the fault simply did not recur. That distinction matters more than it seems, because intermittent faults return at the worst moment.

Read the counters before reading the code. Late slots, decode failures, retransmissions and released contexts are all counted already. Those numbers narrow the problem faster than any amount of reasoning about what should happen.

Question the timing early rather than late. The real-time loop page describes failures that present as protocol errors, and they are the ones that consume whole days. Checking the late slot counter takes seconds and rules out a large class of cause.

Finally, be willing to go down a layer. Most people stop at the message trace, because it is comfortable and readable. The problems that take longest to solve are the ones where the trace is correct and the samples are not.

  • Change one thing between runs : A run takes minutes and gives ambiguous evidence, so two changes make the next result unreadable.
  • Version control the configuration : Hundreds of parameters, and the difference between working and broken is often exactly one of them.
  • Record the samples during a failure : A recording turns a live problem into an offline one that can be examined repeatedly.
  • Reproduce before fixing : A fault seen once and repaired leaves you unable to tell whether the repair worked.
  • Read the counters before the code : Late slots, decode failures and released contexts are already counted, and they narrow the problem fastest.
  • Check the timing early : Failures that present as protocol errors consume whole days, and the counter takes seconds to read.
  • Be willing to go down a layer : The longest problems are the ones where the message trace is correct and the samples are not.

Reference

The list below is where the procedures and the tooling come from. The specifications define what should happen, and the project documentation describes the traces and simulators used to find out what did.