A capture on disk is the last buffer in the chain. It is also the only one that outlives the session, which makes its failures the most expensive. I have thrown away recordings that were technically perfect, because nobody wrote down the sample rate. Each section below covers one part of getting a capture onto disk and keeping it useful afterwards.
- Executive Summary
- Why is recording a rate problem before a capacity problem?
- What is actually inside an IQ file?
- Why is a raw file without metadata worthless?
- What does SigMF standardise?
- How do you sustain the write without dropping?
- Can IQ be compressed?
- How do you make a recording that is still useful next year?
- Reference
Executive Summary
The table is a lookup. Sustained write rate is what matters for a long capture, and it is not the number on the box.
Medium |
Sustained write |
61.44 Msps, 246 MB/s |
200 Msps, 800 MB/s |
What to watch |
|---|---|---|---|---|
Hard disk |
Drive- and workload-dependent; measure long sequential writes and service gaps. |
Requires more than 245.76 MB/s sustained with stall margin; do not assume it. |
No. |
Seeks and competing work can reduce throughput. Preallocation can help but does not guarantee the rate. |
SATA SSD |
Model-dependent after write-cache exhaustion; the SATA link also limits throughput. |
Possible only when the measured sustained rate and stalls meet the requirement. |
No. The interface itself is the limit. |
Cache capacity and post-cache rate vary with the model, free space and workload. |
NVMe, PCIe Gen3 x4 |
PCIe generation describes the link, not the drive's sustained write rate. |
Check the actual device and filesystem under sustained load. |
Yes, provided the sustained figure is checked. |
Measure the drive after the cache is exhausted, not in the first ten seconds. |
NVMe, PCIe Gen4 x4 |
Peak transfer figures do not establish long-duration write performance. |
Verify sustained rate and stall margin on the installed system. |
Multiply by channel count and verify sustained service for the total rate. |
Drive behaviour, temperature, link sharing, filesystem and host processing can each limit recording. |
Capacity, for reference |
200 Msps at 16-bit IQ is 2.88 TB per hour. |
885 GB per hour. |
2.88 TB per hour. |
Decimation reduces rate when the retained bandwidth permits it; compression savings depend on the data. |
Why is recording a rate problem before a capacity problem?
A drive can have enough space for an hour of samples and still lose data during the first minute. Capacity determines how long recording can continue; sustained write service determines whether the stream survives.
For uncompressed data, calculate R = fs * C * B. The terms are complex samples per second per channel, channel count and bytes per complex sample. Signed 16-bit I plus signed 16-bit Q occupy four bytes. One channel at 61.44 MS/s therefore produces 245.76 MB/s, or 884.736 GB per hour.
At 200 MS/s, the same representation produces 800 MB/s and 2.88 TB per hour. These are decimal units and payload-only calculations. Complex float32 doubles both figures. Filesystem space, metadata, other applications and the chosen safety reserve also affect usable recording duration.
Consider a writer that normally keeps up but stops for 50 ms. At 800 MB/s, that pause needs 40 MB of reachable free buffering. If the writer resumes at exactly 800 MB/s, it cannot remove the backlog while input continues. It needs extra throughput before another pause occurs.
Short storage tests can measure RAM caching or an SSD's fast write region rather than long-run performance. Run the actual recorder long enough to reveal slower operation, including temperature and available-space effects. Record queue occupancy and service gaps as well as average throughput.
A storage interface specifies a possible transfer path, not a guaranteed sustained drive rate. The particular device, filesystem and workload decide whether either example rate is practical.
Budget rate and duration separately. A large disk does not compensate for a sustained write deficit.Measure stalls as well as averages. The writer needs buffering and recovery throughput for temporary interruptions.
What is actually inside an IQ file?
An IQ file may contain nothing except numbers written one after another. Before interpreting a spectrum, establish how those bytes become samples; a plausible-looking plot does not prove that the decoder chose correctly.
Specify the component type, bit width, byte order and component order. A common single-channel layout stores I0, Q0, I1, Q1 and so on. But a raw file has no universal convention. Packed converter words, unsigned bytes and planar channel layouts require different readers.
The diagram below gives an explicit example: two signed 16-bit complex samples in little-endian order. Sample zero is 1 - j, and sample one is 2 - 2j. Each pair occupies four bytes, regardless of the converter's effective resolution.
Figure 1. Sample interpretation depends on both the component layout and byte order.
- The blue boxes contain in-phase components; the orange boxes contain quadrature components.
- The byte sequence occupies eight bytes for two complex samples.
A 12-bit ADC can deliver values in a 16-bit storage word. Record the alignment and scaling rather than assuming that every stored bit represents independent converter information. Similarly, floating-point host samples may already include driver scaling. UHD distinguishes host and over-the-wire formats, so a transport format does not uniquely identify the file format.
For a headerless, one-channel file with four-byte complex samples, file_size / 4 gives the stored sample count. A remainder indicates truncation or a different layout. Multiple channels require their ordering and count too. Check a known test pattern before trusting a large capture.
Describe stored bytes explicitly. A filename ending in .iq does not define a format.Separate storage width from ADC resolution. Scaling and packing determine what the numeric values mean.
Why is a raw file without metadata worthless?
The bytes may survive while the measurement becomes impossible to reproduce. A raw recording can still reveal some signal structure, but missing acquisition details can prevent reliable timing, frequency or level measurements.
The sample rate defines both the time axis and the baseband frequency axis. Centre frequency places that baseband on the RF spectrum. Without those values, a peak's bin number does not identify its RF frequency, and a burst's sample count does not establish its duration.
Amplitude needs additional context. Record gain settings, AGC state, antenna path, attenuation and any calibration. AGC does not automatically make level recovery impossible, but unrecorded gain changes can make it ambiguous. Clipping destroys information that a later calibration cannot restore.
Time needs a reference point as well. State whether a timestamp identifies the first sample, a host receive call or a file creation event. Only the first directly describes capture time, and its accuracy still depends on the radio clock. Record reference lock and any known timing uncertainty.
Keep changes tied to sample positions. A retune, gain change or loss of samples may happen halfway through a file. One value in a text filename cannot describe that history. For a discontinuity, preserve the first valid timestamp after the gap and the missing interval when known. Do not silently join the two sides into an apparently continuous recording.
Also retain the purpose of the measurement and the receiver configuration. A technically readable file may still be unsuitable for a later analysis if its bandwidth or calibration assumptions are unknown.
Metadata defines the measurement axes. Sample rate, RF frequency and scaling turn numbers into interpretable observations.Record changes where they happen. Time-indexed configuration and gap information preserve the acquisition history.
What does SigMF standardise?
A private text file can document a capture, but another tool may not understand its field names. SigMF gives recordings a shared description so software can agree on basic interpretation without guessing.
The core arrangement pairs a binary .sigmf-data file with a JSON .sigmf-meta file. The metadata contains a global object, a captures array and an annotations array. Global fields describe the dataset; capture entries describe acquisition segments; annotations identify regions of interest.
For example,
The specification checked here is SigMF v1.2.6. It also defines extension namespaces and Collections for relating recordings. Use documented extensions for additional structured information rather than inventing fields inside the reserved core namespace.
Format validation and measurement validation remain different jobs. A metadata document can be syntactically correct while claiming the wrong rate. Compare its values with the actual receiver configuration, file length and a known signal. A field's presence is not evidence that its value was measured correctly.
Before committing to a workflow, make a short capture and open it in an independent reader. Check sample values, duration and frequency placement. Then test how that reader handles the specific metadata and extensions required by the experiment. Keep a small known-good example with the recording software so future changes can be checked against it.
A shared format reduces interpretation errors. It cannot correct inaccurate acquisition metadata.Test the intended reader. Verify the complete capture workflow before producing a large dataset.
How do you sustain the write without dropping?
A receive thread should not wait for every disk operation to finish. Separate acquisition from storage with a bounded queue, then measure whether the writer drains that queue fast enough during the full capture.
The diagram below separates three events: receiving samples, accepting a write and completing the required durability operation. Buffered writes may complete while data remains in the operating system's page cache. Device caching adds another layer.
Figure 2. A completed application write and a durable recording are separate conditions.
- The bounded ring absorbs finite storage stalls and exposes a measurable occupancy.
- The writer must handle errors and finish pending operations before reporting a completed recording.
Use appropriately sized sequential writes and reusable buffers. Preallocation can reduce allocation work, but does not guarantee sustained speed or contiguous physical placement. Handle short writes and asynchronous failures. Do not reuse an outstanding write buffer before the API permits it.
Buffered and direct access have different tradeoffs. Direct access may reduce page-cache effects but has platform-specific alignment and support constraints. It does not by itself promise power-loss durability. Choose an explicit flush policy, check its results, and include its cost in testing.
Use fio or an equivalent tool on a dedicated test file to characterize the intended workload. Match write size, concurrency and buffering policy. Then run the real recorder, because conversion, metadata handling and other processing share the host. Test long enough to expose storage cache exhaustion and thermal changes.
Keep recording queues bounded. On overload, stop cleanly or record the loss explicitly.Define what completed means. Flush and error-handling requirements belong in the recording procedure.
Can IQ be compressed?
Some captures compress well, while others remain almost the same size. The deciding factor is redundancy in the stored samples, so a compression ratio measured on one waveform is not a storage guarantee.
Lossless compression reproduces the original stored bytes after decompression. Noise-like samples with heavily populated low bits may offer little redundancy. Unused storage bits, repeated patterns or strongly correlated samples can offer more. Test representative captures, including the least compressible conditions expected during the experiment.
Online compression also consumes compute and memory bandwidth. At an input rate of 800 MB/s, the compressor must continuously accept that rate, regardless of its output size. A ratio of 0.6 output bytes per input byte would reduce storage traffic to 480 MB/s. A ratio near 1 provides little relief, and format overhead can make some outputs slightly larger.
Reducing component width is different from lossless compression. Converting signed 16-bit I and Q into 8-bit components halves payload size, but requires a scaling rule and can add quantization error or clipping. Record the conversion so later readers do not interpret reduced precision as the original measurement.
Decimation removes samples after suitable filtering. It can reduce rate substantially when the wanted signal occupies a smaller bandwidth, but it discards information outside that retained band. Choose it from the measurement requirement, not only from the available disk capacity.
Chunked compression can support bounded working memory and partial reads. Preserve each chunk's sample range and verification information. Document the codec and container, and confirm that the intended analysis software can recover the original sample sequence.
Plan for poor compression. A recorder must survive realistic worst-case input rather than a favourable average ratio.Separate compression from information reduction. Quantization and decimation change what the recording preserves.
How do you make a recording that is still useful next year?
A useful archive must preserve more than a readable file. It needs enough context to repeat the interpretation, identify missing data and establish which processing steps produced the samples now on disk.
Give each recording a stable identifier and retain its metadata beside it. Store the receiver model, software version, clock configuration and relevant processing settings. Record calibration references and any conversion, filtering or decimation applied before writing. Keep derived files distinct from the original acquisition.
Validate file length against the stored format and sample count. For a continuous single-channel capture, duration equals sample_count / sample_rate. Compare that with timestamps, allowing for explicitly recorded gaps. A correctly sized file still does not prove that its contents are continuous or unclipped.
Compute a checksum after finalizing the dataset and verify it after copying. A checksum detects byte changes; it cannot prove that acquisition lost no samples. Preserve the acquisition error log and a clear complete, partial or interrupted status as separate evidence.
Large captures are easier to manage as documented segments. Cut at sample boundaries and give every segment its starting sample position and timing context. Keep a manifest that orders the segments and associates their metadata. Do not rely on directory listing order to reconstruct the stream.
Finally, reopen a copy using the intended analysis tool. Inspect a known signal, timing and channel order. Keep a small readback example with the archive, and verify restored copies periodically. Reproducing a short analysis from the saved files is stronger evidence than merely seeing that a large file exists.
Verify bytes and measurement history separately. Checksums and acquisition logs answer different questions.Test readback before the session is forgotten. Missing context is easiest to recover while the equipment and configuration remain available.
Reference
The list below is where the format details come from. The specification is the authority for anything about SigMF, and the ShareTechnote pages carry the stages that produce the samples being recorded.
- SigMF Specification : The Signal Metadata Format - the data and metadata file pair, the global, captures and annotations objects, and the core namespace fields.
- sigmf.org : SigMF project - the overview, tooling and the datatype naming convention such as cf32_le.
- SigMF Python Library : Reading and writing the pair from Python, including archive handling.
- UHD Configuring Devices and Streamers : Ettus Research - host and transport sample formats and format conversion.
- fio Documentation : Sequential write workloads, buffered/direct access and synchronization options. Relevant options were consulted rather than the complete manual.
- ShareTechnote - Memory and Buffering : The buffer chain this page extends onto disk, and the sizing rule used for the writer ring.
- ShareTechnote - Where the Processing Runs : The CPU budget that compression would have to come out of.
- ShareTechnote - ADC and DAC : What 8-bit against 16-bit actually costs in dynamic range.
- ShareTechnote - RF Front End : Gain, AGC and calibration context needed for interpreting recorded amplitude.
- ShareTechnote - Syncrhonization : Absolute time for a capture, and the lock status worth recording beside it.