Skip to content

I/O and DataLoader

Recording discovery and errors

Dataset construction reads directory names, array WAV headers, and the existence of mandatory files. It does not load waveform payloads or sample timestamp tables. Existing WAV files determine the recording index; ZIP archives, results, and MATLAB code are outside it.

An existing selected array directory without its array WAV is excluded and reported in missing_audio with MissingAudioWarning. The reader does not assume every recording has every array, or fabricate warnings for absent directories.

Condition Result
Missing array WAV in an existing array directory Warning and exclusion from the index
Invalid header or missing mandatory file for an existing WAV LocataError during construction
Corrupt TXT rows, invalid values, or mismatched row counts LocataError on first relevant access
Missing optional annotation None, independently of split
Malformed optional file that is read LocataError, including path and cause

Mandatory files are the array WAV, its audio timestamp TXT, array-position TXT, and required_time.txt. Available source files are handled independently. Source WAV and clock content validation is deferred when source audio is disabled and no source VAD needs that clock. Existing malformed data is never silently dropped or reported as an absent annotation.

Partial reads and bounded worker caches

Windows seek within the WAV and read only the requested frames using SoundFile. The reader never loads a whole recording merely to slice a window.

Audio timestamps and VAD are streamed and validated in full on first access in each worker. The reader records file offsets every 4096 rows, then reads near the requested interval. It preserves the actual clock values, including nonuniform sampling, without replacing them with an unverified uniform approximation. Smaller pose and required-time tables are read per item and cropped by time.

Each worker has an LRU limit of cache_size=16 indexed files and cache_bytes=16777216 estimated index bytes by default. An index exceeding the limit is used for the current access without being cached. Eviction, or setting either limit to zero, causes full validation to run again on subsequent access. Temporary indexing memory and returned samples are separate from this cache budget.

The dataset retains no open file handles. Pickling and process changes reset the cache. It writes no disk cache or index under the corpus. Input files must remain unchanged during the lifetime of a dataset instance.

Collation

collate_locata(samples) accepts a nonempty batch with the same array, sample rate, channel count, and waveform dtype. Mixed batches raise a clear ValueError.

Batch field Representation
waveform [B, C, T_max], padded with zeros on the right
lengths [B], int64, original waveform lengths
audio_mask [B, T_max], bool, true only for original audio frames
sample_rate One common int in Hz
audio_time, time_origin Lists with one original entry per sample
required_time, array_pose Lists of original typed annotations
sources, metadata Lists retaining each sample's source mapping and identity

audio_mask=False describes padding only. It does not describe silence, source activity, or valid_flag. Annotation lengths and source counts may vary; the collator does not interpolate, pad ground truth, or create dense labels.

Datasets return CPU tensors. Transfer waveforms explicitly, for example batch["waveform"].to(device), after loading a batch.

Multiple workers and samplers

Both dataset views and the top-level collate function are pickle-compatible. They support standard Subset, SequentialSampler, RandomSampler, and DataLoader shuffle.

Use multiprocessing_context="spawn" for multiple workers, and place loader creation and iteration behind a main guard:

from locata_torch import LocataDataset, collate_locata
from torch.utils.data import DataLoader


def main():
    dataset = LocataDataset(root="/path/to/LOCATA", arrays=("eigenmike",))
    windows = dataset.windows(num_samples=48000, hop_samples=24000)
    loader = DataLoader(
        windows,
        batch_size=4,
        shuffle=True,
        num_workers=2,
        multiprocessing_context="spawn",
        collate_fn=collate_locata,
    )
    batch = next(iter(loader))
    print(batch["waveform"].shape)


if __name__ == "__main__":
    main()

An executable example is available at examples/read_locata.py in the repository. Worker-zero and explicit-spawn checks are recorded in validation. See the PyTorch DataLoader documentation for the upstream multiprocessing contract.