First PyPI release plan¶
Target and current status¶
The first public release includes the reader, dataset path configuration,
and an explicitly invoked LOCATA downloader. The code license is
Apache-2.0. Keep the distribution name locata-torch, import name locata_torch,
and CLI name locata-torch.
Reviewed on 2026-10-07. The reader, root lookup, explicit downloader, CLI,
Apache-2.0 license, package metadata, and validation/publication workflows are
implemented. 0.1.0 is published on
PyPI following the successful 0.1.0rc1 TestPyPI rehearsal. Both OIDC uploads and
protected-environment approvals completed. Public wheel/sdist hashes matched
validated CI artifacts. Independent pip and uv consumers passed, including
existing-data integration and spawn workers. The repository is
taishi-n/torchlocata, and the PyPI owner is taishi-n. The
validation record records executed checks and the unverified
final-release corpus payloads. Release notes describe the release.
The getting-started guide gives the pip/uv installation and
usage instructions for the published release.
Ideas from existing tools¶
These comparisons inform the proposed design; no external reader or downloader code has been copied. Rolling documentation was checked on 2026-10-07.
| Reference | Observed pattern | Proposed use in locata-torch |
|---|---|---|
| Torchvision CIFAR10, 0.29 | Explicit root, opt-in download=False, and reuse of downloaded data |
Keep familiar explicit paths and reuse; expose downloading as a separate operation before creating workers |
| TensorFlow Datasets, API page updated 2024-04-26 | data_dir, TFDS_DATA_DIR, versioned datasets, and separate preparation and reading steps |
Provide an environment setting, a stable dataset release identity, and a preparation step; keep Dataset construction read-only |
| Hugging Face Hub, current cache guide | Shared local storage with revision-specific snapshots | Reuse one downloaded release across projects and virtual environments; a simple release directory is sufficient for two archives |
| Pooch, 1.9.0 | File registries with known hashes, environment-controlled storage, and retries | Pin official URLs, sizes, and checksums; use a dataset release ID rather than the Python package version as the storage key |
The upstream code licenses checked were Torchvision: BSD-3-Clause, TensorFlow Datasets: Apache-2.0, Hugging Face Hub: Apache-2.0, and Pooch: BSD-3-Clause. Pooch's license is version-pinned; the others are upstream branch snapshots checked on the review date, not licenses assigned to this library.
Pooch is a useful alternative for ordinary whole-file downloads. Its documented HTTP downloader does not promise persistent partial-file resume. The proposed first implementation uses focused standard-library HTTP, hashing, and ZIP code to control resume and installation, rather than assuming Pooch supplies those guarantees. No generic download backend or corpus plugin interface is needed.
Add platformdirs
as a small runtime dependency for platform-standard persistent storage. Keep
cross-process locking with filelock
in the optional download extra. HTTP, hashing, ZIP64 extraction, and the CLI can
use urllib.request, hashlib, zipfile, and argparse. Current upstream
platformdirs
and filelock
license files state MIT. The minimum platformdirs 4.3.8 and resolved 4.12.3 are
MIT. Filelock 3.20.0, the minimum, is Unlicense; resolved 4.0.12 is MIT. Reader
use requires no download extra; filelock is imported by the downloader only.
Dataset path contract¶
Lookup for LocataDataset(root=None, ...):
- An explicit
root, includingstrorPath. LOCATA_ROOT, pointing directly to an unpacked root containingdev/oreval/.- The managed final-release root beneath the configured storage directory.
An explicitly supplied or environment-supplied invalid path must raise a useful
error; it must not silently fall back. An empty environment value is invalid.
Expand ~, resolve the selected path once in the constructor, and retain the
absolute path for pickling and spawned workers. Do not search the working
directory or guess a user's existing corpus location.
Managed storage has a separate meaning and lookup:
data_dirpassed todownload_locata, or CLI--data-dir.LOCATA_DATA_DIR.platformdirs.user_data_path("locata-torch", appauthor=False, ensure_exists=False).
LOCATA_ROOT never redirects a download into an existing corpus. The downloader
returns its actual installed root, which can be passed explicitly to the reader.
An existing challenge-era root remains usable without a managed manifest.
When using the default managed root, requested splits must have completed
installation records. Missing managed data produces an error with the download
command, without creating directories or performing network I/O. User-provided
roots retain the reader's existing WAV-based selection behavior.
For an existing shared dataset:
For another Python tool:
uv add "locata-torch[download]"
export LOCATA_DATA_DIR=/path/to/locata-store
uv run locata-torch download --split dev
uv run locata-torch path
from locata_torch import LocataDataset, download_locata
root = download_locata(split="dev", data_dir="/path/to/locata-store")
dataset = LocataDataset(root=root, split="dev", arrays=("eigenmike",))
Reader-only consumers install locata-torch without an extra.
download_locata(split="dev", *, data_dir=None) -> Path accepts "dev", "eval",
or a nonempty sequence of those names; it returns the common root after all
requested splits succeed. CLI --split may be repeated. locata-torch path
prints the resolved reader root to stdout and reports errors to stderr. Calling
the downloader without its extra gives an installation instruction. There is no
implicit download option on the Dataset constructor.
Distribution and storage¶
Download only the pinned official final release,
v1 dated 2020-01-31, DOI 10.5281/zenodo.3630471. Its
metadata API supplies these byte sizes,
MD5 checksums, and HTTPS file URLs:
| Split | Archive bytes | Official MD5 |
|---|---|---|
| dev | 6,207,354,195 | d5a5417c3f6b2ed0e43581dd06f504e6 |
| eval | 13,045,105,034 | 46709713350bc16c106e788920d30a8b |
Bounded HTTP range reads of the official ZIP64 end records and central directories
confirmed the dev/ and eval/ top-level paths. Their declared uncompressed
sizes were 27,136,519,350 and 57,426,914,435 bytes. Keeping both ZIPs and extracted
splits needs about 103.8 GB before filesystem overhead; recommend at least 120 GB
free for a fresh installation. These are directory-metadata observations, not
validation of audio or annotation contents. Full archives were not downloaded.
The observed file lists include VAD in both splits, unlike the local challenge snapshot. This release uses the existing snapshot for reader integration and local HTTP/synthetic ZIP fixtures for downloader validation. No real dataset download is performed during verification. Final-release audio/VAD payload validation remains a documented limitation rather than a publication gate. Downloading transfers an entire split; reader task/array filters do not reduce network transfer.
Storage, entirely outside a user-provided corpus:
<data_dir>/
├── archives/zenodo-3630471/ # Verified dev.zip/eval.zip and resumable .part files
├── datasets/zenodo-3630471/ # Returned LOCATA root, containing complete splits
├── state/zenodo-3630471/ # Per-split installation and attribution manifests
├── staging/ # Incomplete extraction, outside the readable root
└── locks/ # Per-split process locks
The release key is independent of the library version, so upgrading the library does not duplicate the corpus. Downloaded archives are retained for offline re-extraction. The initial API does not delete data or provide automatic repair. Manifests record the DOI, exact URLs, expected and observed checksums, sizes, installation state, inventory, and dataset citation/license information. The reader writes no indexes, caches, locks, or manifests beneath a dataset root.
The dataset metadata identifies ODC-BY 1.0. Keep dataset attribution and license notices in the download record and documentation, separate from Apache-2.0 for this code. Fetch official archives directly; do not bundle dataset assets in wheels or maintain a redistributed mirror.
Download and extraction requirements¶
- Validate the split and destination before network activity. Check available space using pinned size information, then the verified ZIP directory before extraction. Acquire one lock per split and install multiple splits in a fixed order. Use finite timeouts, bounded retries, and progress reporting to stderr.
- Stream to a persistent
.partfile with bounded memory, recording its release, URL, and expected hash. Request identity encoding so offsets refer to archive bytes. Resume with HTTPRangeonly after checking206, the exactContent-Range, and total size. A server returning200must cause a safe restart of the owned partial file, never append a whole response. Validate truncated/oversized responses and handle416, rate limits, and interrupted transfers explicitly. Range support was observed, but must be checked on every response; it is not guaranteed. - Rehash retained bytes when resuming. Verify the final file's exact size and official MD5 before renaming it to a verified archive. MD5 is the publisher's integrity check; HTTPS and the pinned source remain necessary. A mismatch must fail before extraction and leave the installed dataset untouched.
- Use Python's ZIP64-capable zipfile with streamed member reads and CRC validation. Before writing, reject absolute paths, traversal, symlinks, unexpected top-level paths, and duplicate or case-colliding destinations. Preserve file contents; do not repair clocks, preprocess WAVs, or remove optional files.
- Extract into a new owned staging directory on the destination filesystem. Publish a complete split with an atomic rename, then commit its manifest atomically. Record the prepared state first so a crash between these operations can be recovered by verifying the inventory without overwriting data. Failed extraction remains invisible to the default reader; retries can reuse the verified archive.
- Repeated requests reuse completed managed installations after inventory checks. Unknown existing directories, altered installations, or incompatible manifests produce an error. Never merge into or overwrite an existing dataset, and do not silently redownload, replace, or delete user files. Successful splits remain usable if a later requested split fails.
Implementation sequence and acceptance criteria¶
Each implementation step follows specification, failing tests, minimal code, and checks. Update the README and API reference as proposed features become real.
| Step | Deliverable | Acceptance criteria |
|---|---|---|
| 1. Paths and licensing | Apache-2.0 LICENSE and SPDX metadata; optional root; storage resolution; path CLI |
Tests cover explicit/env/default priority, invalid and empty settings, ~, no implicit I/O, and consistent paths under spawn. The base install works without download dependencies |
| 2. Downloading | Pinned archive registry; public download function and CLI; streaming, resume, hashes, locks | Synthetic HTTP-server tests cover interrupted transfers, correct and ignored ranges, bad range/size/hash, HTTP failures, retries, reuse, and concurrent calls. Tests make no external requests |
| 3. Installation | ZIP64 extraction, space checks, staging, manifests, and crash recovery | Tiny forced-ZIP64 fixtures cover corrupt members, unsafe paths, collisions, insufficient space, interrupted extraction, existing directories, and byte-for-byte preservation. No partial split becomes readable |
| 4. Existing-data integration | Read-only verification of the available dev/eval snapshot | Read representative static, multiple-source, moving-source, and moving-array recordings; verify partial reads, clocks, collation, missing files, and spawn. Use local fixtures for transport/ZIP/VAD tests; do not download official archives. Clearly record unverified final-release payloads |
| 5. Distribution and automation | Complete package metadata and URLs; tested wheel/sdist; validation and publication workflows | Install built artifacts outside the checkout in fresh base and download-extra environments; test the CLI, import, sample reading, and spawn. Run supported-Python checks, lint, format, ty, build, strict Zensical, and link checks; inspect artifacts and README rendering |
| 6. Release | TestPyPI rehearsal, then first PyPI publication | Check version/tag agreement and the recorded validation boundaries. Publish only the artifact set that passed validation; confirm an independent consumer can install and use the published package |
Keep routine CI offline with small WAV/TXT/HTTP/ZIP fixtures. Extend the current Python 3.10/3.12 checks with supported-version and minimum-dependency coverage; test platform-specific paths, locking, and spawn on Linux, macOS, and Windows. The configured matrix covers Python 3.10–3.14, macOS/Windows spawn and locking, and Python 3.10 with minimum direct dependencies. CPU PyTorch builds avoid CUDA dependencies on Linux test runners. Locked quality/build checks run on macOS. No official archives are downloaded in this release's verification or CI. Synthetic VAD tests do not count as real final-release VAD validation.
PyPI release operation¶
Package metadata identifies taishi-n and the repository, issue tracker, and
hosted documentation URL; keep dataset citation distinct from authorship. Retain
py.typed, declare the download extra and CLI entry point, and include LICENSE
in wheel and sdist. Recheck name availability immediately before setup: PyPI's
JSON endpoint returned 404 for locata-torch on 2026-10-07, which does not reserve
the name.
Follow the official uv packaging guide
and Python packaging metadata guide.
Use uv build --no-sources, inspect both artifacts, and verify installation
without checkout-local imports. For a TestPyPI rehearsal, install dependencies
from their normal index and the exact test artifact with --no-deps, avoiding a
mixed-index dependency search.
Use prerelease versions for TestPyPI rehearsals and a stable v0.1.0 tag for
the first PyPI release. Route prerelease and stable tags explicitly; do not let
every v* tag publish to PyPI. Use a tag-triggered validation job and a separate
publish job that consumes its already tested artifacts. Configure
PyPI Trusted Publishing
with GitHub OIDC, id-token: write only in the publish job, and a protected pypi
environment. A manual validation run should not publish. For the first project,
a pending publisher
requires the actual repository owner/name, workflow filename, and environment.
TestPyPI needs its own configuration. Avoid storing a long-lived upload token.
Account configuration¶
The public repository taishi-n/torchlocata and local origin now exist. Package
environments pypi/testpypi are restricted to v* tags and require review by
taishi-n; github-pages allows v* tags and main for explicit documentation
updates. GitHub Pages is enabled with source
GitHub Actions and HTTPS. The site inherits the account's existing custom
domain, making its canonical URL https://taishi.org/torchlocata/. The stable
release workflow deployed it successfully; overview, tutorial, and API pages
returned HTTP 200. No account-wide domain setting was changed.
The first release registered pending publishers independently on PyPI and TestPyPI with these values:
| Field | Value |
|---|---|
| Project name | locata-torch |
| Owner | taishi-n |
| Repository | torchlocata |
| Workflow filename | release.yml |
| Environment | pypi on PyPI; testpypi on TestPyPI |
PyPI ownership is under taishi-n. TestPyPI is a separate service and requires
its own account and publisher configuration; a PyPI account does not create it.
No upload token is stored in this repository. See the official
new-project publisher instructions.
Release steps¶
- Finish local validation and record results. Commit reviewed changes and push
main; require all remote CI jobs to pass. - Set the package version to
0.1.0rc1, update the lockfile/release notes, and tag the matching commitv0.1.0rc1. The release workflow reruns the full CI, builds one artifact set, inspects metadata/contents/hashes, and exercises fresh base, download-extra, and sdist installations. Only those validated artifacts reach TestPyPI after environment approval. - Rehearse a consumer outside this checkout. Install normal dependencies from PyPI, then install the exact TestPyPI version without dependency resolution:
python -m pip install torch numpy soundfile platformdirs filelock
python -m pip install --index-url https://test.pypi.org/simple/ --no-deps "locata-torch[download]==0.1.0rc1"
locata-torch --help
Point it at existing data using LOCATA_ROOT=/path/to/LOCATA and run the
recording/window/spawn examples. Do not fetch the corpus during rehearsal.
4. Set the stable version to 0.1.0, update release records, and create the
matching v0.1.0 tag. The separately validated stable artifact set goes to
PyPI after environment approval; the same built documentation is deployed to
Pages after publication succeeds. Do not rebuild distributions in publish jobs.
5. Verify PyPI's exact version and wheel hash, install it in an independent
environment, and run the documented existing-data workflow. Mark publication
complete only after these checks. Commit the completed validation record and
use the manual docs.yml workflow to deploy the updated documentation from
main; this does not rebuild or republish package artifacts.
scripts/check_release.py rejects version/tag mismatches, unsupported prerelease
types, corpus/generated assets, missing licenses or typing metadata, and incorrect
dependencies. It records SHA-256 for the wheel/sdist; the publish job checks those
hashes again. A manual workflow run validates without publishing. Account setup,
remote execution, and public installation must be recorded as pending until
actually verified.