github dathere/qsv 24.0.0

pre-release3 hours ago

[24.0.0] - 2026-10-15 ⚖️ The "Evidence Act" Release 🏛️

Highlights

~160 commits since 23.0.1, and the theme is EVIDENCE. The 2019 Foundations for Evidence-Based Policymaking Act asks US federal agencies to build policy on evidence, to make their data open by default and machine-readable, and to describe all of it in a public data inventory. Seven years on, the hard part is rarely the law. It is the plumbing: the evidence sits in SAS, Stata and SPSS files, and the inventory has to pass a validator. This release works on both.

The inventory first. Metadata is what makes data FAIR (Findable, Accessible, Interoperable and Reusable), and an inventory entry only does that job if it validates: a catalog record data.gov cannot harvest leaves the dataset unfindable, however open the data itself is. The metadata data.gov harvests is DCAT-US, and DCAT-US v3 quietly stopped being JSON-LD this spring. profile was still emitting the prefixed JSON-LD form, and data.gov's own validator rejected it outright. It now emits canonical, unprefixed v3, and --validate actually bites. JSON Schema treats format as an annotation by default, so a malformed date used to pass locally and then fail at data.gov; formats are now asserted. Unknown keys are reported, --strict aborts only on Required findings, and the validator is tested against GSA's own 23 conformance fixtures. The F and the I in FAIR start with metadata that passes.

Then the evidence itself. Program evaluations, surveys and statistical microdata, the raw material of the Act's learning agendas and its statistical agencies, live largely in SAS, Stata and SPSS files. 23.0.1 brought readstat to read them. 24.0.0 makes it a round trip:

  • The new writestat command writes a CSV back out as SPSS (.sav, .por), Stata (.dta) or SAS transport (.xpt).
  • readstat --dictionary writes a JSON Schema data dictionary from the file's own metadata (variable labels, value labels, declared missing values, display formats) with no LLM involved. qsv validate and viz smart read it like any other dictionary.
  • Hand that dictionary to writestat and the file comes back with its metadata: decoded labels return to their codes, and missing-value codes become declared missing values again. Round trips are identical for every SPSS .sav fixture and for the SAS transport fixtures.
  • Where a format cannot hold something, writestat refuses rather than silently dropping it. --lossy writes the file anyway and warns about each thing it left out.

In evidence, an empty cell is not one thing. "Refused", "Don't know" and "Not applicable" are different answers, and SAS, Stata and SPSS keep them apart as user-defined missing values. readstat now preserves them for all three formats with --sentinels-as value|label, labels SAS's from .sas7bcat format catalogs, and, when you don't ask it to keep them, no longer drops them silently: it reports how many were written as empty cells and where. Checked against upstream's 578-file test corpus, its counts match pyreadstat's exactly. readstat can also read just part of a large file (--select, --offset, --limit, --sample). Thanks to @jrothbaum for a run of fast upstream fixes in polars-readstat-rs that made much of this possible.

Evidence also needs numbers you can defend. moarstats gains 17 statistical measures (56 → 73), each checked against scipy. Among them are L-moments, lag-1 autocorrelation, the Hoover index, Cramér's V, OLS regression and Benford MAD, Nigrini's first-digit test of Benford's law, a classic data-fabrication signal. The same scrutiny found that Jarque-Bera, the Bimodality Coefficient and kurtosis were computed from the wrong inputs; all three now match scipy, so their values change. frequency --lmt-threshold had applied its limits to the wrong columns ever since the option was added, and weighted ties and the "Other" row were mishandled; all fixed. stats and frequency no longer invoke undefined behavior on malformed CSV, and stats caches no longer leak a private file's values through world-readable permissions. Evidence you cannot reproduce is anecdata.

Machine-readable by default applies to tools too. Every command now describes itself: <command> --help --format json|md prints its JSON tool definition or help Markdown, and --export-tool-definitions writes them for every installed command. The definitions are generated by the binary from its own usage text, so they always match the version and feature set you are running. Agents and pipelines get the same truth humans do.

And it's faster. frequency is 1.5× faster by default (2× with --weight) with lower peak memory; stats --infer-dates is 1.43× and sniff 1.81× faster, after we removed regex lock contention in qsv-dateparser and csv-nose. The un-indexed count regression we flagged in 23.0.1 is fixed (#4603): the row count is multithreaded again. viz --geojson census:state adds automatic US state boundaries (USPS codes, FIPS, or full state names), and new Linux ARM64 musl and Windows ARM64 GNU prebuilts round out the platforms.

From the statistician's .sav file to the inventory entry on data.gov, qsv now carries the evidence and the metadata that makes it evidence. That metadata is what makes data FAIR, and Findable, Accessible, Interoperable & Reusable (FAIR) Data is AI-Ready Data: the same labels, codes and declared missing values that keep a statistician honest keep an AI agent from guessing. Evidence that is FAIR serves Humans, Agents and Machines alike, as we march towards Accelerated Civic Intelligence (ACI).

Important

The x86_64 prebuilts now require a CPU with AVX2 (x86-64-v3). Older Intel Pentium/Celeron/Atom parts and VMs with a conservative CPU model (kvm64/qemu64) will die with Illegal instruction, and qsv --update would install a binary that cannot run, or update itself back. Check with grep -o avx2 /proc/cpuinfo | head -1 before updating. qsvdp and the aarch64, ppc64le and s390x prebuilts are unaffected. See Breaking Changes.

Don't miss a new qsv release

NewReleases is sending notifications on new releases.