github sponge-b0b/arid v2.0.0

latest releases: v2.2.3, v2.2.2, v2.2.1...
one month ago

Arid 2.0.0 — Cleaner Contracts, Better Automation

Arid 2.0 is a major contract, automation, and integration release built around one principle:

Cleaner contracts. Better automation. Same focused detector.

The detector itself is intentionally not redesigned. Arid 2.0 continues to find exact normalized Python duplicate code and report DUP001; the major changes are in machine-readable contracts, CI/agent workflows, baseline maintenance, project control, and the supported Rust API.

For ordinary CLI users who do not consume Arid's machine contracts or Rust internals, upgrading from 1.2 may require no migration work.

See the Arid v2 migration guide for the complete compatibility and migration details.

Highlights

Report schema v4

JSON reports now use report schema v4.

The principal changes are:

  • top-level version is renamed to schema_version
  • required tool_version
  • required complete
  • required resolved analysis metadata
  • required structured errors
  • required path-independent finding fingerprint
  • occurrence distribution mixed is renamed to hybrid

Structural context and scope continue to use mixed when a finding genuinely spans multiple structural categories.

The current schema is published at:

schemas/report-v4.schema.json

The historical report-v3 schema remains published unchanged.

Stable finding identity

Every v4 finding includes a versioned fingerprint:

arid-finding-v1:sha256:...

The fingerprint identifies normalized duplicate content independently of path, physical line number, occurrence order/multiplicity, structural metadata, output format, and worker mode.

SARIF 2.1.0 output exposes the same identity through:

partialFingerprints["aridFindingFingerprint/v1"]

Focused reporting with whole-project detection

Repeatable --focus <PATH> limits which duplicate groups are reported without narrowing what code is compared.

Arid still performs whole-corpus detection, then baseline enforcement, then focus filtering. A focused finding retains all occurrences, including occurrences outside the focused path.

arid . --focus src/package

Keep-going analysis

--keep-going allows independent source read/parse/normalization failures to be collected while valid files continue through detection.

Incomplete analysis is never reported as successful:

  • report-v4 uses complete: false
  • structured source errors are included in errors
  • process exit status remains 2
  • incomplete reports cannot be emitted as SARIF
arid . --keep-going --json

Baseline lifecycle operations

Baseline schema v1 remains unchanged and existing baseline files remain valid.

V2 adds:

arid . --baseline-status arid-baseline.json
arid . --prune-baseline arid-baseline.json

--baseline-status classifies accepted, active/new, and stale duplicate debt.

--prune-baseline removes stale acceptance only; it never silently accepts new debt.

Explicit project and configuration control

V2 makes project context inspectable and explicit when needed:

arid . --config path/to/pyproject.toml
arid . --no-config
arid . --project-root path/to/project
arid . --show-config
arid . --list-files

Legacy nearest-config behavior remains available when no explicit selector is supplied.

Virtual Python source

--stdin-path <PATH> analyzes exactly one virtual Python source supplied through standard input using the same parser and normalizer as disk-backed source.

cat src/example.py | arid . --stdin-path src/example.py

An equivalent disk path is replaced for the scan; otherwise the virtual source is added when allowed by the resolved project context. Arid never writes the virtual source to disk.

Multiple outputs from one scan

Repeatable --report FORMAT=PATH writes supplemental text, JSON, Markdown, and SARIF outputs from one in-memory report without reparsing or redetecting:

arid . \
  --format text \
  --report json=artifacts/arid.json \
  --report markdown=artifacts/arid.md \
  --report sarif=artifacts/arid.sarif

Findings-only exit policy

--no-fail-on-findings maps a complete findings-only exit 1 to success 0.

Operational or incomplete status 2 is never masked.

Capability discovery

arid --capabilities emits deterministic build capability JSON without requiring project discovery.

Its machine contract is published at:

schemas/capabilities-v1.schema.json

Fatal JSON-mode operational errors use the contract at:

schemas/error-v1.schema.json

Official GitHub Action

Arid now ships an official composite GitHub Action.

- uses: sponge-b0b/arid@v2.0.0
  with:
    paths: .

The Action installs the exact released Arid version encoded by the tagged Action metadata, runs one Arid scan, exposes core metrics as outputs, can write a job summary, and can produce SARIF when configured.

Narrow supported Rust API

Arid remains primarily a CLI application. V2 deliberately narrows the supported Rust surface to the crate-root application API:

use arid::{
    Cli,
    ColorEnvironment,
    ExitStatus,
    RunContext,
    RunResult,
    run,
    run_with_context,
};

Implementation modules and detector/report internals are no longer semver-supported public API.

Compatibility with Arid 1.2

V2 preserves:

  • exact normalized duplicate semantics
  • DUP001
  • normal CLI invocation and existing option names
  • [tool.arid]
  • normalization behavior
  • source suppression through # arid: disable / # arid: enable
  • baseline schema v1 and existing baseline files
  • serial default execution
  • numeric --workers N
  • --workers auto
  • default exit meanings 0 / 1 / 2
  • pre-commit integration
  • supported release platforms

The intentional migration surfaces are concentrated in report JSON, SARIF identity, and the supported Rust API. See the migration guide for concrete before/after examples.

Validation

V2 was exercised through targeted contract validation and a real-world campaign covering Black, Django, mypy, Rich, Unicode/space paths, and feature-composition workflows.

For equivalent settings, canonical duplicate groups from v2 were compared directly against qualified Arid 1.2.0 across Black, Django, mypy, and Rich with no detector-semantic regression found.

Real-world validation also covered:

  • file and directory focus while preserving whole-corpus context
  • baseline enforcement before focus filtering
  • virtual-source replacement without disk mutation
  • keep-going with a controlled malformed source
  • large-corpus multi-output on Django
  • worker determinism
  • the published official GitHub Action

Performance

The v2 performance campaign uses the same pinned benchmark corpora and Hyperfine methodology as the qualified Arid 1.2 campaign.

Against current stable Pylint 4.0.6, serial Arid v2 measured:

Requests:  191.19x faster
Pydantic:  219.06x faster
Polaris:   249.68x faster

The qualifying medium and large corpora therefore remain far above Arid's 10x performance floor versus isolated Pylint duplicate detection.

A paired, reversed-order v1.2→v2 regression investigation measured the serial v2 overhead at low single digits across the canonical corpora:

Requests:         +1.3% default / +0.6% Pylint-compatible
Pydantic:         +0.3% default / +1.4% Pylint-compatible
Polaris:          +2.3% default / +3.1% Pylint-compatible
duplicate-heavy:  +2.0% default / +1.3% Pylint-compatible

No product optimization was justified by the qualification evidence.

See Arid v2 performance report for methodology and detailed results.

Install

With uv:

uv tool install "arid==2.0.0"
arid --version

Or with pip:

python -m pip install "arid==2.0.0"
arid --version

Expected output:

arid 2.0.0

Machine-readable contracts

Published Arid-owned schemas are under schemas/:

  • report-v4.schema.json
  • error-v1.schema.json
  • capabilities-v1.schema.json
  • baseline-v1.schema.json

report-v3.schema.json remains published as the historical v1 report contract.

SARIF remains SARIF 2.1.0 and uses the official SARIF schema.

Migration

Start with the Arid v2 migration guide.

The shortest version is:

  • CLI-only users may need no change.
  • JSON consumers must migrate report v3 → v4.
  • occurrence distribution mixed becomes hybrid; structural mixed remains.
  • findings gain stable fingerprints.
  • SARIF gains versioned Arid finding identity.
  • baseline-v1 files remain valid.
  • Rust embedding code must use the supported crate-root application API.

Arid 2.0 keeps the detector focused and familiar while making its contracts and automation substantially more useful to developers, CI systems, coding agents, and external tooling.

Don't miss a new arid release

NewReleases is sending notifications on new releases.