Arid 2.0.0 — Cleaner Contracts, Better Automation
Arid 2.0 is a major contract, automation, and integration release built around one principle:
Cleaner contracts. Better automation. Same focused detector.
The detector itself is intentionally not redesigned. Arid 2.0 continues to find exact normalized Python duplicate code and report DUP001; the major changes are in machine-readable contracts, CI/agent workflows, baseline maintenance, project control, and the supported Rust API.
For ordinary CLI users who do not consume Arid's machine contracts or Rust internals, upgrading from 1.2 may require no migration work.
See the Arid v2 migration guide for the complete compatibility and migration details.
Highlights
Report schema v4
JSON reports now use report schema v4.
The principal changes are:
- top-level
versionis renamed toschema_version - required
tool_version - required
complete - required resolved
analysismetadata - required structured
errors - required path-independent finding
fingerprint - occurrence distribution
mixedis renamed tohybrid
Structural context and scope continue to use mixed when a finding genuinely spans multiple structural categories.
The current schema is published at:
schemas/report-v4.schema.json
The historical report-v3 schema remains published unchanged.
Stable finding identity
Every v4 finding includes a versioned fingerprint:
arid-finding-v1:sha256:...
The fingerprint identifies normalized duplicate content independently of path, physical line number, occurrence order/multiplicity, structural metadata, output format, and worker mode.
SARIF 2.1.0 output exposes the same identity through:
partialFingerprints["aridFindingFingerprint/v1"]
Focused reporting with whole-project detection
Repeatable --focus <PATH> limits which duplicate groups are reported without narrowing what code is compared.
Arid still performs whole-corpus detection, then baseline enforcement, then focus filtering. A focused finding retains all occurrences, including occurrences outside the focused path.
arid . --focus src/packageKeep-going analysis
--keep-going allows independent source read/parse/normalization failures to be collected while valid files continue through detection.
Incomplete analysis is never reported as successful:
- report-v4 uses
complete: false - structured source errors are included in
errors - process exit status remains
2 - incomplete reports cannot be emitted as SARIF
arid . --keep-going --jsonBaseline lifecycle operations
Baseline schema v1 remains unchanged and existing baseline files remain valid.
V2 adds:
arid . --baseline-status arid-baseline.json
arid . --prune-baseline arid-baseline.json--baseline-status classifies accepted, active/new, and stale duplicate debt.
--prune-baseline removes stale acceptance only; it never silently accepts new debt.
Explicit project and configuration control
V2 makes project context inspectable and explicit when needed:
arid . --config path/to/pyproject.toml
arid . --no-config
arid . --project-root path/to/project
arid . --show-config
arid . --list-filesLegacy nearest-config behavior remains available when no explicit selector is supplied.
Virtual Python source
--stdin-path <PATH> analyzes exactly one virtual Python source supplied through standard input using the same parser and normalizer as disk-backed source.
cat src/example.py | arid . --stdin-path src/example.pyAn equivalent disk path is replaced for the scan; otherwise the virtual source is added when allowed by the resolved project context. Arid never writes the virtual source to disk.
Multiple outputs from one scan
Repeatable --report FORMAT=PATH writes supplemental text, JSON, Markdown, and SARIF outputs from one in-memory report without reparsing or redetecting:
arid . \
--format text \
--report json=artifacts/arid.json \
--report markdown=artifacts/arid.md \
--report sarif=artifacts/arid.sarifFindings-only exit policy
--no-fail-on-findings maps a complete findings-only exit 1 to success 0.
Operational or incomplete status 2 is never masked.
Capability discovery
arid --capabilities emits deterministic build capability JSON without requiring project discovery.
Its machine contract is published at:
schemas/capabilities-v1.schema.json
Fatal JSON-mode operational errors use the contract at:
schemas/error-v1.schema.json
Official GitHub Action
Arid now ships an official composite GitHub Action.
- uses: sponge-b0b/arid@v2.0.0
with:
paths: .The Action installs the exact released Arid version encoded by the tagged Action metadata, runs one Arid scan, exposes core metrics as outputs, can write a job summary, and can produce SARIF when configured.
Narrow supported Rust API
Arid remains primarily a CLI application. V2 deliberately narrows the supported Rust surface to the crate-root application API:
use arid::{
Cli,
ColorEnvironment,
ExitStatus,
RunContext,
RunResult,
run,
run_with_context,
};Implementation modules and detector/report internals are no longer semver-supported public API.
Compatibility with Arid 1.2
V2 preserves:
- exact normalized duplicate semantics
DUP001- normal CLI invocation and existing option names
[tool.arid]- normalization behavior
- source suppression through
# arid: disable/# arid: enable - baseline schema v1 and existing baseline files
- serial default execution
- numeric
--workers N --workers auto- default exit meanings
0/1/2 - pre-commit integration
- supported release platforms
The intentional migration surfaces are concentrated in report JSON, SARIF identity, and the supported Rust API. See the migration guide for concrete before/after examples.
Validation
V2 was exercised through targeted contract validation and a real-world campaign covering Black, Django, mypy, Rich, Unicode/space paths, and feature-composition workflows.
For equivalent settings, canonical duplicate groups from v2 were compared directly against qualified Arid 1.2.0 across Black, Django, mypy, and Rich with no detector-semantic regression found.
Real-world validation also covered:
- file and directory focus while preserving whole-corpus context
- baseline enforcement before focus filtering
- virtual-source replacement without disk mutation
- keep-going with a controlled malformed source
- large-corpus multi-output on Django
- worker determinism
- the published official GitHub Action
Performance
The v2 performance campaign uses the same pinned benchmark corpora and Hyperfine methodology as the qualified Arid 1.2 campaign.
Against current stable Pylint 4.0.6, serial Arid v2 measured:
Requests: 191.19x faster
Pydantic: 219.06x faster
Polaris: 249.68x faster
The qualifying medium and large corpora therefore remain far above Arid's 10x performance floor versus isolated Pylint duplicate detection.
A paired, reversed-order v1.2→v2 regression investigation measured the serial v2 overhead at low single digits across the canonical corpora:
Requests: +1.3% default / +0.6% Pylint-compatible
Pydantic: +0.3% default / +1.4% Pylint-compatible
Polaris: +2.3% default / +3.1% Pylint-compatible
duplicate-heavy: +2.0% default / +1.3% Pylint-compatible
No product optimization was justified by the qualification evidence.
See Arid v2 performance report for methodology and detailed results.
Install
With uv:
uv tool install "arid==2.0.0"
arid --versionOr with pip:
python -m pip install "arid==2.0.0"
arid --versionExpected output:
arid 2.0.0
Machine-readable contracts
Published Arid-owned schemas are under schemas/:
report-v4.schema.jsonerror-v1.schema.jsoncapabilities-v1.schema.jsonbaseline-v1.schema.json
report-v3.schema.json remains published as the historical v1 report contract.
SARIF remains SARIF 2.1.0 and uses the official SARIF schema.
Migration
Start with the Arid v2 migration guide.
The shortest version is:
- CLI-only users may need no change.
- JSON consumers must migrate report v3 → v4.
- occurrence distribution
mixedbecomeshybrid; structuralmixedremains. - findings gain stable fingerprints.
- SARIF gains versioned Arid finding identity.
- baseline-v1 files remain valid.
- Rust embedding code must use the supported crate-root application API.
Arid 2.0 keeps the detector focused and familiar while making its contracts and automation substantially more useful to developers, CI systems, coding agents, and external tooling.