Changelog — v2.0.0 (14 Sep 2026)
New skills — collection grows to 36
- Third Party Risk Management: vendor tiering, framework-mapped due-diligence questionnaires, SOC 2 Type II report review (exceptions, CUECs, carve-outs, bridge letters), DPA/sub-processor review, DORA Art. 30 contract addenda, OAuth-aware monitoring and offboarding — 96% vs 64% (+32)
- 🇬🇧 Cyber Essentials / CE Plus: built on the current Danzell question set (v3.3, April 2026) — auto-fail rules (cloud MFA, 14-day updates), scoping trees, CE+ audits, pricing, bundled insurance, PPN 014/MoD mandates — 80% vs 16% (+64); new UK region filter
- Sarbanes Oxley ITGC: the four ITGC domains, top-down scoping, RCMs and test scripts, the AS 2201/Reg S-X deficiency ladder, filer/EGC/IPO logic — 92% vs 52% (+40)
October pre-cycle — five skills updated for September go-lives (primary-source verified):
EU CRA (Art. 14 reporting live Sept 11 + full ENISA SRP filing workflow), FedRAMP (precise Rev5 machine-readable mandate: Sept 30 intake cutoff, annual-assessment trigger), CMMC (verified Task Force status; advisory-only legal chain), SWIFT CSP (corrected v2026 split 26+6 and architecture taxonomy; v2027-unconfirmed honesty), DORA (CTPP EU-subsidiary deadline ~Nov 2026; third-country branches).
Benchmark — 180 test cases / 902 assertions across 36 skills: 89% vs 57% (+32 pts, +284).
Five re-runs on harder September-event assertion sets dropped honestly (SWIFT 96→56, FedRAMP 96→64) while deltas widened — baselines collapse on post-cutoff events. Zero grading amendments; one eval failure exposed and fixed a pre-existing SWIFT skill error (rerun-2026-10/ERRATA.md).
Engineering — eval page now fully generated from grading artifacts (36 uniform accordions, 180 transcripts; 88 legacy eval dirs backfilled); bump_version.py one-command releases (used for this one); repo hygiene cleanup.