20.20.0 (2026-10-08)
Features
- harbor: prompt hill-climb task on a text-to-SQL dataset (#16458) (bd78550)
- mcp/sql: allowlist the annotation, evaluator, prompt, split and job tables (#16788) (a37b841)
- mcp: add GraphQL schema, query, and mutation tools (#16828) (9351580)
- playground: add claude-haiku-5-5 and remove retired Anthropic models (#16870) (b978e7c)
- pxi: make PXI's MCP code mode configurable for Harbor benchmarks (#16845) (7b43f9b)
- sandbox: add Docker Sandboxes provider (#16538) (a3966b2)
Bug Fixes
- agents: upgrade pydantic-ai to 2.52 and drop the Anthropic max_tokens override (#16765) (bea333f)
- build: keep harbor-verifiers out of the root project's dependencies (#16806) (239b843)
- cost: update built-in model token prices (#16692) (20b9159)
- cost: update built-in model token prices (#16835) (91ea995)
- cost: update built-in model token prices (#16868) (41c0537)
- deps: update arize-phoenix-evals to 3.9.1 (#16866) (808a068)
- evaluators: ignore stale input mappings (#16789) (4458a37)
- experiments: measure base and compare runs the same way in run metric comparisons (#16761) (3fc54c3)
- mcp: improve analytics SQL execution and dialect guidance (#16879) (cb4aef5)
- Remove SciPy and scikit-learn from runtime dependencies (#16795) (ee5ca59)