deciding-with-confidence
Answers yes/no, multiple-choice and rubric-score questions about a piece of text
with a probability for every option and a calibrated confidence, in the request
and response shapes of OpenAI's Decisions API. Claude Haiku 5.5 does the
judging; the skill estimates the probabilities that a purpose-built decision
model would read off its logits.
echo '{"input": "I was charged twice.", "questions": [{"type": "choice", "name": "dept",
"instructions": "Which department?", "choices": [{"value": "billing"}, {"value": "shipping"}]}]}' \
| python3 scripts/decide.py run -It runs through the Anthropic API (pip install anthropic and a key), the
claude CLI, Amazon Bedrock, or parallel subagents inside Claude Code: the
deciding-with-confidence plugin installs a decider agent pinned to Haiku
5.5 alongside the skill.
On a 104-item eval it reached 0.846 accuracy against 0.894 for a purpose-built
decision model, with matching calibration error. Its wrong answers come at low
confidence, so thresholding on confidence and on agreement between samples
catches most of them. See SKILL.md for the procedure and
references/method.md for the measurements.
assets/eval.jsonl includes 64 items from the banking77 test split (PolyAI,
CC-BY-4.0); each item names its source.
Skill folder: deciding-with-confidence
Release of deciding-with-confidence version 0.1.1
📥 Download & Install
⬇️ Download deciding-with-confidence.zip
To install:
- Click the download link above (ignore the "Source code" archives below - they're auto-generated by GitHub)
- Go to Claude.ai Skills Settings
- Upload the downloaded ZIP file
- Requires paid Claude Pro or Team account
See official documentation for more details.
Recent Changes
eb772c5 deciding-with-confidence 0.1.1: correct the miss-confidence and escalation claims