github craft-ai-agents/craft-agents-oss v0.14.0

4 hours ago

Features

  • Decision model (opt-in) — Settings → AI now has a Decision model card. It connects a small, fast model that answers typed questions (pick one of these, score this, yes or no) with probabilities in about a quarter of a second, instead of writing text. Use Jev through TypeSafe, OpenRouter or Vercel AI Gateway with your own key, reuse the key of an existing OpenRouter or Vercel AI Gateway connection, or run Laya on your own machine for free (laya-serve, with a built-in health check). It is off by default, and nothing is sent anywhere until you turn it on. Each feature below has its own switch under the card's Advanced settings, grouped into Agent, Conversations, Permissions, and Automations and tasks, with a tooltip saying what the model is asked. A decision can add a prompt, a hint or a skip, but it never approves anything, and when the model is unsure, slow or unreachable the app behaves exactly as before. Every call is recorded in ~/.craft-agent/logs/decisions.jsonl with a hash of the input, never the input itself.

  • Guarded permission mode — A new mode between Ask to Edit and Execute. The agent works as in Execute, but Bash commands, MCP writes and non-GET API calls are checked first, and you are asked before anything that is hard to undo, leaves the project or reaches other people and services. File writes outside the working directory and the session's plans and data folders always ask, without a model call. It is offered while the decision model and its Guarded mode switch are on; if the check can no longer run, a Guarded session behaves like Ask to Edit. Execute is unchanged and never consults the decision model, and a plan approved from Explore returns the session to Guarded if it came from there.

  • The agent can make quick judgments with decide — While the decision model is on, the agent has a decide tool for classifying, routing or scoring up to 200 items at once, returning probabilities rather than prose, in a fraction of the time and cost of a language model call.

  • Smarter conversations — With the matching switches on:

    • Adaptive thinking lowers the thinking level for simple turns ("thanks!", a quick question), never above the level you set.
    • Mid-turn messages decide whether a message you send while the agent works corrects the running task (steered in) or is a separate follow-up (queued), and merge a message you split over two sends into one request.
    • Smart titles skip small talk ("hey 👋") and title the session when you ask for something, then refresh the title when the conversation moves on. A title you set yourself is never replaced.
    • Suggestions point the agent to the skill or inactive source your message needs, such as your calendar source for "find a free slot on Saturday". Nothing is enabled on the model's behalf.
    • Turn outcome moves a session to Needs Review when the agent ends its turn waiting for you (a question, a decision, a blocker), not when it just offers more help. In tasks, a subtask that asks for input or gives up counts as failed.
    • Risk badges show what a permission prompt would do (deletes, sends, publishes, credentials, system, spends).
    • Large results skip the multi-second summary of a big tool result when its beginning and the saved file are enough.
  • Semantic automations and labels — Label auto-rules can ask a question instead of matching a pattern ({ "semantic": "Is the user planning a trip?" }, or craft-agent label auto-rule-add --semantic), and automations can carry a condition (semanticCondition) that must hold before they run. Skipped runs are recorded as skipped in the automation history. For tasks, a verifier reply without a VERDICT: line is read by the decision model before the app asks again, and a failed run repairs only the subtasks the reply blames. Inferred verdicts show their confidence in the task editor.

Improvements

  • Ask to Edit runs read-only MCP tools without prompting — Fixes #236. Ask to Edit prompted for every MCP call, even list_* tools that Explore mode runs silently. It now applies the same read-only rules as Explore, including the patterns in your workspace and source permissions.json; writes still ask. Thanks to @siyant for the request.

  • Token Optimization update prompt — rtk versions before 0.44.0 corrupted the output the agent read (grep lines with a colon, wc reading from <, crashes on piped output). Older versions are no longer used; while Token Optimization is on, a dialog explains what breaks and offers the update command, and Settings → AI → Performance shows the outdated state. Commands starting with command or builtin are never rewritten, rtkExcludeCommands in config.json is honoured, and the agent is told how to get exact output.

  • "Always Allow" is hidden when it would remember nothing — Prompts from Guarded mode and commands without a safe key no longer show a button that had no effect.

  • Leaner app bundle — Two unused MCP server bundles left over from the Codex and Copilot backends were removed.

Bug Fixes

  • "Always Allow" remembers exactly what you approved — It remembered the first word of the command, so approving one git command also allowed git push --force, git reset --hard or git commit -m x && rm -rf ~ for the rest of the session. It now remembers the subcommand (git commit, aws s3 ls), the exact command for interpreters and runners (python x.py, npm run build), or the hosts for curl and wget. "Always Allow" on file writes and API calls, which silently did nothing, now works.

  • Explore mode's read-only checks are tighter — MCP tools such as delete_account, send_thread_reply or update_spreadsheet ran in Explore because read verbs were matched anywhere in the name, and gh api -X DELETE, gh api -f (an implicit POST), sed -n -i, sed -n 'w file' and sort -o passed as read-only. All of these are now blocked in Explore.

  • spawn_session cannot start a more permissive session — The mode came from the model, so an Ask to Edit session could spawn an Execute session without asking. A spawned session now gets at most its parent's mode. On Pi connections, a required prompt no longer becomes an allow after a source is activated, and a tool call that needs a prompt is denied when no one can answer it.

  • Plans and sign-ins no longer read as "declined" — Partially addresses #1004. Submitting a plan or starting a source sign-in interrupted the tool while it was still running, so the transcript recorded "The user doesn't want to proceed with this tool use": the plan step showed as failed, and after an unanswered sign-in the agent told you that you had declined it. The tool's real result is now recorded first. Thanks to @eruditavanitas for the report.

  • Long-running tools no longer look like a loop — Fixes #1008. A tool call that ran for minutes added a new "Running …" row every 30 seconds, which read as the agent calling the tool again and again. The heartbeats now only update the running tool's elapsed time. Thanks to @SzeJim for the report.

  • Large MCP results are summarized before the agent reads them — Only API sources had big results saved to a file and summarized for the model. Results from MCP servers and session tools reached the model in full, and a summary was generated afterwards only for the transcript. They are now handled before the model sees them, and no summary is made for results it has already read.

  • Source activation continues the turn — Activating a source mid-turn stopped the session when the turn had no text (for example an attachment-only message). The turn now continues with the previous message. The automatic retry after an activation no longer appears as a duplicate message, keeps corrections you sent during the turn, and keeps the turn's thinking level.

  • No more phantom "background task finished" prompts — An idle session could be woken with "the background agent you launched has finished" for tasks it never launched, and Bash tasks showed 0 s in list_background_tasks. Background tasks are now attributed from the SDK's structured events.

  • Stop no longer re-runs your message — Pressing Stop in a Claude session could send the same message again instead of stopping.

  • Thinking level changes apply mid-session — On Claude, changing the thinking level in an ongoing session took effect only after the session restarted; it now applies from the next turn.

  • Quick successive sends no longer fail — The session file was rewritten on every message, and two messages sent quickly could fail with a file error.

  • Automations cannot trigger themselves — An automation that labelled its own session could fire again for that label. Automations now never run for an event about a session they created, and chains of automation-created sessions stop at three levels.

  • source_test no longer reports a source without a credential as connected — A source marked authenticated whose token was missing passed the check; it now tries a refresh and otherwise asks you to sign in again.

  • CLI — craft-agent automation keeps an automation's semanticCondition, conditions and Telegram topic when editing, and new task labels always get valid ids.

Don't miss a new craft-agents-oss release

NewReleases is sending notifications on new releases.