github codewhale-hq/Codewhale v0.10.1

3 hours ago

Codewhale is the public product from Shannon Labs. The codewhale
command, npm package, and release-asset names remain lowercase technical
identifiers. The legacy npm package deepseek-tui is deprecated and
receives no further releases. Users coming from v0.8.x legacy deepseek /
deepseek-tui names should migrate with docs/REBRAND.md.

Install

Recommended — official GitHub release

New macOS/Linux install (checksummed binaries from this release):

curl -fsSL https://codewhale.net/install.sh | CODEWHALE_VERSION="v0.10.1" sh
"$HOME/.local/bin/codewhale" --version

For Windows, use the matching installer or archive below. For an existing
direct install, run codewhale update; it prints the executable path and
keeps newer builds. If the install directory is occupied by a different build,
use the fresh-directory migration in the installation guide.

Secondary packaging — npm and Cargo

npm install -g codewhale
# or build from source
cargo install codewhale-cli --locked

The wrapper downloads the matched codewhale and codew command assets
from this Release. Both contain the same compiled runtime.

Docker / GHCR

docker run --rm -it \
  -e DEEPSEEK_API_KEY="$DEEPSEEK_API_KEY" \
  -v codewhale-home:/home/codewhale/.codewhale \
  ghcr.io/codewhale-hq/codewhale:v0.10.1

The image exposes the same runtime as both codewhale and codew. The
latest tag is also updated on release.

Cargo (Linux / macOS)

cargo install codewhale-cli --locked

The Cargo package installs codewhale. Cargo cannot create a second command
alias from one binary target; users who want the shorter spelling can add a
codew symlink to that installed executable. The npm, Homebrew, archive,
shell-installer, and container channels install both command names directly.

Manual download — platform archives (recommended)

Each archive below contains the same runtime under the codewhale and
codew command names, plus an install script:

Platform Archive Install script
Linux x64 codewhale-linux-x64.tar.gz install.sh
Linux ARM64 codewhale-linux-arm64.tar.gz install.sh
Android ARM64 (Termux) codewhale-android-arm64.tar.gz install.sh
macOS x64 codewhale-macos-x64.tar.gz install.sh
macOS ARM codewhale-macos-arm64.tar.gz install.sh
Windows x64 (installer) CodeWhaleSetup.exe NSIS setup
Windows x64 codewhale-windows-x64.zip install.bat
Windows x64 (portable) codewhale-windows-x64-portable.zip —
Windows ARM64 codewhale-windows-arm64.zip install.bat
Windows ARM64 (portable) codewhale-windows-arm64-portable.zip —

Unix (Linux / macOS):

tar xzf codewhale-<platform>.tar.gz
cd codewhale-<platform>
./install.sh

Windows:

  • For the installer path, run CodeWhaleSetup.exe; it installs
    codewhale.exe, codew.exe, and codewhale.bat under
    %LOCALAPPDATA%\Programs\CodeWhale\bin, adds that directory to the
    current-user PATH, and creates a Start Menu shortcut that prefers
    Windows Terminal (wt.exe) when it is installed.
  • Extract the archive for your machine: codewhale-windows-x64.zip or
    codewhale-windows-arm64.zip
  • Double-click codewhale.bat (not the raw .exe) to launch
  • Run install.bat to copy the binaries and launcher to %USERPROFILE%\bin
  • Add %USERPROFILE%\bin to your PATH

The portable Windows archive skips the install script — extract and run codewhale.bat from any directory. The NSIS installer is currently unsigned and may trigger Windows SmartScreen until a signing certificate is wired into the release pipeline.

Each platform also has bare, unarchived codewhale-<platform> and
codew-<platform> assets. The seven codewhale-tui-<platform> filenames
attached to v0.9.5 are byte-identical compatibility copies used only to let
already-installed v0.9.4 clients discover and cross this single-binary
transition; current installers do not expose a third runtime. The legacy npm
package deepseek-tui is deprecated and is not republished. For migration
from v0.8.x legacy binary names, see docs/REBRAND.md.

Verify (recommended)

Download the checksum manifests from this Release and verify:

# Linux — archive bundles
sha256sum -c codewhale-bundles-sha256.txt --ignore-missing

# Linux — individual binaries
sha256sum -c codewhale-artifacts-sha256.txt --ignore-missing

# macOS
shasum -a 256 -c codewhale-bundles-sha256.txt --ignore-missing
shasum -a 256 -c codewhale-artifacts-sha256.txt --ignore-missing

What's in v0.10.1

Contributor integration and reliability

  • OAuth retry diagnostics omit provider-controlled error fields; account replacement stops before retired Weixin receipts exceed their storage limit, preserving uncertain work for local review.
  • Resuming a session through a symlink to the same workspace no longer shows a false workspace-change warning.
  • Runtime clients can read one tool call's actual workspace changes and reviewed skill details (thanks @gaord, #6817 and #6869). In-flight snapshot pairs remain pending; missing objects and corrupt repository metadata are distinguished.
  • Search accepts valid preferred locales, and image dimensions describe the same bytes sent to the model (thanks @asto18089, #6860 and #6858). Automation deletion keeps its definition until cleanup succeeds, and compaction preserves its original summary anchor (#6864 and #6857).
  • Config/status/permission commands share portable contracts while the host retains mutation authority; queue workers acknowledge a scheduled retry for temporary first-claim contention and fail honestly on corruption (thanks @aboimpinto, #6832).
  • Indefinite questions, approvals and elevation waits survive the TUI watchdog. Answers get time to resume the current turn; settled requests disappear by identity. Thanks @7jrxt42BxFZo4iAnN4CX for #6872.
  • Configured approval expiry belongs to the held Engine request; hiding or covering its card cannot restart the deadline, and a late queued answer cannot approve an expired call.
  • Tool discovery keeps the highest-ranked matches when a result batch exceeds the existing cache bounds, preserving search order and the 16 KiB limit (adapted from @AdityaVG13's #6393).
  • Model-switch receipts now translate their session-only saving note in every complete locale pack; the three save commands remain directly usable (thanks @Lstarsky0, #6875).
  • Route-save receipts and /workspace replies are translated in every complete locale pack, the Operate descriptions in fourteen packs match the current English, pt-BR, es-419 and ca regain accents six strings had lost, and long CJK text can be shortened between Han and kana characters (thanks @Lstarsky0, #6884, #6885, #6886, #6888, #6887 and #6882).
  • Long sessions keep less in memory: superseded journal entries beyond a retained window move out of the live session into a per-session archive under sessions/.journal-archive/ (kept until the session is deleted; /branch can still restore an archived entry), ended shell operations in the work graph are capped, and session metadata larger than the first read no longer forces a full-file read. Not yet addressed: usage fingerprints in session metadata still grow without a bound, snapshots are still cloned on the UI task, and resume still reads the session file twice (thanks @7jrxt42BxFZo4iAnN4CX, refs #6842).
  • Weixin bridge threads belong to the account and chat that created them:
    resuming or listing another chat's thread is refused. An explicit /new after
    an account replacement keeps the old account's private receipt and never
    replays its prompts automatically. Replies never reuse another bot account's
    Weixin context token, and /threads lists this chat's own older threads even
    when other chats have newer ones.
  • Windows messaging bridges retry a briefly locked record replacement
    (EPERM/EACCES/EBUSY) with bounded delays and never delete the previous record
    first.
  • The website and README installation guides are generated from one shared
    source, and the critical Simplified Chinese guides were re-reviewed against
    current English.
  • A Codewhale sign-in that expired, or an account service that is unreachable,
    no longer blocks turns on your own provider key. The turn uses your local
    settings and Codewhale says once how to restore account preferences.
  • codewhale login keeps a DeepSeek route you chose, or one with a local key,
    instead of switching it to the managed Codewhale provider.
  • The bundled computer-use plugin is 0.12.1, reconciled with canonical source
    724f9c422db1a51880dc81154dae70816c475ed8 while retaining Core's embedding
    manifest and version contract. Its image renderer lock now carries Sharp 0.35.5.
  • The bundled first-party catalog pins marketplace revision
    6512f1dfaa91ee287e9f81ebabaf4909e8a371a3 and lists all 19 reviewed plugins,
    up from 6. Every catalog plugin still installs disabled and untrusted until
    you review it.
  • Every committed npm lockfile is audited in CI, including build tooling. The
    VS Code extension packages with @vscode/vsce 4 on Node 22 while it still
    compiles and tests on its Node 20 runtime. The website keeps one reviewed,
    hash-verified depth guard for an unpatched braces advisory
    (GHSA-vfj7-8cjw-p6xm); its raw audit findings are retained, not hidden.
  • Pasting through Windows Terminal (Ctrl+Shift+V) no longer drops emoji and
    other characters outside the Basic Multilingual Plane. Release binaries and source builds carry
    a small patch to crossterm 0.29.0 in patches/ (the change proposed upstream
    as crossterm-rs/crossterm#1073); a real Windows Terminal paste of 24 Unicode
    lines is now a release gate.
  • Experimental extension host: reviewed Native plugins can register avatar
    packs with ctx.avatars.registerPack (/pet avatar [key], /pet action|view <name>). /plugin import dsh <package-dir> reviews a DeepSeek Harness
    bundle and now lists the Native host code it contains; approve <dir> <content-hash> installs it disabled and untrusted. Trusted host plugins stay
    trusted across a restart, enabling a second host plugin activates it without
    /plugin reload, and the review no longer crashes the TUI.

Codewhale v0.10.1 focuses on reliability and first-run behavior. Turns that
stall now say so, approvals keep what you approved, plugin suggestions are
quieter, and Fleet runs can be checked before they spend anything.
Scripts that parse --json output should read the Breaking for scripts
note below before upgrading.

Added

  • /plugin doctor reports stale built-in records and snapshots, and
    /plugin doctor --fix retires them. A user plugin, a snapshot a running
    process names, and a snapshot inside the grace window are kept. The
    previous state.json is kept as state.json.pre-gc.

  • Reviewed plugins can declare named OpenAI-compatible OAuth routes. The host
    owns PKCE, refresh and credential storage, and checks the review at each
    request (docs/PLUGIN_PROVIDERS.md, #6805).
    The provider capability advances plugin review policy to v5 (v6 with the
    extension host): older receipts require explicit review again.

  • OrcaRouter account sign-in uses PKCE and saves the same durable API key as
    manual setup; its live catalog keeps chat-capable rows and stated pricing
    and modality facts (#6867).

  • Experimental TypeScript extension host. New in this release and off by
    default: turn it on with [features] extension_host = true. A plugin that
    declares a native TypeScript or JavaScript entry (the Cordis / DeepSeek
    Harness plugin model) can contribute tools, slash commands, pre-execution
    policy hooks, additive prompt sections, reviewed skill roots and plugin-local
    JSON state. Its code runs in a separate
    Node process, never inside Codewhale (Bun is an opt-in through
    [extension_host] runtime), and every call goes through Codewhale's own gate:
    an extension tool is never treated as read-only, so it always meets the
    approval gate (Full Access, Bypass or a session grant for that reviewed plugin
    can satisfy it), a command runs only when you type it, and a plugin must be
    reviewed and enabled first. The host is sandboxed on macOS (Seatbelt) and on
    Linux with bubblewrap, with no direct network and no reads of the protected
    credential locations (Codewhale's secret stores and the Codewhale, Codex and
    DSH homes). A Native extension starts only after that sandbox is verified at
    launch: if bubblewrap is missing or its namespace probe fails, activation is
    refused with the concrete error. Codewhale's own pinned built-in host modules
    keep a diagnosed exception, and every effect still needs a Rust operation
    ticket. On Windows a Native extension starts only in a freshly created Less
    Privileged AppContainer whose only capability is registryRead (Winsock
    cannot start without it), after Codewhale checks its token and a real file
    and network probe; otherwise activation is refused.
    Other files you
    can read, such as project .env files, stay readable, and plugins sharing the
    host can interfere with each other, so enable only plugins you have reviewed
    (docs/EXTENSIONS.md).

  • Extension slash commands: a plugin registers one with ctx.commands.register.
    It runs only when you type it, and it can show text, or submit a prompt as
    your next message that then goes through the ordinary turn and its approvals;
    it cannot call the model or a tool itself. A command cannot take a built-in
    command's name or another plugin's, and a markdown command with the same name
    wins. A command has 30 seconds, and commands are TUI only: the Runtime API
    does not list or run them.

  • Native extension authoring now includes policy listeners through
    ctx.on('tools/pre-execute', ...), prompt sections through
    ctx.prompt.registerSection, and persistent plugin-local state through
    ctx.storage. Listeners may deny, ask, revise input or annotate; Rust
    rechecks revised calls and owns every approval. Prompt sections cannot
    replace the system prompt, and state exposes no session history or secret
    API. These services share the experimental, off-by-default host's reviewed
    owner lifecycle and are withdrawn when their owner unloads or is revoked;
    saved state remains available when the plugin reloads.

  • Extension tool input is checked against the JSON Schema the tool registered,
    before an approval card appears and again before the call reaches the host. An
    invalid call comes back to the model as an error naming what to correct, and
    the host never sees it. A schema that cannot be compiled, or that refers
    outside itself with $ref, is refused when the tool registers.

  • An extension tool can call eligible core tools with core/call while the
    ordinary turn is running that extension tool. Rust supplies a temporary
    invocation ticket and still owns planning, hooks, permissions, approval and
    execution. The ticket is bound to the plugin owner, host process and live
    invocation, and expires when that invocation ends or is revoked. Commands,
    timers and plugin startup cannot use this path. Approval grants are scoped
    to the reviewed extension build; a grant to the model does not cover it.
    Shell and network calls require a person's prompt even with a remembered
    grant; a permission mode that cannot present that prompt refuses the call.
    MCP tools, Computer Use, other extensions, agent/workflow launches and
    permission-changing tools remain unavailable through this path. A cancelled
    invocation withdraws its pending approvals instead of deciding them for you.

  • Extension host identity is checked at startup: its reported trust tier and
    built-in module digests must match the process Rust launched. Built-in host
    code and third-party plugins have separate process and owner namespaces;
    the table pins the MCP protocol module and the execution harness. The
    optional Host MCP backend uses the official TypeScript SDK for framing,
    with Rust retaining credentials, network/process access, approval and exact
    operation tickets. All 30 existing recorded comparisons pass locally.
    The Rust backend stays the default until platform and rollout gates pass;
    its protocol adapters remain through the compatibility window.

  • A plugin can declare several native entries (native.paths, up to 64). They
    activate in order under one owner, so one disable, review change or crash
    takes all of them down. If one entry fails to activate, only that entry is
    withdrawn: entries that already activated stay registered, and the plugin
    fails only when no entry activates.

  • Plugin settings and context: [plugins."<name>".config] in your own
    config.toml is passed to the plugin's apply(ctx, config) and checked
    against its exported Config schema; a project's .codewhale/config.toml
    cannot set it, a change applies at /plugin reload, and /plugin show lists
    the keys, never the values. Tools and commands also receive the workspace of
    the session that called them and a private data directory for the plugin under
    ~/.codewhale/extension-host/data/plugins/; no home, credential or other
    workspace path is passed.

  • codewhale --enable <feature> and --disable <feature> work from the
    codewhale command; before, they were rejected. Put the flag before any
    subcommand.

  • The extension host bundle carries MIT-licensed code (Cordis, Schemastery and
    others). Its third-party licence notices are generated from what the bundle
    actually contains and written beside the bundle when it is unpacked, and
    THIRD_PARTY_NOTICES.md lists the packages.

  • Extension host runtime: Bun 1.4.0 or newer is an opt-in; Node stays the default. [extension_host] runtime = "node" | "bun" | "auto" chooses: unset means node (or bun when the table sets only a bun path), bun runs only Bun, and auto prefers a supported Bun and uses Node when none is found or the Bun host fails to start. An explicit choice never falls back, and restarts keep the runtime the session first started on; a runtime binary replaced mid-session is refused. A configured node or bun path is used as given and fails loudly rather than falling through to PATH, and a runtime found on PATH inside a node_modules directory or the working directory is skipped, not run. codewhale doctor and /plugin show which runtime runs the host, why, and how its memory is capped. Under Bun the host never auto-installs packages and ignores .env and bunfig.toml in its data directory. Extensions cannot run native code inside the host on either runtime: bun:ffi, Bun.FFI, node:ffi, the SQLite builtins, process.dlopen, process.execve, Worker threads and ShadowRealm are refused, and the host will not start if one of these restrictions does not hold. Known limit: that lockdown covers the native-code entry points found so far (Bun 1.4, Node 22 and 26); one a newer runtime adds is not covered until it is added. The 1 GiB memory cap uses RLIMIT_DATA on Linux, the Job Object on Windows and, for a Bun host, a kernel jetsam limit on macOS; a Node host on macOS is checked at each heartbeat. Bun is qualified on macOS only: the Rust host tests, the memory cap included, have not run a Bun host on Linux or Windows (CI runs them with Node, and the JS host suites under Bun 1.4.0 on Linux). host/hello reports the runtime and version that actually run.

  • docs/features.toml lists every user feature with its status, first
    release, docs page and owning code. A test fails when a [features] flag
    and its row disagree, or when a listed docs page or code path is missing.
    The configuration reference now lists the verify_tool, vision_model and
    extension_host flags, and config.example.toml lists code_mode.

  • A reusable GitHub Action runs PR reviews with a configurable model endpoint
    and a checksum-pinned Codewhale binary. It checks the PR's eligibility and
    exact revisions before inference; fork events receive no model key
    (#6780).

  • An optional, off-by-default shadow Decision Gate classifies the first request
    of a turn — tool need, answerability from context, intent — through the
    existing System One client and only logs a typed recommendation; it never
    changes routing or delays the model call. Configure it with
    SUPERFAST_ENABLED and SUPERFAST_PROVIDER (docs/CONFIGURATION.md)
    (#6603,
    #6604, thanks @Andrea-Bruno).

  • tool_call_after hooks for shell tools receive DEEPSEEK_TOOL_EXECUTION_RECEIPT: the command that actually ran after admission, its working directory, how it ended, and bounded stdout/stderr previews, so a hook can record exactly what executed (#6689, requested by @wuisabel-gif).
    Supported settled local shell calls also send a versioned JSON document on
    stdin to foreground or background observers, with session/tool-call IDs,
    nullable exit status and output truncation flags. Existing environment fields
    stay available; unsupported execution paths send no document, and observer
    output cannot change the completed call
    (#6582).

  • Runtime API: turns now record what they produced. Each item and turn
    carries typed artifact references (path, kind, size, revision, and a
    restore point when file-revert would accept one) for files a tool wrote,
    spilled tool output, and media.

    • Once the post-turn snapshot settles, the turn also lists what changed in
      the workspace while it ran, including shell and sub-agent writes. It then
      publishes turn.artifacts.
    • GET /v1/threads/{id}/turns/{turn_id}/artifacts lists a turn's
      references, and .../artifacts/{artifact_id} reads one from the
      workspace, the post-turn snapshot or the session artifact directory.
    • Spills from unbound runtime threads were unreadable before; they are now
      readable.
    • The legacy artifact_refs field is now filled with the workspace files a
      tool call wrote, so Preview in current desktop builds shows them
      (#6653).
  • Runtime API: git stage, unstage, discard and commit accept optional
    expect preconditions (full HEAD id, an index token, per-file rev, or a
    whole-tree revision, all read from GET /v1/git). When the repository
    changed since the client read it, the write does nothing and answers 409
    git_state_changed with the current state, so a Review sheet can no
    longer stage bytes, discard edits or commit an index the user never saw.
    Requests without expect behave as before. Path writes now use literal
    pathspecs, as the docs already said, and files[].path is
    workspace-relative in a subdirectory workspace. Guards cover executable
    mode, submodule HEAD and changes outside that workspace; oversized or
    budget-limited stat-only revisions cannot authorize a guarded write.
    Status stays best-effort: unreadable paths, symlinked ancestors, special
    files and broken nested repositories withhold only affected row tokens
    and the whole-tree revision. Normal Git filters and untracked settings
    apply, with directory inventories scoped to the rows or paths that need
    them. Untracked workspace files and renames leaving the workspace remain
    visible, and broken HEADs cannot satisfy an unborn-HEAD guard
    (#6647).

  • Runtime API: POST /v1/threads/{id}/fork-at-turn forks a thread at a named
    user turn, keeping that turn and every turn before it. The receipt matches
    /undo and returns the first dropped prompt so a client can put it back in
    the composer. Naming the turn replaces a client-computed depth, which could
    fork the wrong prefix. The fork leaves the workspace and any running turn
    untouched
    (#6580, thanks @gaord).

  • Official model routing: /router (also /model router) sets up the Auto
    router with presets: Jev (TypeSafe's decision model, via OpenRouter or a
    TypeSafe key), your provider's fast tier, Off, or Custom. Jev and Fast each
    make one test call before you save, the Turn Inspector (Ctrl+Alt+O or
    /turn inspect) shows the router's choice, cost and latency, and a failing
    router is shown as failing
    (#6525).

  • ChatGPT sign-in uses Codewhale-owned protected credentials and a model roster
    fetched for the signed-in account. Account, workspace and issuer changes
    cannot reuse another roster; a missing or stale roster offers no models.
    Authentication failures and subscription limits do not silently fall back
    to a paid API route.

  • Code mode composes MCP and plugin tools and is on by default:
    execute_tools programs can call MCP tools, and each nested call passes the
    same approval gate as a direct call, pausing the program for approval when
    needed. Every nested call keeps its receipt, including calls that finish
    before a deadline, and [features] code_mode = false stops offering
    execute_tools up front (it stays reachable through tool search).
    codewhale mcp list
    and codewhale doctor warn when a user MCP server duplicates the built-in
    Computer Use bundle
    (#6562,
    #6509).

  • Tsubasa is a bundled OpenAI-compatible provider row: base
    https://api.tsubasa.sh/v1, key TSUBASA_API_KEY, model tsubasa-pro
    (also tsubasa-fast). Its window is 32,768 tokens; set
    providers.tsubasa.context_window to match
    (#6695, thanks @cenab).

Changed

  • The pinned prompt header follows the turn at the top of the viewport,
    handing over to the previous prompt as you scroll, and clicking it jumps back
    to the message it names
    (#6830, thanks @SparkofSpike).
  • On Windows, every PowerShell command the shell tool starts now passes
    -ExecutionPolicy Bypass for that process only, so a local Restricted or
    AllSigned policy no longer blocks multi-line commands. (.ps1 script tools
    still start as powershell -File without the flag and remain subject to the
    local policy.) A policy set by
    Group Policy still wins and the command is refused with PowerShell's own
    message. Scripts that a command invokes run under the same process-scoped
    setting. To let the machine or user policy apply instead, start Codewhale
    with CODEWHALE_POWERSHELL_EXECUTION_POLICY=inherit: it then omits
    -ExecutionPolicy, so a Restricted or AllSigned policy refuses the
    multi-line commands that need a temporary script; any other value keeps
    the default. Not verified on a native Windows machine
    (#6745).
  • The website's not-found page now uses the Codwhale poster and typo joke,
    with English/Chinese recovery links to home and docs
    (#6419,
    #6420).
  • Script tools can no longer approve themselves or replace built-in tools
    (founder decision D4). A # approval: auto line in a script under
    ~/.codewhale/tools (or [tools].plugin_dir) is ignored: the tool follows
    the session's approval setting like a script with no approval: line, and
    the runtime log and /plugin tools name each script that still declares it.
    A [tools.overrides] entry of type = "script" or type = "command" keyed
    by a built-in tool name is refused, and the built-in stays active; a status
    line names the key once per session, and the runtime log records it. type = "disabled" still turns a
    built-in off, and script or command overrides under a new name still work.
    To keep a wrapper such as an audited shell, disable the built-in and give
    the wrapper its own name
    (configuration).
  • Complete the session slash-command group’s shared command boundary, including
    /structcopy, so all seventeen commands can compile independently of the TUI.
    Host operations remain behind capability interfaces; command behavior and
    upstream tool-execution identity safeguards are preserved
    (#6792,
    #6145).
  • The declared minimum Rust version is now 1.89. It said 1.88, which could not
    build Codewhale: a locked dependency (serde-saphyr) needs 1.89, the version
    CI's minimum-version job builds. The install guides and the npm wrapper's
    build-from-source hints now say 1.89 too.
  • A deferred tool's first call now runs when its arguments already carry every
    required field and no field the schema does not declare, instead of costing a
    retry turn. Malformed calls still get the schema, and approvals, deny lists
    and Plan mode apply first. Argument types are still checked by the tool
    itself. Sub-agents preload tools named in allowed_tools
    (#6494,
    #6437).
  • Automatic Git status and review reads need Git 2.31 or newer, because they pin
    Git's runtime configuration overrides. With an older Git they refuse with
    "upgrade Git (2.31 or newer)". This covers the read-only Git tools
    (git_status, git_diff, git_log, git_show, git_blame, verify); Git
    writes you ask for keep their existing behavior
    (docs/dependency-maintenance.md).

Fixed

  • A failure Codewhale can name is no longer labelled an internal fault. An HTTP
    400/405/409/413/422 rejection, an out-of-credits 402, the context-budget stop
    and a turn's own step or wall-clock ceiling now carry an input or budget
    label, and a bare ERROR from a provider is reported as an unreadable error
    instead of a warning. Refs #6843.
  • A transient upstream failure reported as an error frame inside a successful
    response is retried within the stream retry budget. When the budget is spent
    the turn fails once with an error card instead of an amber warning that
    promised a retry. Refs #6795.
  • Diff lines and tool output wrap at grapheme boundaries, so emoji families,
    skin tones and variation selectors no longer split across lines
    (#6829, thanks @Lstarsky0).
  • Twelve translated packs now translate the context inspector's making-room
    and anchors rows, the Ctrl+O hint and the /turn inspect and /advisor
    descriptions instead of showing English
    (#6831, thanks @Lstarsky0).
  • Optional MCP servers can be found before they connect. tool_search matches
    the query against configured, enabled server names (or an mcp_<server>_
    prefix), connects up to eight matches within the existing boot wait, and
    returns their real tool schemas. Before, a lazily started server exposed no
    tools, so the model could never trigger its connection
    (#6828). Known limit: a
    query that names neither the server nor mcp does not wake it.
    codewhale mcp connect, validate and tools run their own connection and
    do not attach to a running session.
  • /undo and /restore <N> refuse while a turn is running in the workspace,
    instead of rewriting files under it.
  • --resume <id> after a crash recovers that session's interrupted turn from
    its crash checkpoint, as --continue does.
  • /resume <file> keeps an imported session on the current provider route
    instead of a default one.
  • When the stall watchdog recovers a turn, the Engine's turn is ended too, so
    the next message is accepted
    (#6800).
  • A turn you cancel, or one the stall watchdog recovers, releases input at once.
    A message still waiting to reach the engine is abandoned and its unsent text
    returns to the composer, instead of holding input for up to 60 seconds
    (#6800).
  • A transient upstream failure that a gateway reports as an error frame inside
    a successful response ("Provider returned an empty response") is retried
    under the stream retry budget when nothing had streamed; authentication and
    invalid-model frames still fail at once
    (#6795).
  • Long sessions no longer start every turn late. Once a workspace held more
    than 50 undo snapshots (around the sixteenth turn), each new snapshot rebuilt
    the whole snapshot history before the provider request, about 2.8 s per turn
    in a small workspace. Old snapshots are now dropped half a window at a time.
  • Resuming a crashed session from inside the TUI recovers its interrupted
    turn too.
  • Release builds compile on Rust 1.99.
  • TUI undo and retry rewind the Engine conversation and saved session before
    replacement inference. If the conversation or its settings change while
    undo is being prepared, it refuses without overwriting that newer state or
    changing the draft. Retained compaction summaries survive the rewind and
    cannot be selected as editable user prompts
    (#6788).
  • /edit replaces the exchange it revises through the same rollback: the old
    prompt and its answer leave the transcript, the model's context and the
    saved session before the edited prompt is sent.
  • When the rollback that /edit needs is refused, your revised text goes back
    to the composer with edit mode still on, instead of being dropped. Opening
    /edit while you are editing a queued follow-up returns that follow-up to the
    queue first.
  • A session can still be resumed by its id when a stray copy of its file sits
    in the sessions directory; before, the id was reported as ambiguous.
  • Agents follow the Permissions you choose while they run: switching to
    Full Access reaches an agent that is already working, instead of leaving
    it with the Auto-Review guardian that denied it. Tightening reaches it too.
    An agent's answer to an approval prompt is delivered while the main
    session is busy, and a forced delete inside the workspace is no longer
    treated as a catastrophic system delete. On Linux, Full Access switched on
    at runtime cannot lift the kernel no-new-privileges flag set at startup,
    so sudo and setuid helpers still fail; start with
    sandbox_mode = "danger-full-access" or CODEWHALE_NO_NEW_PRIVS=0 if
    agents need them (#6787).
  • Stream limits and transport settings share a typed [stream] configuration
    table, including retry budgets, TCP keepalive and HTTP/2 keepalive. Explicit
    values take precedence over legacy [tui] aliases; omitted values preserve
    existing defaults and environment behavior. A configured HTTP/1 pin stays
    pinned during recovery. /config stream_chunk_timeout_secs ... --save
    updates the canonical setting so it survives reopening a configuration
    that already had a timeout. Transport changes apply when a client is built
    (#6700).
  • Switching providers keeps a model set only in the root default_text_model
    with the provider it belongs to. Switching away and back (/provider in the
    TUI, or the desktop app's model chip) used to land on the provider's catalog
    default such as gpt-5.6, and a pass-through provider switched to in between
    inherited the other provider's model, or a stale legacy root model that the
    root default had been shadowing
    (#6693).
  • A turn no longer stops after an hour of work. The cumulative per-turn wall
    clock is unlimited by default, like model steps; set
    [tui].turn_wall_clock_secs to cap it.
  • OpenRouter turns are priced again instead of always showing "rate
    unavailable". The model-list refresh no longer fails as a whole on one bad
    row: ~ "latest" aliases are dropped, routers that publish a -1
    variable price are listed without a rate, and other malformed rows are
    skipped and counted. Main interactive turns now honor a [[custom_models]]
    rate the same way background work does, a declaration without rates no
    longer hides the catalog price, and an explicit declared rate prices a
    vendor-pinned OpenRouter route
    (#6690).
  • codewhale exec can take prompts past the OS argument limit (~128 KiB on
    Linux, where larger prompts failed with Argument list too long before
    Codewhale started): --prompt-file <PATH> reads the prompt from a file, and
    --prompt-file - reads it from stdin; a positional - stays literal text
    (#6688).
  • Keep reasoning effort labels, including xhigh and ultra, visible beside
    short model names in the 80-column terminal footer.
  • First-run TUI startup keeps explicitly configured providers, models and
    endpoints (including environment overrides and a legacy top-level
    base_url), even with a missing key, instead of adopting a detected local
    Ollama model. The default_text_model line written into the generated
    first-launch config is not treated as a choice, so starting Ollama after the
    first launch is still detected. Usable configured routes skip the provider
    picker and record the provider step as configured, without claiming a
    successful check; missing-key recovery focuses the configured provider even
    on a fresh home. Switching from that recovery picker (or after Esc) to a
    provider that has its key clears the launch "needs a key" state, so the
    footer names the route instead of "model not connected" and a detected local
    Ollama can no longer take over the provider just chosen. Empty sessions no
    longer show a full context window from the startup prompt; a submitted first
    turn or reported usage still shows real pressure.
  • /trust on|off changes file-tool trust for this session. Add --save to
    persist workspace trust for project skills, commands, hooks, MCP servers,
    and project context. Skipped-skill notices stay out of conversation titles
    and user bubbles, and trust changes append a correction to the model's history.
  • Stored-secret cleanup skips busy Runtime stores and reports unreadable paths
    while continuing through other files. Configured storage roots may be symlinks;
    discovered symlinks are not followed. Cleanup also masks current configured
    credentials, and doctor's bounded scan reserves capacity for saved transcripts
    and checkpoints separately from Runtime receipts.
  • Hooks: tool_call_after and on_error now get a shell command's exit code
    in DEEPSEEK_TOOL_EXIT_CODE on Runtime API threads as well as in the TUI,
    and for a failing command as well as a passing one, so exit_code
    conditions match. The Runtime API path passed no exit code at all, and a
    command that exited nonzero, timed out, or was killed reached hooks with no
    exit code on either surface. The new DEEPSEEK_TOOL_STATUS says how the
    command ended (completed, failed, timed_out, killed)
    (#6582).
  • Scrolling a long transcript repaints the reasoning reveal hint in place
    instead of rebuilding the transcript tail. Wheel and scrollbar input use
    interactive frame pacing, collapsed tool groups reuse cached summaries, and
    height-only resizes keep wrapped history. Width changes still rewrap content,
    and other per-frame work still grows with session length
    (#6652).
  • A settled transcript frame reuses its existing render plan instead of
    rebuilding or moving unchanged history on every poll. New output, selection,
    resize and other real changes still invalidate the affected work
    (#6652).
  • Workflow progress rows and working, verification and translation indicators
    use the pinned Codewhale terminal component kit. Quiet and reduced-motion
    working/verification indicators hold its semantic current-work mark. Working
    and verification keep their existing delay and cadence, while translation
    uses the same earned marker and five-frame-per-second cadence. The Engine
    still supplies workflow state, clocks, localized text, custom themes and
    terminal adaptation. The composer, transcript viewports, dock tabs,
    posture/metrics rows, pending-input cards, approval band, Ocean background
    and the whale's Braille raster now also render through the pinned kit.
  • Calm transcript previews keep the latest three rows of live thought and a
    single duration header when it settles. Successful tool headers are quieter;
    failed generic and MCP calls keep a bounded head-and-tail excerpt, with full
    output available in details. Existing collapsed groups show their hidden
    count and reveal cue, and explicit thought expansion stays available.
  • Composer history preserves multiline entries and literal quoted text while
    retaining the old history file during migration. An oversized message stays
    in the composer if its paste file cannot be saved, with a localized notice;
    mouse placement and quoted file mentions handle multibyte text and spaces
    (#6776).
  • The notice for an oversized message that cannot be saved as a paste file is
    now a short line that fits the footer ("Not sent: over N characters; paste
    file not saved."), in every supported language. The full reason, including the
    write error, goes to the transcript once.
  • Model-facing guidance names callable tools and explains discovery before a
    deferred read. The stopship workflow's scout activates the deferred
    grep_files with one required tool_search call, then gathers bounded
    grep_files evidence from the current workflow source paths
    (#6747).
  • Simplified and Traditional Chinese permission, provider and session wording
    now follows the current behavior. Chinese guides correct hook receipts,
    trusted skills, MCP startup and Windows limitations, preserving the
    contributor translations and a shared glossary.
  • Unknown website locale segments, including dotted missing paths, return a
    real 404 with noindex metadata and no home-page canonical URL. The existing
    layout guards now have regression coverage across all registered locales
    (#6786).
  • The terminal caret no longer blinks at the hidden composer while a picker,
    settings screen or other view covers it; it returns when the view closes
    (#6545).
  • On Linux and macOS, a background: true shell's process group is now
    stopped when the TUI dies without cleaning up (SIGKILL, a crash, or the
    signal exit path). The shell and the processes it started used to keep
    running as orphans. Processes that move to their own process group or
    session, staged services you choose to keep, and tty: true background
    shells are not covered. Under the bwrap sandbox, the sandboxed command
    now also exits when bwrap does
    (#6654).
  • Idle task workers stop periodically reloading a settled, unchanged empty
    queue, and task listings reuse an unchanged store snapshot. New queue writes
    and local notifications still wake workers; recent or unreadable metadata and
    failed claims keep bounded retries. This reduces repeated disk work when
    several TUIs share a data directory without changing task ownership or
    cancellation (#6573,
    #6728).
  • Idle task workers check the queue file's metadata about once a second instead
    of five times a second, and no longer reloads the store when nothing changed.
    An in-process notification still wakes a worker at once; a
    write by another Codewhale process that shares the data directory is noticed
    within about a second instead of about 200 ms. A worker with pending work, a
    retry deadline or a failed claim keeps the short interval
    (#6728,
    #6573).
  • After five seconds without session activity, the UI loop polls every 250 ms
    instead of 48 ms and the automation-panel scan backs off to 15 seconds. After
    thirty quiet seconds, the Git probe runs one git status every 15 seconds
    instead of the full probe every two seconds, unless the Git panel is showing.
    User input and engine events restore the active cadence. This reduces idle
    work; it does not remove every cost that grows with session history
    (#6728).
  • When the task store is busy, the message now names the process that holds its
    lock and how long it has held it ("lock held by pid 1234 for 12s"), when that
    can be read. A repeated "Task claim unavailable" error is logged once per
    episode, then at debug level
    (#6573).
  • On Windows, Fleet can read a running worker's log while its writer holds the
    file open, and task-store contention can still read the lock-holder record.
    These paths previously failed because of whole-file sharing and locking.
  • File writes, including legacy File and write_file calls, stop if the
    original contents cannot be read; legacy writers also reject non-UTF-8
    originals, keeping undo and diffs from recording an empty original file.
  • Resuming a session preserves user messages that quote a compaction marker.
  • Requirements allow-lists now check unset approval and sandbox defaults,
    accept equivalent approval aliases, and lock permission posture when sandbox
    modes are constrained so saved Full Access and YOLO cannot bypass them.
  • Logout completes xAI OAuth revocation and attempts every credential deletion
    before reporting remaining credentials with a non-zero exit status.
  • ChatGPT and xAI sign-in now show the account the selected route actually
    uses, and sign-in names the account it replaces. /auth chatgpt and
    /auth xai-device switch the running session; signing in from a separate
    shell requires restarting an already-open session. Subscription-limit errors
    name the account that made the request and explain how to switch, without
    printing credentials (#6715).
  • doctor --fix keeps temporary files modified within the last hour.
  • Provider streams stop with a clear error if an SSE line exceeds 8 MiB.
  • Recursive rlm_query turns inherit the parent's remaining deadline,
    bounded by the existing child budget, including asynchronous lock waits,
    startup, model requests and Python evaluation. Expired work is refused before
    dispatch; a timeout returns the last response with an incomplete-result error.
    Nested code and status receipts accompany the enclosing tool result on success
    or failure and survive its session save/reopen. In-flight receipts still need
    that result to be handed back before they become durable
    (#6511).
  • The TUI keeps redrawing while its terminal is unfocused. v0.10.0 held
    every frame on focus loss, so on Windows Terminal, macOS and other
    terminals a window sitting behind another one looked frozen until you
    clicked back into it. Frames are now held only on GTK/VTE terminals, which
    replay the damage themselves when they return. That hold was added for the
    MATE flicker in #6311
    (#6651).
  • TUI /undo now restores only the files the undone tool call or turn
    changed, and refuses, changing nothing, when one of them changed since.
    Before, it checked out the whole snapshot tree, which also reverted later
    edits to other files. It also finds restore points older than the newest
    100 snapshots, and a forked session can undo the turns it inherited. It
    waits for a post-turn snapshot that is still being written, leaves a
    changed symlink or other non-regular path in place (and says so) instead
    of refusing the whole undo, and writes nothing to the snapshot store when
    it refuses outside trusted mode
    (#6644).
  • On a fixed model, each Ctrl+T press now moves to a different thinking
    tier. Ctrl+T and the /model picker could list tiers the route treats as
    the same one — DeepSeek's medium is high, and Z.ai GLM-5.2's low is
    high — so some presses changed nothing. They now share one list per
    route with one entry per effective tier. Under Auto model routing, Ctrl+T
    walks the picker's Auto list instead of a longer private one; the tier is
    still settled when the turn is routed, so two neighbouring choices can
    land on the same tier. DeepSeek now reports minimal, xhigh and ultra
    as the low, high and max it sends, in /status, receipts and
    /effort, and Kimi Code K3 no longer offers off, which it always ran as
    low (#6650).
  • A top-level base_url or api_key in config.toml now means one thing
    everywhere. Every reader used its own rule for which routes inherited it,
    which is how a DeepSeek endpoint became the Xiaomi MiMo route's and failed
    with DeepSeek's 401. Old files keep working: the keys are read as
    [providers.deepseek] (or the vendor whose official host they name), the
    next save moves them there with a one-time backup and a one-line note, and
    a value that disagrees with its table is left for codewhale config migrate --prefer to settle. codewhale config doctor shows what is in use
    (#6394).
  • Setting up a bundled OpenAI-compatible host from /provider now shows
    that host's key console, docs link and guidance on the setup form, the same
    way the built-in providers' key entry does; before, the descriptor file
    carried them and nothing read them. AICraft gains all three
    (#6616, thanks @BX166).
  • Runtime API undo now restores the files of the turns it undoes, for every
    kind of thread. Before, a new thread's turns were not linked to their
    workspace snapshots, so patch-undo returned 201 and rewound the
    conversation but left the files changed, and file-revert refused. Each
    turn record now lists its snapshots in workspace_snapshots, the engine runs
    under the thread's own id across restarts, and a fork owns the turns it
    inherited. patch-undo restores only the files the undone turns changed,
    including every write in a turn, and leaves later edits by the user or
    another thread alone. Every tool call that may write is bounded by its own
    snapshots, so a file another thread or an editor changed while the turn ran
    is never reverted as the turn's, and a turn's restore points survive the
    snapshot count cap. When it cannot restore files (including a write to a
    gitignored path) it now refuses with 409 and an error.code instead of
    returning 201. A nameless PUT /v1/sessions updates the document the
    thread is bound to
    (#6621).
  • A Runtime thread that is not bound to a saved session keeps one engine
    session id, its own thread id, across its first turn, eviction and
    restarts. Before, each engine spawn generated a new id, so the thread's
    tool-output spills, snapshot tags and shell jobs were scattered across a
    different sessions/<id>/ per spawn
    (#6659).
  • Saved sessions no longer go orphaned, and existing orphans are repaired at
    launch. Deleting a session now sets aside the Runtime store it was actually
    bound to, and switching a conversation to another host's store sets aside the
    store it left, unless another session still uses that store or it holds
    work. Exporting a thread (POST /v1/sessions) updates the thread's own
    document instead of creating a new one each time. A thread whose document was
    appended to elsewhere keeps loading. A thread whose document was deleted or
    rewritten loads from its own turns instead of failing, and threads write
    their files under the session they belong to. Each launch and codewhale serve repairs the store in the background: a store holding threads with no
    session gets a "Recovered:" session for each thread, and unreadable
    documents, unused empty stores and old files no session names move to
    sessions/.set-aside/ with a manifest. Nothing is deleted. codewhale doctor
    reports the last repair, and codewhale doctor --repair-sessions [--dry-run]
    runs one now. PUT and DELETE /v1/sessions refuse a session that another
    Codewhale process has open
    (#6144).
  • Runtime API: the thread event stream now ends with a typed stream.end
    frame (reason and resume cursor) whenever the server closes it, including
    on shutdown, so clients can tell a Runtime error from a dropped connection
    and resume from the right event. A replay worker crash no longer ends
    history early and skips events without a detectable gap, and the mobile
    page resumes from its last event instead of replaying from zero, with stream
    statuses in the configured UI language. Cross-origin clients can read the
    stream-end and replay-progress capability headers, and the SDK exports
    isThreadStreamEnd for TypeScript narrowing.
  • codewhale exec --auto no longer exits 141 with no output when a child
    it writes to, such as a stdio MCP server, closes its pipe early. Headless
    exec now ignores SIGPIPE while it runs, as the interactive TUI already did,
    and exec ... | head still ends quietly. One-shot codewhale exec no
    longer prints DeepSeek's raw <||DSML|| calls> tool-call markup as its
    answer: the markup is removed, and an answer that was only a tool call
    fails at once with the reason and a pointer to --auto, instead of asking
    the model again and blaming an incomplete provider response.
  • Resuming a session keeps its "Resumed:" confirmation on screen instead of
    replacing it with "Make room automatically: on" when nothing was switched.
  • Tool output is no longer cut off where you can't get it back. Search
    answers, test runs, git, verifier and web results reach the model whole up
    to one budget sized to the model's context window, and anything beyond it
    can be read back with retrieve_tool_result, including read-only tools
    executed alone or in parallel. Restored raw results use the
    active route's same inline budget; existing recovery receipts stay intact.
    Tiny budgets use a compact retrieval reference instead of a long artifact
    path and instructions; a complete reference is kept even when it alone
    exceeds the allowance.
    Native search now asks for answers up to 8,192 tokens (was 2,048–4,096)
    and waits long enough for
    them to arrive, and an answer the provider still cuts short is marked as
    cut (#6508).
  • The installation page is generated from docs/INSTALL.md, so the website
    and the guide can no longer disagree; broken anchors and unsafe links fail
    the build (#6450).
  • codewhale config set refuses a value of the wrong type for a known setting
    (a word for an on/off switch, text for a number, a choice outside the list)
    instead of saving it (#6568,
    thanks @dajiaohuang).
  • Receipts: /receipts, codewhale receipts [ID|--last] [--format md|json],
    and GET /v1/threads/{id}/receipt (plus a per-turn form) list what a session
    did, one line per action: files changed with line counts, commands with exit
    codes, web and MCP calls, agents, approvals and who gave them, and failures.
    They also count what ran without asking and name the posture each turn ran
    under, read from the turn's own record. A call Codewhale blocked before it
    started (Auto-Review or guardian, a tool policy, a refused sandbox
    escalation, invalid input, a missing tool) is listed as blocked, with the
    reason, and is not counted as run or as ran without asking. Only
    Codewhale's own refusal text counts: an MCP server, a fetched page, or a
    program cannot make a call that ran read as blocked. A terminal
    session's receipt also lists the files a command changed in each turn,
    from the workspace snapshots taken before and after it (not ignored files
    or anything outside the workspace), with control characters in paths
    escaped so a file name cannot forge a receipt line. All three read the
    records Codewhale already keeps and say what those records do not hold
    (docs/RECEIPTS.md). audit.log is not that record: it
    logs security events, and it logs an approval only when one is requested,
    which under Full Access is almost never.
  • Approvals now record who decided: you, a session rule, or the posture. An
    automatic approval used to be saved exactly like one you gave, and an app
    approval that expired was saved as your denial. GET /v1/approvals now
    returns decided_by.
  • Network audit lines now go to the same audit.log as every other audit
    event ($CODEWHALE_HOME included), and test runs no longer append to your
    real one.
  • In the terminal UI, Auto-Review verdicts now reach audit.log, as
    /permissions said they did; before, they were written only when
    CODEWHALE_TOOL_AUDIT_LOG was set. Headless exec and Runtime API sessions
    do not write them yet.
  • A turn that stops producing output now reports itself: the turn loop records
    its phase and last progress, and an overdue phase surfaces instead of
    hanging silently until the stream idle timeout. A delegated agent's final result is
    never dropped when the host is busy, so a finished child no longer leaves a
    ghost Running row behind (#6184).
  • Git commands run by tools never stop to ask for a password, passphrase or
    host-key confirmation inside the terminal, and git_fetch has a timeout
    (#6184).
  • A provider response that ends cleanly with no text and no tool call is
    retried before the turn fails, and the failure names how many retries ran
    (#6310).
  • The context meter, the point where Codewhale makes room, preflight,
    /context and turn receipts show one pressure number instead of disagreeing
    (#6407).
  • When Codewhale makes room, the note it leaves for the next turn is written in
    its own words, under fixed headings: the objective, the user's corrections,
    permissions and limits, changed files, what is still running and the checks
    left. A message that quotes the note's opening words now stays in the
    conversation instead of being dropped. If the summary request itself is too
    large, older history is dropped but the previous note is kept, and making
    room stops rather than write a new note without it.
  • Continuing a conversation that is already open no longer adds a second
    thread, and a fork keeps its own session file, so autosave on one side no
    longer leaves the other unloadable
    (#6406, thanks @gaord).
  • Upgrading Codewhale no longer turns off the built-in Computer Use. Each build
    writes the built-in bundle to its own directory, so an upgrade used to present
    it as never reviewed and disabled. Now the review and enablement carry to the
    new build when its capabilities are unchanged. Changed capabilities show
    capabilities-changed and wait for review, and a revoked trust never carries
    (#6303).
  • "Allow for this conversation" records a grant for that tool and argument
    class instead of switching the whole thread to Full Access, so the call you
    just approved is no longer failed by a Permissions change. An approval
    also survives a Permissions change that only widens what is allowed, grants
    end when a thread is archived or deleted, and web.run open grants are
    scoped by host. Full Access covers MCP tools that declare themselves destructive in
    every host, including codewhale exec
    (#3866).
  • web.run retries a refused page once with a browser user agent, and one
    site's failure no longer fails the whole call or drops its search results.
  • Web search no longer stops when DuckDuckGo cannot be reached. After a
    connection error, timeout or non-2xx answer, the search tries Bing before
    giving up, and the result says it used Bing. A DuckDuckGo that hangs no longer
    uses up the whole time budget: when Bing may follow, DuckDuckGo gets 60% of it
    and Bing the rest. This also covers the DuckDuckGo step after a configured API
    provider. A custom search_base_url still has no public fallback and keeps
    the full budget. Bing sees the query only after DuckDuckGo has failed or
    returned nothing it could read
    (#6746).
  • Hooks treat bash, Bash and exec_shell as one tool in tool_name
    conditions, so the documented example fires.
  • macOS no longer reports Codewhale's ordinary heap as GPU (IOAccelerator)
    memory.
  • Code highlighting uses less memory, and long transcripts, the pager and the
    session picker do less work on the event loop; session previews load in the
    background (#6014).
  • The composer's send cue follows the draft, not a paste in progress
    (#6397).
  • Voice status is localized, ASCII-mode markers are distinct, and the cursor
    honours NO_COLOR (#5846).
  • /cache, /stash, /config, session prune, metrics --since and the
    lane start/lane stop --json flags handle their edge cases.
  • Plain codewhale exec (no --auto) runs one Engine turn with the same
    system prompt as every other run, so --json no longer changes the model's
    instructions. It offers no tools unless a flag grants them (--auto,
    --yolo, --allowed-tools, or resuming a session): --max-turns,
    --disallowed-tools, --append-system-prompt, --sandbox and
    --output-format stream-json no longer turn a chat call into a tool-using
    agent, and tool-only flags passed without a grant print a warning. A reply
    cut off at the provider's output limit is continued instead of failing the
    run, for at most 8 model steps unless --max-turns sets another limit.
    Breaking for scripts: the --json one-shot receipt no longer has
    stop_reason, adds the agent receipt's fields (prompt, tools,
    outcomes, status, termination_reason, error_category), and leaves
    out usage when the turn never settled
    (#6510).
  • codewhale review of a plain diff uses the same review prompt as
    review --pr and prints the structured review as Markdown, or the model's
    prose when it ignores the JSON format
    (#6510).
  • A recursive rlm_query that runs out of rounds returns its last answer
    marked [rlm_query incomplete: …] instead of an empty string, its model
    calls appear in the parent turn's record, and its history is no longer
    trimmed (#6511).
  • Starting without a network connection no longer drops images you attach to a
    model that accepts them. The offline model list lagged behind providers and
    listed Claude and others as text-only; it can now only say a model takes
    images, never that it refuses them, and a provider that does refuse gets one
    resend without the image and a message saying so
    (#6396).
  • Model prices that a flat rate would get wrong now show as unknown after the
    model list refreshes online, as they already did offline: DeepSeek (priced by
    time of day, or from Codewhale's reviewed table on hosted providers), Grok and MiniMax M3 (rates rise on long prompts), Xiaomi MiMo
    (pay-as-you-go and Token Plan keys look the same), GLM-5.3 and the Alibaba
    Model Studio plans (quota or credits, not per-token) and StepFun. DeepSeek's output limit stays at its
    published 384K instead of the refreshed 393,216, and the reasoning controls
    Codewhale records for Grok, MiniMax, Qwen 3.8 Max, Muse Spark and Step 3.5
    (defaults, always-on thinking, extra effort tiers) survive the refresh.
    Committed corrections now fail at load if their provider or model is missing
    from the bundled seed; partial live refreshes may still omit those rows
    (#6396).
  • The offline model list is now generated from Models.dev instead of edited by
    hand, so a fresh install without network sees the same limits, image
    support and reasoning controls as an online one. Claude, GPT-5.5, Kimi and
    several OpenRouter models now show image input offline, and a model whose
    catalog row offers an on/off switch next to its effort tiers keeps Off in
    /effort and Ctrl+T instead of rounding it up
    (#6396,
    #6612).
  • Cost estimates for GPT-5.6 and GPT-5.6 Sol use OpenAI's current rates of
    $4 input, $0.40 cached input and $20 output per million tokens. They were
    still at the older $5, $0.50 and $30
    (#6396).
  • codewhale integrations dsh now reads current DeepSeek Harness prerelease
    versions such as 0.1.7-alpha.2 as semver. They were reported as offline
    and refused; they now show as stale-version (launchable once connected,
    clearly unverified). Only 0.1.0-rc.6 is still the verified version, and text
    that is not a version is still offline.
  • Rule and chain-segment source labels in the project instructions the model
    receives (<project_rule source=...> and the scoped-instructions comments)
    are repo-relative instead of absolute. Moving or recasing a checkout no longer
    changes the pinned system prompt prefix or adds a spurious context update
    (#6739, thanks @asto18089).

Security

  • On Windows, a broad Node kill (Stop-Process -Name node,
    Get-Process node | Stop-Process, taskkill /IM node.exe, and wrapped,
    aliased or nested-shell forms) is held by a built-in safety floor: it is
    refused in Full Access, Auto-Review and Never, and asks in other modes. With
    the npm launcher such a command ended this and every other npm-launched
    Codewhale session without cleanup. Stopping a server by PID or port still
    works, and the Windows installer and archives have no Node launcher
    (#6827, thanks @jayanthvee).
  • code_execution and js_execution now start their Python and Node
    interpreters through the same permission-aware launcher as other commands, so
    the session's execution policy applies to them. Before, an approved call ran
    the interpreter directly, and its working directory was the only boundary
    (#6820, thanks @Guan0923).
  • A saved task without an auto_approve field is no longer treated as
    auto-approved. Updates accept HTTPS URLs only from the release-host
    allow-list, and the npm release-asset check bounds every request.
  • Workspace profile and skill discovery, pasted-image and screenshot writes,
    configuration redaction receipts and legacy configuration migration reject
    linked paths where they could redirect access outside their intended roots.
    Additional .codewhale writers use the existing confined filesystem helpers.
    The VS Code file opener checks the file's real path, and PowerShell temporary
    scripts receive unpredictable names and owner-only permissions.
  • Update development tooling to Undici 7.29.1 or newer for
    GHSA-w293-vg96-wgc3,
    brace-expansion 1.1.21/5.0.12 for
    GHSA-q2hr-2g5m-vwhr,
    and the VS Code extension's markdown-it to 14.3.2 for
    GHSA-253c-mchw-3w2r.
  • Harden local runtime browser sessions, fleet SSH trust, agent continuation
    ownership, task gate approval, plugin tool registration, bridge action tokens,
    and release metadata credential forwarding. Browser sessions recover across
    reloads and new tabs, and stream tickets retry after transient failures.
    Fleet SSH known-host checks support OpenSSH 7.x and later. SSH host configs
    using host_key_fingerprint must migrate to known_hosts with verified host
    keys; the unsupported fingerprint field is now refused when an SSH worker
    starts, before any connection is made, instead of being silently ignored.
  • Harden workspace instruction, note, and anchor file access with shared
    no-follow reads and writes. Compaction loads pinned anchors only from
    trusted workspaces. Validate registry skill names before selecting cache paths.
  • Tighten approval, execution and endpoint boundaries. Computer control
    consent and script calls need an exact decision a person gave on that call's
    own approval card. Automatic Git status and review reads go through the
    sanitized review command and stay pinned to the checked Git executable.
    Config backups, config dumps, MCP listings and notification payloads share
    one sensitive-key vocabulary; structured exports now also redact keys ending
    in key. Python a child model writes during a recursive rlm call is
    admitted round by round like an inline repl block, and Python kernels are
    discarded when permissions narrow. Gate commands and the cargo test runner
    start inside the session's shell sandbox. Mutable session artifacts are
    written and read through pinned session-relative paths. The stdio bridge's
    Runtime child chooses and reports its own loopback port.
  • Chat bridges (Feishu, Telegram, WeCom, Weixin) accept an approval decision
    only from the person who started that turn, for the Runtime's current
    pending approval. After upgrading, approvals for turns that were already
    running are decided from the TUI, and WeCom /allow no longer takes
    remember.
  • More file and network paths are confined. A task or automation created by a
    session without full shell access no longer inherits the host's shell default.
    The commit planner, oversized-paste backup and project harness notes do not
    read or write through links that leave the workspace. On Unix, audit and
    approval logs are created owner-only and are not opened through a link;
    Windows is unchanged. Lane ids must be plain names. A session file is refused
    when it records a different session id than its name. The updater does not
    follow a redirect from HTTPS to plain HTTP and stops reading a response larger
    than 512 MB.
  • Nothing is written through a .codewhale that is a link. The workflow run
    journal no longer creates .codewhale/ or its ledger when it is opened, only
    when the first record is written, and reads and appends through the same
    confined helpers as other state. The sub-agent coordination lock refuses a
    linked .codewhale or state directory before it creates anything.
  • On Unix, the runtime log is created owner-only and is not opened through a
    link; an existing log from an earlier version is tightened. The
    .reconcile.lock and current.json.lock lock files are created owner-only
    too.
  • Inline ```repl blocks in a reply now follow the same rules as the
    code_execution tool. They never run in Plan mode or when
    code_execution is not on the turn's tool surface, and they are admitted
    like a code_execution call with the same code: before-tool hooks,
    disallowed-tool rules, Auto-Review (including its reviewer) and repository
    law all apply, then the same approval under the session's permission
    posture (a blocked, denied or unanswerable call skips them with a visible
    note). Only a fence that opens its own line counts, here and in the rlm
    tool's child rounds, so a reply that mentions a fence in prose runs nothing.
    An rlm(...) call from a fence gets a one-shot child answer rather than a
    child that runs its own code. The docs no longer describe the REPL kernel
    as sandboxed.
  • codewhale serve --mcp now exposes only read-only tools by default
    (file_read, search). Tools that write files or run commands are
    neither listed nor run unless the operator sets require_approval = false
    in the MCP server config; a client's own approved flag no longer counts
    as approval. At startup the server names on stderr each configured tool it
    withholds and the setting that allows it.
  • A Bash entry in the disallowed tools now also blocks start_mcp_server
    and start_registry_mcp_server, which launch local processes.
  • An MCP tool call whose connection closes before it answers is no longer
    sent again on a new connection, since the server may already have run it.
    The same holds for a server error reply that mentions an expired or invalid
    session. The call reports an unknown outcome and the next call reconnects.
    Only a call the HTTP transport turned away for a stale session, before the
    server handled it, is still retried once.
  • Durable tasks and scheduled automations created by the agent no longer
    carry more authority than the session that created them: requested
    auto_approve, trust_mode and allow_shell are capped at what that
    session holds, a task's workspace and an automation's directories must be
    reachable from the session, and a task's turn runs under the posture pinned
    when it was created instead of re-deriving one from the legacy bit or a
    legacy mode alias.
  • The approval card for creating a task or creating/updating an automation
    now lists the requested trust mode, shell, auto-approve, mode and
    workspace as labeled lines.
  • "Approve for session" on a shell command now covers only a simple, known
    command family (git status still covers git status -s). Compound,
    wrapper, interpreter, unrecognised and option-configured commands are
    granted as the exact command, and interact/wait calls on a running shell as
    the exact call.
  • Workflow plan approval flags shell, network and file-write capability for
    the Bash, Web and File tool families regardless of letter case.
  • An agent session can no longer update, resume or run an automation whose
    stored auto_approve, trust_mode or allow_shell is more than the
    session holds, unless the same update lowers those fields.
  • The approval summary every Runtime client receives for task and
    automation create/update now names the requested trust mode, shell,
    auto-approve, mode and workspace, from the same fields the TUI card shows.
  • A shell session grant keeps its command family only when every option is a
    known value-free option such as --release or --porcelain; other options,
    in any spelling, grant the exact command. Families whose arguments are what
    runs or is installed (go run, deno run, cargo run, make,
    git bisect, git submodule, package installs) grant the exact command.
  • Workflow plan approval recognises tool names in any quote style and with a
    pattern suffix ('Web', `File`, Web(*)), and flags rlm, Git and
    GitHub, computer and browser control, and task/automation tools with the
    capabilities they carry.
  • Model-run code no longer inherits provider credentials or other secrets
    from Codewhale's environment. Python code_execution (and every other
    Python runner), gate_run task gates, custom and completion verifier gates,
    persistent terminal sessions, and run_tests now start from the same
    scrubbed child environment as exec_shell; variables a gate declares are
    still passed through. gate_run also runs /bin/sh -c instead of a login
    shell, so profile files cannot re-export what was removed.
  • A workflow start with verify: true now needs approval even for a
    read-only plan, because the completion gates run workspace build and test
    scripts.
  • Git commands no longer hand provider credentials to programs that a
    workspace's git config can start (core.fsmonitor, hooks, filters) when
    they run through Codewhale's shared git helper: those start from the scrubbed
    environment plus the ssh agent, global config location and author identity.
    Sub-agent worktree provisioning and delivery do not use it yet. The read-only git_status,
    git_diff, history and verify tools also disable the workspace's
    fsmonitor, hooks and clean/process filters.
  • Language servers started for post-edit diagnostics, and every Python and
    Node command constructor, start from the scrubbed child environment.
  • Proxy URLs passed to model-run children (HTTP_PROXY, HTTPS_PROXY,
    ALL_PROXY, FTP_PROXY, CARGO_HTTP_PROXY) keep their host and port but
    lose any user:password@ part.
  • Behaviour change: the scrubbed child environment now keeps non-secret build
    configuration — CARGO_* (except registry and token-like keys),
    RUSTFLAGS, RUST_LOG, RUST_BACKTRACE, VIRTUAL_ENV, JAVA_HOME, Go
    paths, NVM_*/NODE_OPTIONS and CA-bundle variables — so run_tests and
    gates keep the user's target dir, job limit and toolchain. Connection
    strings such as DATABASE_URL are still dropped; set them in the project's
    own config, or wrap the command in a script that sets them, when a build
    needs them.
  • Deny rules in the permission settings now hold for these ways of hiding a
    command word until the shell runs it: a variable ($v), a substitution, a
    glob or brace list, escaped ANSI-C quoting, or a shell reading its script
    from a pipe, here-string or process
    substitution. Rules in a separate execpolicy.toml file are applied when
    the command runs, so there such a command is refused at run time rather than
    before an approval prompt. While any deny rule is configured in the permission
    settings, such a command is refused instead of being checked against text
    the shell will rewrite (an
    approval-always policy asks instead). This also applies to common forms such
    as source "$HOME/.cargo/env", eval "$(pyenv init -)" and $PYTHON -m pytest; name the command directly to run it. The same holds when the text
    does not parse cleanly (an unterminated quote, substitution or heredoc, a
    case inside $( … )) or when a parse budget runs out. Commands after
    if, then, while, do, ! and similar words, function f { … }
    bodies, and find -exec payloads are now checked like any other command.
  • Wrapper commands are unwrapped by their real option grammar, so
    chroot DIR cmd, sudo --user NAME cmd and timeout -s SIG N cmd expose
    cmd (and any -c payload) to deny rules. Options missing from a wrapper's
    table are read both with and without a value, BSD and macOS options are
    covered (env -P, xargs -J, chroot -u), and more wrappers are recognized
    (caffeinate, arch, sandbox-exec, nsenter, unshare, runuser,
    flock, watch, wsl, noglob, nocorrect), including .exe spellings.
    Code passed as a string to trap, su -c, flock -c, script -c,
    watch, cmd /c and PowerShell -Command is checked as a command line.
  • (( … )) arithmetic no longer reads << as a heredoc that hides the lines
    after it, and an xargs or find -exec replacement string used as the
    command or as shell code (xargs -I{} sh -c {}) counts as known only at run
    time.
  • task_shell_start and task gate commands are now checked against shell deny
    and ask rules, like any other shell command.
  • An "approve for session" grant for a shell command covers only flag variants
    of that command as written; a command with options before the subcommand,
    a chain, or nested code matches only an identical repeat.
  • The read-only shell surface for sub-agents rejects a word that starts with
    an unquoted *, whose matches could be read as options; quote the pattern
    or give it a path prefix (src/*.rs).
  • A trusted or allow prefix such as git status no longer covers options
    placed before the subcommand (git -c key=value status,
    git --exec-path=… status), nor a command that runs nested code or whose
    command word is resolved at run time. Such commands ask instead. Typed deny
    rules also match a path-qualified command word (/bin/rm).
  • Commands containing parentheses are no longer auto-approved as parallel
    read-only commands, since some shells treat them as glob qualifiers or
    command substitution.
  • Sub-agent worktrees stay under the per-repo .codewhale-worktrees/<repo>/
    root: an absolute worktree_path is now held to the same containment as a
    relative one, with symlinks resolved before the check. Any start that asks
    for a worktree keeps the approval card, even for a read-only role. A
    worktree_base starting with - is refused, and git worktree add now
    receives its path and base after --.
  • Fleet and reasoning-router names must be plain file names (optionally
    origin/name); a name with path separators or .. is refused before any
    file is looked up, through one shared check in the workflow crate.
  • pandoc_convert and image_ocr apply the same read deny-list and
    credential-store checks as read, through one shared helper, and pandoc
    always runs with --sandbox. This needs pandoc 2.15 or newer; an older
    pandoc gets an upgrade message instead of a conversion.
  • Computer Use: screenshot and zoom output paths must be .png/.jpg/.jpeg
    files inside the recordings directory, and zoom always crops the last
    captured raster instead of a caller-named source file.
  • Skill registry sync refuses an index key that is not a single path-safe
    name before it is used as a cache directory, the same check an installed
    skill name already gets.
  • A project's .codewhale/config.toml can no longer set notes_path; it is
    ignored with the other user-only keys. The note tool and /note refuse a
    notes file that is a symlink or whose directory resolves outside the
    workspace, and a note is flushed to disk before the tool reports success.
  • Diffs and shows that read repository content (codewhale review, the
    git_diff/git_show tools, verification, delivery, task attempt records,
    @diff and the runtime API diff routes) share one flag set: no external
    diff, no textconv, and submodules compared by commit only, so a submodule's
    configured filters and diff drivers never run. Reads of the working tree
    (git_diff, git_show, git_blame, commit planning, verification, task
    attempt records, @diff/@git) also run through the hardened review
    command, so the repository's clean filters do not run either. git_blame
    and read-only shell git blame skip textconv, and read-only shell git
    reads neutralize the repository's clean filters.
  • git_fetch refuses refspecs that would write a local branch or tag: a
    src:dst destination must be under refs/remotes/, and the two-word
    tag <name> form is rejected.
  • Opening a URL on Windows goes through the URL protocol handler instead of
    cmd /C start, so characters in the URL are not interpreted by the shell.
  • Skill downloads and the skills registry index are read through a streaming
    size cap shared with the MCP HTTP transport, instead of buffering the whole
    body before checking its size.
  • Durable runtime thread, turn and item ids carry a full UUID instead of 32
    random bits, so two records can no longer collide and overwrite each other.

Removed

  • Flags, settings and tool parameters that did nothing are gone
    (#6516). --output-mode
    is hidden. It is still accepted, prints a warning, and is ignored.
  • The dispatcher no longer exports DEEPSEEK_* copies of its CODEWHALE_*
    variables. A DEEPSEEK_* variable you set yourself is still read.
  • lane start and workflow run --runtime vm|ci are rejected before a lane
    is created. Older lane records for those runtimes still load.
  • The control socket's relaunch verb is removed; it always returned an
    error.
  • The speech tool drops stream. stream=true used to fail; it is now
    ignored, a complete audio file is written, and the result no longer carries
    "stream": false. The finance tool drops market, and a call that still
    passes it has it ignored.
  • [context].enabled, the seam-manager keys and
    tui.terminal_probe_timeout_ms no longer load; old configs that carry
    them still start. The [workshop] docs now describe bounded spillover
    instead of a synthesis sub-agent.
  • About 2,650 lines of workflow code that nothing ran are deleted: the replay
    executor, the review-repair loop and experimental search. The
    replay_diverged status they produced goes with them. The isolated
    Runtime Chat prompt and the legacy YOLO alias list each have one owner now
    (#6517).

Experience

  • Typing a first message with no model connected leaves a line in the
    transcript that says the message was not sent and opens the provider picker.
  • First run picks a chat-capable Ollama model instead of the alphabetically
    first tag, and says plainly when no model is available yet.
  • codewhale doctor leads and ends with one verdict and the next step, and
    gives the update command for how you actually installed Codewhale.
    Command-line usage and errors say codewhale.
  • The approval card leads with a plain summary of the action, such as
    "Run cargo test", and shows workspace-relative paths. The footer labels
    its values.
  • /status warns when the session's pinned model is no longer in its
    provider's live model list
    (#6035).
  • Error messages give one true sentence and one next step. The TUI's English
    copy says agent, Fleet, Permissions and Work consistently, help lists one
    summary per row, provider rows without a key say "needs key", /setup says
    what it sets up, and the pet tank rests when it is offline.
  • ACP clients can see the Permissions setting the server started with,
    including Full Access and how to turn it on, but cannot select it
    (#6310).
  • GET /v1/commands tells clients each command's argument shape, so they do
    not re-derive composer behaviour from the usage string
    (#6230).

Fleet and agents

  • codewhale fleet run <spec> --check runs every validation a real run would
    and stops there: nothing is created, launched or spent.
  • A queued agent says why it is waiting, for example when launches are
    throttled after provider rate limits, and when it stops waiting
    (#6277).
  • Read-only agents can run chained inspection commands (a leading cd,
    &&, ;, echo separators, 2>/dev/null), and a refused command now
    names the rule it broke and what to do instead. Durable Fleet workers
    accept the same read-only commands as in-session agents. An agent's time
    budget starts when it launches, and a queued agent that never gets a slot
    says it never started
    (#6015).
  • Stopping an agent that writes files keeps and names the work it had
    changed, as a budget stop already did
    (#5529).
  • workflow(fleet:) runs Fleets saved from the Fleet UI, and finds
    workspace Fleets under .codewhale/fleets.
  • The runtime API can stop a delegated agent run from the desktop.
  • A finished agent's answer is no longer cut off. Its row and its completion
    notification show the first sentence of its result instead of a
    ## Summary heading or its last tool, and opening the agent shows the whole
    result, or the full reason it stopped, even when no transcript was captured.
    Each agent also has one name: a workflow task's label or its dispatch name
    appears on the rows, the notification and the runtime API alike, never its
    internal id (#6565).
  • The mobile page shows the thread's agents: a strip naming each one, its
    state, and what it is doing or what it found, rebuilt when the page
    reconnects. Sub-agent prompt caching now counts toward the session, including
    cache-write-only reports. PRICE and /cache show the parent, agents and
    combined hit rates, each labelled and weighted by all input tokens,
    including cache writes,
    and the footer cache N% still means this conversation's own requests
    (#6565).
  • The dock's GIT, FILES and NOTES views are real. GIT shows the branch and
    where it stands against its upstream, the changes (with their paths one
    Enter away), linked worktrees and the last five commits, and it keeps
    updating during a turn while it is open. It says "not a git repository"
    only when that is true, reports failed status probes in the view and composer
    instead of claiming a clean tree, counts every unmerged path as a conflict, and only
    says there are more changed paths when the list is capped. FILES lists the
    files this session edited, including writes without diff receipts, and the
    files it read, with every unique path retained beyond the activity summary's
    twelve-path preview. FILES and its badge refresh from the current session
    even before TASKS opens. Edits with receipts show their size and open their diff.
    NOTES lists your /note notes. The git
    badge, the Git view and the model's git line now share one
    git status --porcelain=v2 call, so there are fewer git processes than
    before (#6565).
  • Background work tells you when it ends, even between refreshes. Batched
    shell notices cover explicitly backgrounded or detached commands; foreground
    results and old completions from before this TUI session stay quiet.
    Switching sessions does not repeat a completion notice. Batched
    notices count completed, failed and stopped work separately; shell commands,
    task prompts and errors stay in the app, away from lock-screen notifications.
    A running dev server no longer holds the notice back. Failed, killed and
    timed-out shells retain their outcomes, and the last eight finished shells
    stay listed per session. Finished tasks no longer look live or reopen the
    dock. Background labels follow the UI language, long agent previews keep
    their full-result pointer, and the quiet indicator uses the engine's clock
    and limit without marking an active tool as quiet
    (#6565).

Plugins

  • Codewhale no longer appends plugin recommendations to your messages to the
    model. Suggestions appear in one place, follow one switch and one budget,
    and never advertise built-in plugins, generic words or plugins for another
    operating system.
  • /plugin dismissals lists the plugins suggestions skip, and
    /plugin dismissals reset [<name>] brings them back.
  • Tools from reviewed plugins that declare themselves read-only no longer ask
    for approval on every call.
  • The bundled Computer Use plugin is 0.12.0, the published upstream release
    8435692 (#6303,
    #5856). On macOS the agent
    uses its own pointer and never drives your cursor. app_script refuses shell
    escapes. Clicks on irreversible actions such as
    pay, send or delete need confirmation. Consent decisions cannot ride inside
    run_actions or trajectory replay, and trajectories redact secure fields.
    Also new: a shared-computer control lease that pauses agent input while a
    person drives, and a browser attach mode for a shared Chromium. The vendored
    README no longer claims delegated agents share the Computer Use session; they
    never receive its tools.
  • The bundled first-party catalog pins marketplace revision
    ae3dd2255a9a266365c6125a084f511eb26bc04d. It lists Computer Use 0.12.0 and
    adds Codewhale for Chrome (Chromewhale) 0.3.0 as a developer preview: you
    load its Chrome extension unpacked, and like every catalog plugin it installs
    disabled and untrusted until you review it.

Release reliability

  • Upload the complete release into a private draft, verify every asset's size
    and SHA-256 digest, then publish. An interrupted retry cannot reuse stale
    same-size bytes. CNB and GHCR version tags follow canonical publication.
  • Ubuntu Lighthouse bootstrap requires trusted SSH source CIDRs or an explicit
    public-SSH opt-in before changing the host; malformed IPv6 and broad default
    networks are refused.

CI

  • State-touching command and session tests seal HOME, USERPROFILE and
    CODEWHALE_HOME onto temporary directories. Unsealed session, snapshot,
    artifact, composer-history, audit and built-in-plugin storage resolves into
    an isolated test root; unrelated worker threads cannot borrow another test's
    home seal. A child-process sentinel checks that known command-dispatch and
    session cases leave their ambient home unchanged.
  • Fork pull requests stay under the macOS runner limit and the Actions cache
    stays under its cap.
  • Release candidates and releases share one parity gate, and a release tag
    without a release-candidate receipt is refused.
  • Budget ratchets block same-repository pull requests unless the pull request
    updates the budget with a receipt.
  • A CodeQL advanced-setup workflow is ready for when the repository switches
    from default setup.

Contributors

  • @AdityaVG13 — supplied the discovery-cache priority correction adapted from #6393, keeping highest-ranked tools through cache overflow. Its broader echo and fork-inheritance draft remains open.

  • @7jrxt42BxFZo4iAnN4CX — reported indefinite questions cancelled by the TUI watchdog and supplied the timer evidence (#6872).

  • @hodeswildsmith455-boop — added OrcaRouter account sign-in with PKCE and its live chat catalog (#6867).

  • @LIghtJUNction — added reviewed plugin-provided AI routes with host-owned OAuth PKCE credentials and request-time authority checks (#6805).

Contributors and issue reporters are credited below, including
@cenab's provider report.

  • @Guan0923 — accepted case-insensitive HTTP(S) schemes in config doctor without rewriting the configured URL (#6819), and routed the Python and JavaScript execution tools through the session's execution policy (#6820).
  • @harryvgiunta — added Yolo-Auto as a bundled OpenAI-compatible host, starting on the vendor's recommended qwen3.8-flash model (#6408).
  • @asto18089 — contributed the integrated runtime liveness, context, search, JavaScript execution, stopship scout and pet repairs, preserving their original contributor commits (#6799); made context rule and chain-segment source labels repository-relative (#6739).
  • @qiuYliangM — made provider-bound project instruction and constitution labels stable across directory moves and kept their absolute paths in operator reports (#6799).
  • @zhuowp — supplied the process-scoped PowerShell execution-policy repair adapted for Codewhale, preserving machine and user Group Policy precedence (#6745).
  • @Andrea-Bruno — designed the Superfast Decision Gate and contributed its off-by-default shadow classifier (#6604, #6603).
  • @aiapienthusiast — added Cheaper Inference to the bundled provider catalog (#6761).
  • @gaord — let a client fork a thread at a named turn (#6580), let undo roll back files for the turn it is undoing (#6483), stopped resume and fork from duplicating threads and sessions (#6406), exposed user-defined provider routes to native clients (#6404), and kept a fork going when a turn lost its tool call (#6664).
  • @Lstarsky0 — moved the docs/work, legal, digest and FAQ pages onto the dictionary spine (#6405, #6417, #6499, #6574), tightened the Chinese-branching ceiling to 18 (#6403), and made Fleet publish without a two-link window (#6431). Also moved the constitution page onto the dictionary spine and kept its install link in the selected locale (#6733), wrapped diff and tool output at grapheme boundaries (#6829), and translated the context inspector rows twelve packs still shipped in English (#6831). Translated the session-only note after model switches across the complete TUI locale packs (#6875). Translated the route-save receipts (#6884) and /workspace replies (#6885), brought the Operate descriptions in fourteen packs up to date (#6886), restored lost accents in pt-BR, es-419 and ca (#6888), let semantic_truncate cut between Han and kana (#6887), and removed the unused session-only route-save choice (#6882).
  • @aboimpinto — restored a green Linux full-workspace test gate without loosening any test, twice (#6581, #6666). Completed the seventeen-command portable session group, including /structcopy (#6793).
  • @dajiaohuang — codewhale config set checks a known setting's value against its schema type before saving it (#6568).
  • @jayanthvee — reported and diagnosed that killing the npm launcher's node.exe ends Windows sessions without cleanup, with reproductions and fix directions (#6827).
  • @cenab — requested the Tsubasa provider row and supplied its endpoint, key and model values (#6695).
  • @BX166 — reported the AICraft provider row missing its key console, docs link and guidance, and supplied the values (#6616).
  • @Water-Run — ingested namespaced model-only catalog entries so models present only in the canonical models map reach the offering list (#6400), and retired the blanket dead-code allowance with its unused feature stages, tightening the budget to match (#6402).
  • @wuisabel-gif — designed the tool_call_after execution-receipt contract and its tests on a reference branch, which landed re-implemented on the current hook seam (#6689, #6713).
  • @SparkofSpike — let making room survive a provider request-body limit (HTTP 413) by shrinking, then replacing, inline images for that one summary pass (#6642). Translated seventeen Tier-2 guides and thirteen developer and internal docs into Simplified Chinese, and connected the localized documentation (#6662, #6663); added regression coverage for rejecting unknown website locales before dictionary lookup (#6786); made the pinned prompt header follow the viewport's turn and jump on click (#6830).

See CHANGELOG.md for full notes and docs/CHANGELOG_ARCHIVE.md for older releases.

Don't miss a new Codewhale release

NewReleases is sending notifications on new releases.