github Dicklesworthstone/ntm v1.37.0

6 hours ago

Spawns can enforce per-agent-type fleet limits that hold across processes, Claude agents launched by ntm now carry the safety policy hook, robot work reads are verified against the tracker, and ntm's integrations with bv, br, cass, cm, caut, caam, ubs, xf, ms, jfp, slb, ru, rano, pt and dcg now read those tools' real output. Also fixes GitHub issues #283, #335, #336, #337 and #338, and wires up several features that were configured or documented but never ran, among them ntm spawn --privacy, session timelines, ntm level, context warnings, compaction recovery and the session coordinator.

Upgrade notes

  • --robot-jfp-install and --robot-jfp-installed are retired. jfp 1.0.x moved skill management to jsm and answers deprecated_command, so both flags could only return NOT_IMPLEMENTED while --robot-capabilities advertised them. They are removed together with --jfp-project, their schema, pagination and capability entries; passing one now fails with INVALID_FLAG (unknown flag). --robot-jfp-status and the prompt list, search, show, suggest and export surfaces are unchanged (83c391ec).
  • Config keys that never had an effect now fail to load. [context_rotation] enabled, warning_threshold, min_session_age_sec, require_confirm and default_confirm_action, and [agents.plugins], had no reader (the live rotation trigger is [rotation] usage_percent_threshold). A config that still sets one fails to load with a line per key; ntm doctor lists them and ntm config migrate deletes them, keeping a backup. Pending context rotations are manual-confirm unless [rotation] auto_confirm is set: [context_rotation] confirm_timeout_sec is now only how long a request stays open, its default rises from 60 to 600 so an unattended request is not re-announced every minute, and 0 means the default rather than "no auto-rotate" (ef529fec, e209108f).
  • Schema migrations 025_send_operations_kind and 026_send_operations_claim_fence apply on first run. They add columns to the durable send-operation table. Rows left in_progress by an earlier version are marked as started, so retrying one of those operation IDs returns OPERATION_OUTCOME_UNKNOWN instead of delivering again; completed receipts are unchanged (c631a703, 209446ea).
  • Spawn admission is serialized across local processes and now covers ntm add. With the default [spawn_pacing] settings (enabled, positive agent_caps), every non-dry-run --robot-spawn, REST spawn, swarm_spawn job, ntm add and ntm scale scale-up evaluates the fleet while holding a per-user lock (ntm/spawn-admission/fleet.lock under the user cache directory) and waits up to 30 seconds for it; a lock still busy after that returns RESOURCE_BUSY with reason spawn_admission_busy. ntm add, scale-up and the dashboard's add action are now checked against the same summed agent_caps budget and pressure checks as robot spawn. A tmux inventory that cannot be read defers the launch (agent_inventory_unavailable) instead of counting as an empty fleet. Count-capped launches against a remote tmux server (--ssh) or on platforms other than Linux and macOS are refused (spawn_admission_unavailable) rather than run without the lock: run the command on the target host, or set [spawn_pacing] enabled = false to keep the previous unchecked behavior (e42387c3, ba1d701b, fa5a66b9).
  • ntm serve, ntm serve --web and ntm web apply the selected config to API spawns. REST spawn and swarm_spawn jobs called the spawn engine with no configuration, which bypassed configured agent commands, model defaults and fleet limits. They now load the file selected by --config / NTM_CONFIG, plus the launch project's overlay, strictly at execution; a missing explicitly selected file or an invalid config fails with INVALID_FLAG before any pane is created. Code that embeds serve.New must call ConfigureSpawnPolicy to get this behavior (77b4e56d).
  • Claude agents launched by ntm now get a PreToolUse safety hook through --settings ([safety] claude_policy_hook, default true), plus the dcg and rch hooks when those integrations are enabled and installed. A custom [agents] claude command whose last shell command is not the claude invocation launches without the hooks and logs a warning naming the two fixes: --settings {{shellQuote .ClaudeSettings}} in the template, or claude_policy_hook = false. If you ran ntm safety install before, run ntm safety install --force: the old script was never registered with Claude Code, and the new scripts delegate to ntm safety check --hook and are registered in ~/.claude/settings.json. ntm safety status reports an old install as "Script present but not registered" (85b8a777, 3d64eea1, 6b9b6512).
  • The pre-commit hook ntm init installs now enforces Agent Mail file reservations. It runs ntm guards check --staged, which refuses a commit touching files another agent holds exclusively and names the holder; if Agent Mail is unreachable it warns and allows the commit unless NTM_GUARD_STRICT=1. Existing managed hooks are not changed automatically: run ntm guards install to regenerate one in place (it keeps the beads sync and UBS steps instead of refusing or replacing the hook). The guard also reaches a token-protected or non-default Agent Mail server configured in [agent_mail], where it previously failed open on every commit (67f13b56, c0a50fb8).
  • Every spawned session now runs the session coordinator inside its monitor. Assignment maintenance (renewing the leases of delivered assignments, releasing leases and claims when the bead closes, recovering work from lost panes) runs every 30s, so a one-shot ntm assign no longer loses its file lock after an hour. Features enabled with ntm coordinator enable (digest, conflict negotiation), mail nudges, [rotation] usage_percent_threshold and [integrations.caam] auto_failover now act in every session without ntm coordinator run, and config is re-read every 15s. [coordinator] auto_assign = true now runs in every session monitor, not only under ntm coordinator run: idle agents in every spawned session get ready beads assigned automatically, subject to the assignment safety policy. Remove the key or run ntm coordinator disable auto-assign if you set it for one-off coordinator run use. [coordinator] conflict_notify now defaults to false, since it mails every holder of overlapping reservations project-wide at high importance, and the session monitor never runs it even when it is set: config files written by ntm config init, config edit or config reset up to v1.36.1 contain conflict_notify = true, which cannot be told apart from an opt-in, so conflict notification stays where it was, under ntm coordinator run, and ntm coordinator status says so. ntm coordinator run and a reservation-maintaining ntm assign --watch take the session over while they run; ntm coordinator status gains a Runtime section (75d083b7, 79646496).
  • --robot-bulk-assign reserves each bead's files by default, like ntm assign. The new --reserve-files defaults to true: explicit --reservation-paths win, otherwise each bead's paths are discovered from its title and description before the claim. A bead that names no files, or an unreachable Agent Mail, refuses that assignment unclaimed with RESERVATION_REQUIRED. Pass --reserve-files=false to opt out; combining it with --reservation-paths or --require-reservation is INVALID_FLAG (38f24ba2).
  • Pipeline resume no longer reruns a command step interrupted by a worker crash. Command launches are journaled before the shell starts. On resume, a command whose process group may still be running or may already have had side effects stops the resume with a message naming the step and recorded process ID; after inspecting it and stopping any surviving processes, use --mode=restart-failed (or --keep-state=false for the whole workflow) to repeat it deliberately. Runs saved by an earlier version with an unfinished command stop the same way instead of rerunning it (0023dfd7).
  • ntm workflow run now acts on on_agent_crash and on_agent_error. Nothing raised those faults before and restart_agent did nothing. The runner now checks every workflow pane on each poll (pane gone or dead, agent CLI exited to its shell, rate limit, auth error, blocking gate) and runs the configured action; unset values default to notify, and the built-in templates set them (specialist-team restarts and pauses, red-green and review-pipeline pause, parallel-explore notifies). Restarts go through the ntm respawn engine, re-send the current stage prompt when the pane had already received it, and are capped per agent per run by [resilience] max_restarts, so a crash loop stops the run. on_timeout = restart_agent is now rejected at validation (d18d55a2).
  • Robot error codes and REST statuses. A retryable work-source change on --robot-assign, --robot-bulk-assign and --robot-spawn --spawn-assign-work is STALE_WORK_COORDINATION (REST 409) instead of INTERNAL_ERROR (GH #283). A project with no .beads returns DEPENDENCY_MISSING with a br init hint from the bv-backed robot analyses and 404 BEADS_NOT_INITIALIZED from REST beads endpoints; an unindexed cass returns DEPENDENCY_MISSING from the robot cass surfaces and 503 CASS_UNAVAILABLE from REST. --robot-mail now fails (exit 1; PERMISSION_DENIED for a rejected token, otherwise INTERNAL_ERROR) when it cannot read the project's agents, instead of reporting success with an empty list. IDEMPOTENCY_CONFLICT and OPERATION_IN_PROGRESS map to REST 409, and SENSITIVE_DATA_BLOCKED to 403, instead of 500. New: DESTRUCTIVE_COMMAND_BLOCKED (REST 403) on robot and REST sends that dcg refuses, and OPERATION_OUTCOME_UNKNOWN (REST 409) for a retried send or interrupt whose earlier delivery may have started (a4a655a9, 34b5dad2, 4f168194, 289cc7b3, d394c3e2, c631a703, 4c8832cf, 209446ea).
  • Robot and REST output fields. --robot-agent-health lists user shells and unattributed panes under a new non_agent_panes map instead of grading them; fleet_health.total_panes counts agent panes only, and overall_grade is "N/A" when no agent pane is graded (GH #335). --robot-cass-insights returns agents[] and workspaces[] buckets and drops topics; REST /cass/insights returns bucket lists, /cass/status reports version and database_size in place of the always-zero index_size, and /cass/timeline rejects ?workspace= with 400. --robot-rano-stats drops the request and byte fields (rano never provides them), adds database, and returns DEPENDENCY_MISSING when the observer database is missing. --robot-dcg-check agent hints drop the never-populated safer_alternative. REST checkpoint verification drops checksums_valid, which was always true; verification details now carry checksum_status (verified, failed or unavailable) and checksums_checked. Pending context rotations no longer carry default_action in --robot-context or ntm rotate context pending (97271c76, 9037eb9d, 4db8ad28, 9f2330f5, c7f2e127, 0b589055, 4cff2514, 70b51d1f, 88d6e9ff, e209108f).
  • Robot flags that are now applied or rejected. --robot-interrupt and --robot-ack honor --type; they interrupted or watched every agent in the session. --robot-send rejects the --robot-route modifiers (--strategy, --route-strategy, --last-agent, --route-type, --route-exclude) with INVALID_FLAG instead of sending to every matching pane. --brief with --fix or --diagnose-pane is INVALID_FLAG. --robot-slb-deny requires --reason, and slb approve and deny need a reviewer session in SLB_SESSION_ID / SLB_SESSION_KEY. --xf-mode takes xf's own modes (lexical, semantic, hybrid, two-tier) and --xf-sort its orders (relevance, date, date-desc, engagement); other values are INVALID_FLAG. --spawn-preset now runs the recipe and cannot be combined with count flags (59c4a5cc, 4cf7c91a, 44e34a46, 2d21c77c, dfaf8eac, 7b9eba41, 5230731a).
  • ntm send, ntm logs and ntm spawn flag changes. ntm send --route now enables smart routing; it was ignored unless --smart was also given, so the prompt went to every matching pane. --smart or --route combined with --all, --skip-first, --project, --batch or --distribute is an error, and --route=explicit requires --pane. --template refuses to send when a placeholder is left unfilled instead of sending literal {{var}} text, and cannot be combined with --batch, --distribute, --project or --codex-goal, which silently dropped it. ntm logs --since is removed; it never filtered anything. ntm spawn rejects an out-of-range --stagger-delay in every stagger mode (44e34a46, 67f13b56, 2c902d45, c3b640f0, d0b39b04).

Added

  • Per-agent-type fleet limits (e42387c3, ba1d701b, fa5a66b9, 77b4e56d): the opt-in [spawn_pacing.agent_type_limits] table (keys claude, codex, gemini, antigravity, grok, omp, opencode) bounds each agent type across every session the tmux client can see, in addition to the summed agent_caps budget. A mixed request is evaluated as one batch before any pane is created; admission reports sorted agent_type_limits rows (running, requested, projected, limit, remaining headroom, blocks_request), and a type-cap refusal uses reason agent_type_limit_exceeded. Count-capped launches hold the cross-process admission lock from fleet inventory through the last launch, so two cooperating processes cannot spend the same headroom; admission.serialized says whether the decision was made under the lock. docs/spawn-agent-limits.md documents the scope and what the lock does not cover.
  • Recipe-backed robot spawns and readiness-gated jobs (5230731a, 35a2b200, 355427be): --robot-spawn --spawn-preset=NAME and swarm_spawn jobs with preset expand the recipe (builtin < user < project, read from the launch directory's .ntm/recipes.toml) into per-instance models and efforts, render every instance before launching, and admit the expanded counts; preset_used names the recipe. swarm_spawn jobs accept launch_ready_timeout to start agents one at a time, each passing its readiness check before the next launches. A job with no working_dir uses the server's selected project, a relative one resolves against it, and queued jobs keep the absolute directory captured at admission.
  • --robot-spawn initial prompts and stagger pacing (d0b39b04, 116098c6) (refs GH #251): --spawn-prompt / --spawn-prompt-file deliver a first prompt after the readiness wait through the same dispatch path as work assignment, prefixed to each work prompt with --spawn-assign-work. --spawn-stagger-mode=none|fixed|smart and --spawn-stagger-delay pace prompts and assignment dispatch; output gains a stagger plan (previewed with --dry-run), prompt_deliveries[] (PROMPT_SEND_FAILED on any failure) and assignments[].delivered_at. A new [spawn] table (stagger_mode, stagger_delay) sets defaults for both ntm spawn and --robot-spawn, which now share one stagger planner, and ntm spawn --json reports the resolved mode and interval. --robot-spawn also accepts --spawn-wait and --spawn-assign-work for Grok panes, which it had refused with NOT_IMPLEMENTED.
  • Retry-safe --robot-interrupt and fenced send operations (c631a703, 209446ea) (refs GH #245): --op-id, and the REST Idempotency-Key, now claim a durable operation as --robot-send does. An identical retry replays the recorded outcome without a second Ctrl+C, a conflicting reuse is IDEMPOTENCY_CONFLICT, and a live claim is OPERATION_IN_PROGRESS; --robot-send-receipt returns interrupt receipts. Send and interrupt operations record their final payload digest, targets and a dispatch boundary under a rotating claim token before any input, so only claims that never started are recovered, and one that may have delivered returns OPERATION_OUTCOME_UNKNOWN with its receipt instead of typing again.
  • [assign.work_source] policy (3bb33ade, a4a655a9, 34b5dad2, edb1a14a) (GH #283): required_ref, require_clean and program_labels restrict what ntm assign, --robot-assign, --robot-bulk-assign, --robot-spawn --spawn-assign-work, the coordinator and ntm work-snapshot treat as dispatchable. A project .ntm/config.toml can only tighten the user's policy. Stale-source refusals carry a work_source_mismatch receipt with the expected and observed tracker JSONL SHA-256 and HEAD.
  • Assignment prompts carry CASS history and CM rules (86372035, 693f775c): ntm assign (and ntm coordinator assign), --robot-bulk-assign, --robot-spawn --spawn-assign-work and the coordinator's auto-assign enrich each new assignment prompt the way ntm send does, with the same precedence (--no-cass > --with-cass > [cass] / [cass.context]; --with-memory or [memory] send_injection), querying by the bead's title, labels and description. Enrichment happens before the durable intent is recorded, so retries and recovery replay the recorded prompt without querying cass or cm again. Each assignment reports cass_injection and memory_injection; a missing or failing cass or cm never blocks an assignment.
  • Context warnings and context-pressure attention events (1207374b, 57305163, 3673e13e): alerts.context_warning_threshold (default 75%) now raises a context_warning alert for each agent pane at or above it, read from the pane's own session transcript; the threshold was documented but nothing produced the alert. It appears in --robot-alerts, robot status, snapshot and markdown, and the dashboard, and resolves when usage drops. The session monitor also appends per-pane context-pressure events to the attention feed every minute, so the context_hot wait condition can fire, and the reservation and file conflict conditions now fire on the coordination problems the projection refresh reports. A pane that escalates from the warning threshold to action-required surfaces at once instead of after the ten-minute dedup window.
  • Compaction recovery runs in the session monitor (c5a76431, 2b27203c, 3b989f01): automatic recovery after an agent compacts its context ran only while a dashboard was open. The resident session monitor now detects fresh provider completion banners and sends the recovery prompt under the [context_rotation.recovery] policy (enabled, prompt, cooldown, attempt limit), rechecking the pane's process and identity and waiting for fresh idle evidence before delivery, and never replaying an uncertain write. The dashboard only displays compactions; its "recovering" indicators, which could never show, are removed.
  • SLB decisions flow back into ntm approvals (04ac1fb9): an approval mirrored into slb is approved or denied when a reviewer decides it in slb (approver slb:<reviewer>, the reviewer's comments as the deny reason), through the same two-person and approver-list checks as ntm approve. ntm approve, REST /safety approvals and the force-release gate all see the decision.
  • Checkpoint artifact integrity (a6d866cf, 88d6e9ff, 10ad3518, 70b51d1f): checkpoints record a SHA-256 and byte count for scrollback, git patches and git status before their metadata is written, and a resave keeps the original fingerprints. ntm checkpoint verify and REST verification check them; restore and import refuse a corrupted payload before replacing a session, launching agents or injecting context, even with force or overwrite; export verifies its source bytes and writes a manifest for what it exports. Checkpoints from earlier versions still load, with an integrity-unavailable warning.
  • ntm agents stats and ntm agents recommend use recorded outcomes (104ba9a3): completed and failed assignments are summarized per agent type across sessions, including assignments already cleared from the ledger, which are now archived to the session's outcomes.jsonl. Stats show the observed success rate, completion count and mean time, and recommend applies the observed rate instead of a seeded prior. Agent types with no outcomes print as unmeasured.
  • affinity routing strategy (44e34a46, 07e130d5): ntm send --route=affinity and --robot-route --strategy=affinity pick the available agent holding live Agent Mail reservations on the largest share of files named in the prompt, and fall back to least-loaded (reported as fallback_used) when no agent holds any. --robot-route accepts --msg / --msg-file, required for affinity, and shows affinity_match per candidate. Pipeline route: steps accept every router strategy except explicit, not only least-loaded, first-available and round-robin.
  • ntm doctor reports Claude hook coverage (116098c6): Safety Defaults gains "Claude agent hooks", listing the PreToolUse hooks ntm-launched Claude agents carry and whether ntm safety install registered the policy hook, and warns when Claude agents would run unprotected. JSON adds safety_defaults.claude_agent_hooks, claude_policy_hook_registered and claude_agent_hooks_warning.
  • --robot-agent-health samples process triage without a resident monitor (50c26991): a standalone call takes one bounded passive pt snapshot through the monitor's own sampler, or reuses a running monitor, and reports pt_observation_mode (monitor or snapshot) with explicit timeout, unavailable and empty states. Shell-only, dead and disabled selections skip pt.

Fixed

  • User shells were graded as stalled agents (GH #335): the tmux projection ran every pane through the agent classifier, so a plain shell that went quiet past the stall threshold became an error, a shell-only session went critical in --robot-status, and --robot-agent-health graded shells 0/F. User and unknown panes now get a neutral busy/idle/active classification and still count in session totals; session health uses agent panes as the error denominator (97271c76, 9037eb9d).
  • ntm assign --auto claimed beads it could not reserve (GH #336): reservation paths came from the bead title only, so a bead that listed its files in the description was claimed, failed with "file reservation is required but no reservation paths were defined", and stayed claimed with no prompt sent. --auto and direct ntm assign <bead> --pane now resolve paths from the live title and description before claiming and refuse a bead with no paths unclaimed. When a description declares owned outputs (## Owned outputs, Files to modify:, **Owned files** and similar), only that section is reserved, including its sub-labels and sub-headings, so files the bead only reads do not become exclusive reservations. Path extraction accepts a Markdown backtick before a path and no longer turns docs/notes.md into a docs/notes/**/* glob (89b32d9a, 15e4dd56).
  • ntm status estimated Claude context from scrollback (GH #338): a Claude pane's scrollback is a sliver of its session, so a six-pane swarm at 28-50% of its window read 4-6%. Status now reads each pane's own session transcript. On Linux, ntm binds a Claude process to its session through Claude Code's ~/.claude/sessions/<pid>.json, checking the process start time, which attributes transcripts correctly when several Claude panes share a directory; previously every such pane got the newest transcript, or none. --robot-context, the snapshot, the coordinator's rotation trigger and native compaction use the same binding (in that layout the trigger never fired and compaction always refused). ntm status --json adds context_source (status_bar, transcript or scrollback_estimate), and text output marks estimates. Current Claude models (Fable 5/5.1, Mythos 5/5.1, Opus 5.5/5/4.8/4.7/4.6, Sonnet 5.5/5/4.6) are listed with a 1M window; Haiku 4.5 and Opus/Sonnet 4.0/4.5 stay at 200K (3da30a92, e4d911a6).
  • --robot-snapshot and --robot-status advertised blocked beads as ready (GH #283): work rows were served on age alone, unscoped by project, and a blocked bead that triage did not list was counted as ready. Both now read work only from the source-verified observation (tracker JSONL digest, checkout HEAD, peer reservations) that the projection refresh just collected; if it no longer matches, one new observation is taken, and any other failure reports work as unavailable with its reason and zero ready. A rejected observation keeps its diagnostics without restoring readiness (67a29ac0, 6034ebfd).
  • Robot assignment and spawn reservations (38f24ba2, c99121e1, bb92f4a2): --robot-bulk-assign reserved files only with --require-reservation plus explicit paths, so two agents given overlapping beads through the robot API could edit the same files with no lock; it now shares ntm assign's discovery (see Upgrade notes) and reports a reservation object per assignment. --robot-spawn --spawn-assign-work --require-reservation without --reservation-paths discovers the bead's paths instead of failing. Robot conflict reports and work-coordination reservation_conflict problems compare glob languages the way the coordinator does, so src/*/main.go and src/service/*.go are reported as overlapping and src/*.go no longer matches src/deep/main.go.
  • Claude turn detection (6eb7a776, 482e4df6): Claude Code now ends a turn with a suffixed line (✻ Crunched for 10s · done 11:09 PM · 1 shell still running) that ntm did not recognize, so idle panes read as working and a failed tool call above the line made a completed turn an error. Inline tool output (⎿ Error: Exit code 144) no longer raises agent_error alerts or ERROR in --robot-activity, and idle Claude panes with queued text, only the footer, or a stale spinner classify as WAITING, matching --robot-is-working.
  • Claude agents never got the safety hooks (85b8a777, 3d64eea1, 6b9b6512, 8ba5eedb, 2db8f1d4): ntm safety install wrote a hook script that Claude Code never ran because no settings file listed it, while ntm safety status called it active; spawn and add passed the dcg and rch hooks in a CLAUDE_CODE_HOOKS variable Claude Code does not read; and ntm send skips dcg for Claude panes on the assumption that the hook covered them. Every Claude launch path, including spawn, add, robot spawn, restart, restore, rotation, swarm and CAAM account recovery, now passes the hooks with --settings, and the installed script delegates to a fail-closed ntm safety claude-hook instead of parsing with jq and allowing everything when jq was missing. Hook and wrapper refusals are logged with their tmux session, pane and pattern, so ntm safety blocked lists them, and both refusals and dcg-blocked sends now count in ntm metrics blocked_commands, which stayed at 0. The hook's tmux session lookup is bounded to 2s so a stuck tmux server cannot hang an agent's tool call. ntm safety install and the REST install endpoint ship one copy of the scripts.
  • dcg did not check robot or REST sends (4c8832cf): only ntm send consulted dcg, so --robot-send, --robot-send --track, --robot-interrupt --interrupt-msg and REST agent sends delivered destructive command lines to non-Claude agents unchecked. They now share the guard, refuse when the check fails or times out, and interrupt refuses before sending Ctrl+C. A dcg probe cut short by cancellation is no longer cached for five minutes as "unavailable", which had switched the guard off for every sender.
  • ntm spawn --privacy kept nothing private (a48a32ab): the setting lived only in the spawn process, so ntm send, ntm serve and the coordinator wrote prompt history, event logs, audit logs and checkpoints for private sessions. Spawn now records the setting on the tmux session before any agent starts and every ntm process consults it; a recorded private answer is never downgraded by a later failed lookup.
  • Recorders that nothing called (77dffca1, 28f35969, 078a2473, 30d309cd, 7780e0bc, aa5515a2, 595a6c6f): ntm level always showed "Commands run: 0"; successful interactive commands are now recorded (robot and JSON calls, shell completion and hidden commands are not). ntm timeline list returned nothing; the per-session monitor now records agent state transitions and persists the timeline, merging checkpoints so a concurrent ntm kill cannot wipe it. The dashboard's process-health panel was always empty; the dashboard now starts the pt monitor when [integrations.process_triage] is enabled, and readers get the monitor that was actually started. The state database was never garbage-collected; ntm serve now collects at start and hourly, and robot commands at most once an hour (tracked by a state.db.gc stamp file). recovery.max_cm_rules and recovery.max_cm_snippets were ignored in favor of fixed 10/3 limits.
  • Events were lost before reaching webhooks or the attention feed (15bcade8, d952e3ea, 3673e13e, ef1c9329, a9dbf877): the session monitor, coordinator and server now persist event-bus activity (agent crashes, restarts, rate limits, session endings, coordinator actuations) to the durable attention feed read by robot attention and incremental snapshots. Short-lived commands drain the events they emitted (spawn, add, kill, assignment) before exiting, and the webhook bridge waits for them before unsubscribing; before, the webhook for those events never fired. Terminal and robot ntm kill now share one implementation, so both reap orphaned agent processes and emit session_killed, and the paired kill and session-end events are recorded as one session destruction.
  • Robot and REST spawn skipped Agent Mail identity registration (fa486a81, 54dc72e6, 6ea1089b): only ntm spawn, add, adopt and relaunch ran the per-pane identity coordinator, so reservation-bearing assignment through the robot API failed with "pane has no canonical Agent Mail identity", pane badges never appeared, and ntm lock in such a session failed with "session has no Agent Mail identity". Robot and REST spawn now register each pane before its launch and the session identity after, and report agent_mail (with agent_map and session_agent) on the spawn envelope; agents[] entries carry pane_id and agent_mail_name.
  • Agent Mail clients ignored the configured endpoint (c0a50fb8, 1dbb0199, d394c3e2): about 20 call sites, including the pre-commit reservation guard, robot bulk-assign reservations, robot mail check, pipeline mail steps, handoff, build slots and scanner notifications, read only AGENT_MAIL_URL / AGENT_MAIL_TOKEN, so against a token-protected or non-default server they got 401 or the wrong endpoint. Every client now starts from [agent_mail] url and token. With no token configured, ntm uses the bearer token am stored in mcp-agent-mail/config.env, and only for a localhost or loopback endpoint. --robot-mail no longer reports success with an empty agent list when it cannot read the project's agents.
  • Robot flags that were accepted and ignored (2d21c77c, 59c4a5cc, 4cf7c91a): --robot-dismiss-alert ignored --session and --all, so --dismiss-all --session=X dismissed every session's alerts; --robot-palette ignored --category, --robot-activity --type, --robot-tokens --agent, --robot-cass-search --since, --robot-markdown --session and --robot-inspect-pane --lines; and the shared --limit was ignored by --robot-triage, --robot-search, --robot-label-attention, --robot-file-beads, --robot-file-hotspots, --robot-file-relations, --robot-beads-list and --robot-files. --robot-interrupt and --robot-ack ignored --type; REST interrupt now accepts agent_types.
  • --robot-metrics reported every agent at zero (916b5d35): its activity fields had no writer and --period was echoed but not applied. Prompts now come from prompt history (robot and REST sends are recorded there too, redacted), crashes and restarts from the session monitor's analytics events, tokens and context percent from transcripts, uptime from the pane shell's start and session duration from tmux. --period bounds the counters and never reaches back before the session's creation, and unmeasured lists what could not be measured on that call. ntm health and the dashboard health badges read uptime and restarts from the same sources instead of an in-memory tracker nothing wrote to.
  • ntm send --template sent literal placeholders (2c902d45): templates rendered once before fanning out, so every pane received "You are Agent #{{agent_num}} ({{agent_type}})" and "br update {{bead_id}}". Templates now render per target pane with its agent number, type and position in the batch, and the new --bead <id> fills bead fields from br show in the session's project. Dry runs, the JSON preview, dcg checks and prompt history use the per-pane text.
  • Smart routing ignored the requested targets (44e34a46, ca2d958c): ntm send --route was ignored without --smart, and --project, --batch, --distribute and --all silently dropped smart routing. Smart routing also ignored variants (--cc=sonnet), --tag, multiple types and every type other than cc, cod, gmi and agy, so a Claude-targeted prompt could go to a Grok pane; it now chooses among exactly the panes a broadcast with the same filters would reach.
  • Quota and account data (eba325ac, 3b55b1bf, e3699309, 4cff2514): ntm's caut parser never matched caut's published v1 contract, so usage-aware health, --robot-monitor, --robot-quota-status and the dashboard Usage & Costs panel saw no provider data; it now follows the contract, queries one provider per call, and refreshes the usage cache when it is older than a minute (its poller had been removed). --robot-quota-check overlays CAAM's live windows the way --robot-quota-status does (refs GH #319) and reports a provider with no data as NOT_FOUND rather than PANE_NOT_FOUND. A caam account in cooldown (cooldown.active, or health.launch_usable=false) is treated as unusable, so coordinator failover no longer switches a pane onto it, and account status surfaces show the cooldown. A failed caam activate reports caam's reason, and the rate-limit fix hint names ntm --robot-switch-account=<provider>.
  • ubs scans that failed were read as clean (b1215710): a failed or refused scan (exit 2) or one that scanned nothing (exit 3) passed the pre-commit hook, and ntm scan --update-beads closed every open ubs-scan bead. They now return SCAN_FAILED or report that nothing was scanned, and close nothing. Findings are also read from scanners[].findings[] (a file with two critical eval() hits had produced 0 findings), and language exclusion passes --exclude-langs. The scan bridge, the REST beads label filter and claim evidence use the br list flags br accepts (--label, one --status per value).
  • bv and br output (54a6847a, 403ccff6, 843b5673, 94e66f60, f4e78848, 0d917335, 540ca2a9, 6fbb660b): bv emits dependency cycles as arrays of issue IDs, and any cyclic graph made --robot-insights fail to decode, so --robot-graph returned INTERNAL_ERROR, the cycle alert never fired and --robot-health lost its bottlenecks and keystones. --robot-suggest, triage project_health and the file-beads, hotspots and relations surfaces decode bv v0.25.2's shapes. ntm work history requests a bounded window (--since, default 30d; the full history exceeded the 10MB output cap) and decodes bv's per-bead history, ntm work burndown renders issue and day counts instead of "0/0 points", and ntm work graph --format=dot|mermaid prints the diagram instead of the JSON envelope. Blocked dependencies are read from the br blocked --json envelope (recovery prompts and robot health had none), scanner beads are created with br create --json so project-prefixed IDs are linked instead of producing a duplicate bead on retry, and dependency neighbors no longer include the queried bead itself.
  • cass and cm (407d4479, ffceede9, c7f2e127, 4db8ad28, 9f2330f5, 9b811053, 166ed3ef, 13ab0e8f): context injection (ntm send --with-cass, ntm cass preview, robot sends) treated cass scores as an absolute 0-1 scale, and default hybrid search scores are about 0.016, so no hit passed the 0.7 threshold and injection added nothing. Scores are now relative to the best hit in each result set. ntm cass timeline and GET /cass/timeline decode cass's range/groups/sessions shape (the table was always empty and REST entries had no timestamps); ntm cass insights and --robot-cass-insights request cass's aggregation buckets and honor --since (default 7d); ntm cass status reports cass's real version, database size and counts. ntm search --session resolves the session to its project directory. Ideation reads cm rule text from content. ntm send cass enrichment honors [cass] timeout (15s when unset) and bounds pipe waits, so a wedged search cannot hold dispatch.
  • xf, ms, jfp, slb and ru (3535d31a, 7b9eba41, bc9ae683, fa78a284, 1d54be4a, dfaf8eac, c0ed58ef, 36554659, 8ffddc76): the palette's xf search (ctrl+k) passed a removed flag and always failed; xf search and health now use --format json and xf stats, and the robot xf flags forward limit, mode and sort. The ms adapter uses ms's global -O json and decodes its results and skill envelopes (skill discovery had failed or returned nothing), and --robot-ms-search returns a typed skills array. jfp list, search and suggest envelopes are unwrapped and counted, and --robot-jfp-update runs jfp refresh. --robot-slb-approve and --robot-slb-deny could never succeed; requests and reviews now go through slb sessions and reject, with PERMISSION_DENIED for a missing session, self-review or bad key and NOT_FOUND for an unknown request, and errors never include the session key. --robot-ru-sync reads ru's data.repos envelope, adds a failed list, maps ru's exit codes to RU_CONFLICTS, RU_PARTIAL_FAILURE and dependency or interrupted-sync errors, and lists a repo once even when ru reports it twice; ru's capability probe reads its help from stderr.
  • dcg rule names and tool capability probes (4cff2514): --robot-dcg-check reported no rule because ntm read fields dcg never emits; it now reports dcg's rule_id as rule_matched and its explanation as the suggestion. Capability probes read --help from both output streams, so ubs, pt and rano report their robot modes.
  • rano and pt (6ecedcc0, c5f3c567, 0b589055, 4df918ca, 7be83421): ntm called rano and pt subcommands that do not exist. --robot-rano-stats and the dashboard Network panel now read rano export --format jsonl from the database set by the new [integrations.rano] sqlite_path (default observer.sqlite), count connect events per pane (joined by pane ID, not title), and show connection counts, last connection and provider tags instead of always-zero traffic. Process triage uses pt agent watch --once --threshold low --format jsonl, never applies anything, and feeds the resident monitor, the dashboard and --robot-agent-health again. ntm serve no longer logs "alert channel full, dropping alert" for every pt alert past the 100th.
  • The web Agents page showed no agents (b92d5fc9): GET /api/v1/sessions/{name}/agents and the legacy agents endpoint answered with no agents while agents ran (one read a table nothing writes; the v1 route was shadowed by another). Both now list live agent panes enriched with projected status, state_reason, last_seen, health and current task. A hung tmux falls back to projection rows with a warning or answers 504, and a stopped session the store knows returns an empty list instead of 500.
  • ntm swarm launched Codex into a rate-limited account (b73b69c8): the swarm launcher never loaded the rate-limit history that ntm spawn and ntm add use; it now waits out each pane project's recorded Codex cooldown before launching.
  • Concurrent agent additions could lose session manifest records (2656f647): per-agent manifest updates were guarded only within one process. Updates, saves and deletes now take a per-session lock across processes (bounded to 5s) and reject malformed manifests or a mismatched session identity instead of overwriting them; non-Unix systems fail closed.
  • Pane badges on tmux 3.4 (9a2f48f0): tmux 3.4 escapes the \x1f field separator in format output, so badge reads returned the whole record and window listing failed, which disabled badges on Ubuntu 24.04's default tmux.
  • Pipeline step failures lost their evidence (af9dd340): prompt and template steps that failed to send or timed out stopped recording the pane's last 50 lines and agent state for ntm pipeline status and the pipelines API; they record them again.
  • ntm personas was registered twice and appeared twice in help and shell completion (520fa799).
  • Docs (e12b2afe, 8065d57a, 375641c1, 448cd9f8, fa3b17db): robot API docs named flags that do not exist (--robot-bead-list, --robot-pipeline-status, --robot-incident, a --query flag, --project for mail check, --limit for xf, --spawn-dry-run), and examples used the deprecated --md-compact and --ack-timeout; they now use the flags the binary accepts. The README names the WebSocket endpoint (/api/v1/ws) and the per-user checkpoint store (~/.local/share/ntm/checkpoints/<session>/). docs/openapi.json is regenerated from the served router.

Changed

  • Assignment matching (f8924140, 83c19bfb): candidates are filtered by confidence before workload balancing, so an idle low-score agent no longer takes work a qualified agent could accept. Batch assignment first admits the highest-priority feasible set of beads, then maximizes total score with the min-cost flow solver, moving flexible work to free specialists instead of leaving them idle.
  • Builds include ensemble (116098c6, 91bc4f27, 77aa2db2): make build, make build-all and the README's from-source install now build with -tags ensemble_experimental, as release binaries do, so a source-built ntm runs ntm ensemble. The container image builds with the digest-pinned Go 1.26.8 Alpine builder and the same tag, and installs Alpine packages from the branch's current index rather than exact builds that can disappear from it.

Performance

  • Session monitor file-change sampling (GH #337): the monitor ran git status --untracked-files=all and stat'd every record every 15s, so one untracked directory of agent scratch output (84k files in the report) cost an idle session about 13% of a core. Untracked directories are now listed per file only while small (1,000 entries each and 5,000 per sample, with the size probes bounded by the same budget and pathspecs batched 500 at a time); a larger directory is tracked as one entry. On a repository with 50,000 untracked files a sample went from about 215ms to about 7ms. The sampler also runs git with --no-optional-locks so it never takes index.lock from an agent's commit, and stats paths against the working tree's top level, which fixes missed repeat edits and commits recorded as deletions when the manifest root is a subdirectory (3ba3702e, 36232c17).
  • Robot commands skip the projection refresh they do not read (e39494fb, 624077a7): every --robot-* call first ran the full projection refresh (tmux captures of every pane, bv triage, br and git queries). Metadata surfaces (capabilities, help, schema, palette, recipes and similar), the tool bridges (cass, ms, xf, jfp, slb, rch, quota, account, ru-sync, dcg, mail, rano and others) and send, interrupt, ack, tail and is-working now skip it. Measured on one host: --robot-palette 4.65-5.25s to 0.15-0.20s, --robot-mail 4.03-4.72s to 0.30-0.36s, --robot-tail 2.44-2.61s to 0.15-0.20s, --robot-is-working 2.31-2.68s to 0.12-0.21s.
  • The projection refresh collects work and tmux state concurrently (cd55bdc0): --robot-status went from 3.57-4.39s to 2.36-2.71s and --robot-snapshot from 4.60-5.37s to 3.64-3.97s on the same host.
  • The session timeline recorder samples agent states every 10s instead of 5s, matching the health-check cadence (00322379).

Dependencies

  • Go modules: charmbracelet/x/ansi 0.11.7 -> 0.11.8, go-chi/chi/v5 5.3.1 -> 5.3.2, mattn/go-runewidth 0.0.27 -> 0.0.30, shirou/gopsutil/v4 4.26.7 -> 4.26.9, golang.org/x/sys 0.47 -> 0.48, golang.org/x/term 0.45 -> 0.46, golang.org/x/tools 0.49 -> 0.51 and modernc.org/sqlite 1.56.0 -> 1.60.1 (c8209333, c399825b). chromedp stays at 0.16.0: 0.17 through 0.20 redesign its API and the e2e web UI test has not been ported (8de65bda).

Don't miss a new ntm release

NewReleases is sending notifications on new releases.