Improvements
-
Large results keep what the agent asked for — With the decision model's Large results switch on, a tool result too big for the conversation was summarized by a separate model, which took 7–8 seconds. The decision model now picks the parts the agent's request needs: the matching messages of a Slack search, the failing lines of a log, one section of a long document. The agent gets those parts and the saved file in about half a second. When nothing matches, or nearly everything does, the result is summarized as before. The decision model now receives the whole result in parts, not just its first 3,000 characters.
-
Source tools say why they are called — The agent now states a one-line purpose for every MCP and API source call, which Large results uses to decide what to keep. Before, the purpose almost never reached it.
-
Adaptive thinking reads the conversation — Besides your message, it now sees the end of the previous reply and the names of attached files (never their content), so a short "here" with an attachment or "yes, do it" keeps enough thinking. Pushback on the previous answer keeps your thinking level, and a request to change or send something hard to undo (update records, deploy, push, send a message) runs at least at High.
-
See what each decision feature does — Under the Decision model card's Advanced settings, each feature shows its checks, changes and failures over the last 7 days. A feature that has run 30 checks without changing anything says so.
Bug Fixes
-
Guarded mode asks when the risk check cannot answer — If the decision model timed out or failed, a Guarded call ran unchecked, as in Execute. It now asks you, and the prompt says the call could not be checked.
-
Fewer decision-model timeouts after a pause — Checks failed far more often when the decision model had not been asked for a while (6–12%, against 1%). The first check after 30 seconds without an answer now gets 2.8 seconds instead of 1.5.
-
One set of checks per message — When an expired sign-in made the app re-send your message, the checks at the start of the turn ran twice. They now run once, and the app's own background-task notices are no longer rated for thinking level.
-
Long sidebar names fit on one line — Long page names and other sidebar items wrapped onto several centered lines. They now end in "…" and show the full name on hover.