github Niko1221/Strata v0.1.40.1
Strata v0.1.40.1

3 hours ago

Hotfix for 0.1.40: a tool call that is only quoted (in the thinking or in a code block) no longer runs, and requests waiting during an engine restart no longer hang. Python server fixes only. The engine is the same as in 0.1.40, so there is nothing new to download. See Updating below: this one time it needs one extra command.

Tool calls written inside the thinking (#804, #1058): 0.1.40 turned a <tool_call> the model wrote inside its thinking into a real call whenever the tool was one of the request's. That rescue is meant for a model that forgets </think> and then writes its call. It also fired on a call that was only an example, even in a reply cut off by max tokens (a quoted rm -rf build became a tool call). A call in the thinking is now a call only when all of this holds:

  • tools are declared and the name is one of them;
  • it starts at the beginning of a line, outside a code fence (``` or ~~~) and outside inline code;
  • its </tool_call> is there, and only whitespace or more calls follow it before the end of the turn or </think>;
  • the turn ended normally. A reply cut by max tokens (or cancelled, or failed) keeps it as thinking text.

Anything else stays in reasoning_content. The same holds when the answer streams. Each rescue is logged, and /metrics counts them in totals.tool_calls_from_reasoning. We ran 37 recorded cases (the 12 real strandings from 47Hunter47's second machine and the 25 cases in PR #525's list) in whole and in chunks, plus fences and tags split across stream chunks: all 16 real calls come through and the 21 quoted ones stay text. Thanks to 47Hunter47 for the 20 test cases and the live runs that showed it.

An engine restart with requests waiting (#1012): after the server restarted a dead engine, requests that were waiting for the engine's control lines could hang for good (we saw five stay at "reading the prompt" for 90 minutes) or fail with list.remove(x): x not in list. The restart now keeps the same lock and wait list, and wakes the waiting requests. A waiting request goes on with the new engine, or ends at once with a 503 saying it was not sent, so a client can retry. A long prompt that was being read when the engine died also ends at once now instead of after 300 s. Thanks to Jackwwg83 for the report and its cause.

Smaller: when the engine exits after an ERR line that no request read (a failed batch window), the server window now says what it was (PR #998, ischencheng). Setup reads a four-part version such as 0.1.40.1 as 0.1.40.

Quoted calls in the visible answer: a <tool_call> inside a code fence (``` or ~~~) or inside inline code in the answer (after </think>) is now text, never a call, also when the answer streams and the fence is split across chunks. Before, a model that showed the call format in a fenced example got a real tool call. Calls at the top level of the answer work as before, including one that follows a sentence on the same line.

Updating (once by hand): we cleaned up the repository history today, so UPDATE.bat's normal pull refuses it this one time. In your Strata folder run:

git fetch origin
git reset --hard origin/main

Then restart the server. Your models, configs, run scripts and the engine are not touched (git does not track them); only changes you made to Strata's own source files would be lost. After this, UPDATE.bat (Linux: ./update.sh) works as before. A zip download of the code needs nothing special. The ready-made engines are the ones from 0.1.40, attached again below.


Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Don't miss a new Strata release

NewReleases is sending notifications on new releases.