Models & Algorithms•SOTAAZ Lab••KR

Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools

Armin Ronacher writes that codemode, where the model writes code that calls tools instead of calling them one at a time, does not yet work with smaller models. We tested it on 30 tasks against a mock issue tracker with Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B in llama.cpp. Codemode got more tasks right at every size (22 vs 18, 28 vs 19, 30 vs 26); only the 9B difference passes our test (p = 0.012), and most of it comes from tool-calling replies cut off at our 2,048-token output limit. On tasks both modes got right, codemode's median tokens per task were 28-45% of tool calling's.

Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools

Does Codemode Work With Small Local Models? Qwen3.5-9B Got 28 of 30 Writing Code, 19 Calling Tools

Most coding agents give the model a set of tools and let it call them one at a time: list the issues, read the result, call the next tool, read that result. Armin Ronacher's "What is Codemode" (6 October) describes another way, a name he credits to Cloudflare. The model gets one tool that runs code, and inside that code the other tools are ordinary functions. It can loop over pages, join results and filter in code, and only what it prints comes back into the conversation.

Near the end he lists what is still unsolved, including "the inability of this pattern to work with smaller models." The post has no measurement behind that sentence. Small models are what people run locally (our guide to running Qwen locally covers which ones fit which GPU), so we measured it.

How we measured

  • Environment: a mock issue tracker built from a fixed seed, with 240 issues and 30 users on five teams. There are five tools: list issues (20 per page, filter by state and label), get one issue, get a user's name and team, list labels, and search.
  • Tasks: 30 questions with exact answers, ten of each kind:

- aggregate: count or sum across several pages, for example "how many open issues labeled bug have more than 5 comments?"

- join: connect issues to users' teams, for example "how many open issues were opened by people on the platform team?"

- single: one or two lookups, for example "what team is the author of issue #42 on?"

  • Two modes:

- Tool calling: the five tools as functions; every result goes into the conversation; up to 40 model calls.

- Codemode: one run_python tool; the five tools are pre-defined Python functions in a separate process with no network or file writes and a 10-second limit; only printed output comes back (first 4,000 characters); up to 10 model calls. Ronacher's implementation runs JavaScript in WASM; we used Python, which small models know better.

- Both system prompts share the same tool descriptions. The codemode prompt also says to do the counting, filtering and joining in code, which is what the method is for; the tool-calling prompt has no matching instruction.

  • Models: Qwen3.5-4B Q4_K_M, Qwen3.5-9B Q4_K_M and Qwen3.8-27B UD-Q4_K_M (Unsloth GGUFs), in llama.cpp 4da6337, temperature 0, reasoning off, 32K context, at most 2,048 output tokens per reply.
  • Grading: the model ends with ANSWER: <value>, compared after normalizing case, spaces and commas. A run with no such line counts as wrong.
  • Rule: for each model, the two modes are compared task by task with McNemar's exact test; a difference counts only if p < 0.05.

We fixed the environment, tasks, prompts, harness and rules before running any model. One change came before the scored runs: in a two-task smoke test, the 9B in codemode guessed that comments was a list of comments when it is a count, and answered 0. Tool calling sees the raw result and cannot make that mistake; codemode only sees the function descriptions. So we added each function's return structure to the shared description. Ronacher makes the same point about MCP: "The outputSchema system in MCP is great for that."

The result

Two bar charts. Left: tasks answered correctly out of 30. Qwen3.5-4B: tool calling 18, codemode 22, p = 0.388. Qwen3.5-9B: 19 and 28, p = 0.012. Qwen3.8-27B: 26 and 30, p = 0.125. Right: median tokens per task on tasks both modes got right. 4B: 5,104 and 1,945. 9B: 4,003 and 1,787. 27B: 6,104 and 1,724.
ModelTool callingCodemodeOnly tool calling rightOnly codemode rightp
Qwen3.5-4B1822480.388
Qwen3.5-9B19281100.012
Qwen3.8-27B2630040.125

On these tasks, the small models did not fail at codemode. Codemode got more tasks right at all three sizes. The 9B difference passes our test; at 4B and 27B we can't separate the two modes with 30 tasks. In no case did codemode come out significantly worse, which is what "doesn't work with smaller models" would predict.

We also asked whether the gap grows as models get smaller. It doesn't in any simple way: codemode was ahead by 4 tasks at 4B, 9 at 9B and 4 at 27B.

By kind of task:

ModelAggregateJoinSingle
4B, tools / code6 / 82 / 710 / 7
9B, tools / code7 / 92 / 1010 / 9
27B, tools / code8 / 108 / 1010 / 10

The gap comes from join tasks. Tool calling got 2 of 10 at both 4B and 9B; codemode got 7 and 10.

Why tool calling failed: it ran out of room counting by hand

Every failed run falls into one of these:

Model, modeWrong answerCut off at the 2,048-token reply limitFinished, but no ANSWER lineHit the call limit
4B, tools4800
4B, code5030
9B, tools3800
9B, code1001
27B, tools4000
27B, code0000

Most tool-calling failures at 4B and 9B were not wrong reasoning. With reasoning off, the model gathered the data with tool calls and then counted in its reply, issue by issue, until the reply hit our 2,048-token limit mid-list. Of the 10 tasks only the 9B's codemode got right, 8 were tool-calling runs cut off this way, so the 9B result depends heavily on that limit. A larger limit might have let some of them finish; we did not test that.

Take "how many open issues were opened by people on the platform team?" There are 150 open issues on 8 pages and 30 users. With tool calling, the 9B fetched all 8 pages and called get_user for each author, then started going through the issues in its reply, "user16 (platform) - count 43", and was cut off before reaching the end. Its 7 model calls used 58,848 tokens in total.

In codemode the same model wrote one script: fetch every page, collect the authors, look up each author's team, count the matching issues, print the number. Its 2 model calls used 2,956 tokens, and it answered 47, which is correct.

That is the real difference between the modes on these tasks: in codemode the counting happens in code, so the model never has to write out 150 items.

Where codemode lost

On single lookups the 4B got 10 of 10 with tool calling and 7 of 10 with codemode. In all three misses the answer was right but the format was not: "The author of issue #42 is on the infra team." with no ANSWER: line. Our grading counts that as wrong, as we fixed in advance. The 9B and 27B had no such misses in codemode.

If we count a finished reply as right when it states the correct answer without the ANSWER: line, a looser rule we added after seeing the results, the 4B codemode score goes from 22 to 25 and nothing else changes. Replies cut off at the limit don't count, since they never reached an answer.

Tokens

On tasks both modes got right, codemode's median tokens per task came to 28% to 45% of tool calling's: 1,945 against 5,104 at 4B, 1,787 against 4,003 at 9B, and 1,724 against 6,104 at 27B. That compares the two medians for each model; the ratio for individual tasks spreads wider. These are the sums over all calls of input plus output tokens, counted the way an API bills them. Tool calling sends every page of results back into the conversation on every later call; codemode keeps the pages inside the code and returns a few lines. How much of that turns into waiting time locally depends on the server: llama.cpp reuses its cache for the unchanged part of the conversation, and we did not measure speed. The context-budget post shows how fast agent conversations fill a small GPU's context.

What this does and doesn't show

  • It shows that on structured lookups where the answer needs many pages and a join, small local models did at least as well writing code as calling tools, and used far fewer tokens.
  • It does not show that codemode beats tool calling in general. Our tasks favor code: the answers are exact, the data is paginated, and joins are mechanical. Much of the 9B gap comes from tool-calling replies that hit our output limit while counting by hand. Ronacher's examples include image generation and MCP servers with inconsistent output, which we did not test.
  • At 4B, three codemode answers were right but in the wrong format. The 9B and 27B had none.
  • Return types are part of the interface. Without the return structure, the 9B guessed a field's type wrong in codemode. If you give a model code access to tools, give it the types too.

Limits

  • One mock environment, 30 tasks, one run each at temperature 0. With 30 tasks, differences of three or four tasks are not distinguishable.
  • A 2,048-token limit per reply. Most tool-calling failures at 4B and 9B were replies cut off at that limit; a larger limit, or reasoning on, might change those results.
  • Python instead of the JavaScript sandbox Ronacher describes, and a codemode prompt that tells the model to do the counting in code.
  • Qwen models only, reasoning off.
  • The return-structure change was made after a smoke test and before the scored runs; the lenient count was added after the results.

The plan, environment, tasks, prompts, harness and every conversation are in the reproduction package below.

Files for this post

Reproduction package

The harness, raw logs and result CSVs behind this post. No account needed; the README has a command that recomputes the tables without a GPU.

codemode-small-models-repro.zip · 582 KB

Download