Gemma 4 vs Qwen3.8: Parallel Tool Calling Changes Agent Scaling
A local RTX 4090 agent benchmark shows Gemma 4 winning small tasks on speed while Qwen3.8 keeps agent turns flat by batching parallel tool calls.
Approximately 11 min read
Inside this experiment
The result, up front.
- What we found
- Gemma was faster on small tasks. At 16 files, the measured times nearly met while Qwen used five model turns instead of twenty.
- Test conditions
- One RTX 4090; tested Q4 artifacts; llama.cpp ce8caa6; 8K context; temperature 0; 2,048 output tokens per turn. Servers were already loaded.
- What this cannot prove
- One recorded run per model and file count, not a general model ranking or a universal crossover point. Both 24-file attempts failed strict task validation.
24 files are not plotted as successful tasks: Gemma stopped before writing; Qwen completed the workflow but omitted filename suffixes and failed strict formatting.
- 2 filesGemma 45.31 sQwen3.811.05 s
- 4 filesGemma 47.42 sQwen3.814.36 s
- 8 filesGemma 49.82 sQwen3.820.06 s
- 12 filesGemma 413.49 sQwen3.822.14 s
- 16 filesGemma 427.81 sQwen3.826.99 s
Successful tasks only. Single-run wall times, rounded as in the original table; these are not averages or cold-start timings. Original table & context ↓
Drawn directly from this article’s published tables at build time. No new benchmark run or synthetic data.
A local agent can look fast for two very different reasons.
One model can generate each step quickly.
Another can take fewer steps.
I ran into that difference while comparing Gemma 4 26B-A4B-it Q4 with a local Qwen3.8-27B Q4 artifact on the same RTX 4090.
On small file-agent tasks, Gemma 4 was clearly faster in wall-clock time. At two files, the task finished in 5.3 seconds versus 11.0 seconds for Qwen3.8.
Then I increased the fan-out.
At 16 files, the gap disappeared: 27.8 seconds for Gemma 4 versus 27.0 seconds for Qwen3.8.
The reason was visible in the tool trace.
Gemma 4 issued file reads serially, one tool call at a time.
Qwen3.8 batched every independent read into one parallel tool-call turn.
The two models were not just solving the same task at different speeds. They were using different orchestration strategies.
That difference matters for local agents because tool scheduling can change how wall time scales even when raw token generation favors the other model.
The benchmark task
I wanted a task simple enough that the tool trace would be easy to inspect.
Each run created a directory of tiny text files:
f01.txt
f02.txt
f03.txt
...
Each file contained one score:
score=95
The model received this instruction:
Use tools to inspect every .txt score file in the directory.
Find the top three scores.
Write summary.txt exactly as three lines in descending order
using filename=score.
Then read summary.txt back to verify it and report completion.
The available tools were deliberately small:
list_dir
read_text_file
write_text_file
There was no search tool in the scaling test. If the model wanted to know every score, it had to read every input file.
I ran the same workflow with 2, 4, 8, 12, 16, and 24 input files.
The harness recorded:
- task success;
- wall-clock time;
- model turns;
- total tool calls;
- largest batch of parallel tool calls;
- duplicate calls;
- finish reason;
- the final written summary.
The timer started after the model server was already loaded. These are agent-task times, not cold-start times.
Test setup
The reported runs used:
| Setting | Value |
|---|---|
| GPU | NVIDIA RTX 4090 24 GB |
| Runtime | llama.cpp |
| Context | 8K |
| Quantization | Q4 |
| Temperature | 0 |
| Max output tokens per model turn | 2048 |
| GPU offload | full model offload where supported |
| Tool executor | local filesystem functions |
| File contents | tiny deterministic plaintext score files |
The Gemma artifact was:
gemma-4-26B-A4B-it.Q4_K_M.gguf
The Qwen artifact was:
Qwen3.8-27B-UD-Q4_K_M.gguf
For Qwen3.8, I am reporting the local artifact name exactly as tested. The GGUF metadata identified the model as Qwen3.8-27B, but did not include an upstream source URL, so I am not attributing that file to a specific public distribution.
The llama.cpp build was from commit ce8caa6, which included recent Gemma 4 CUDA tuning.
While loaded, the server used roughly 17.7 GB VRAM for Gemma 4 and 16.4 GB for Qwen3.8 in this configuration.
A Gemma 4 template issue had to be fixed first
The first Gemma smoke test was invalid for two reasons.
First, I initially grabbed the base checkpoint instead of the instruct checkpoint.
Second, the Q4 instruct GGUF contained a chat template that llama.cpp explicitly warned was outdated.
That matters a lot in a tool-calling benchmark.
With the wrong setup, Gemma produced behavior that looked catastrophic: it ignored a file-reading tool, guessed an answer, and emitted chat-template tokens repeatedly.
That would have made a dramatic benchmark result.
It also would have been wrong.
For the reported results I used:
- Gemma 4 26B-A4B-it, not the base model;
- Google’s current official
chat_template.jinja; - llama.cpp Jinja chat templating.
After that correction, Gemma passed the small functional tool tests cleanly.
This is an important benchmark rule by itself:
If tool calling looks completely broken, verify the checkpoint and chat template before blaming the model.
Small agent tasks: both models worked
Before the scaling run I used four small agent tasks:
- read one file and report a value;
- find a marker under a directory;
- read several score files and identify the largest value;
- edit a configuration file and read it back.
With the corrected Gemma setup, both models passed all four tasks.
The interesting difference was already visible in the multi-file test.
Gemma read the files one at a time.
Qwen3.8 issued several independent reads in the same assistant turn.
For tiny tasks, Gemma’s much faster per-turn generation still dominated the result.
Across those four tasks, the measured total wall time was approximately:
| Model | Passed | Total wall time |
|---|---|---|
| Gemma 4 26B-A4B-it Q4 | 4/4 | 8.1 s |
| Qwen3.8-27B Q4 | 4/4 | 22.4 s |
That made Gemma look like the obvious agent choice.
The scaling test changed the picture.
Fan-out scaling results
Here is the final controlled series with a 2048-token output budget:
| Input files | Gemma 4 time | Gemma turns | Gemma max parallel | Qwen3.8 time | Qwen turns | Qwen max parallel |
|---|---|---|---|---|---|---|
| 2 | 5.31 s | 6 | 1 | 11.05 s | 5 | 2 |
| 4 | 7.42 s | 8 | 1 | 14.36 s | 5 | 4 |
| 8 | 9.82 s | 12 | 1 | 20.06 s | 5 | 8 |
| 12 | 13.49 s | 16 | 1 | 22.14 s | 5 | 12 |
| 16 | 27.81 s | 20 | 1 | 26.99 s | 5 | 16 |
| 24 | failed | 4 | 1 | 35.38 s | 5 | 24 |
The pattern is unusually clean.
Gemma 4’s tool plan effectively becomes:
list
read 1
read 2
read 3
...
read N
write
verify
final
Qwen3.8’s plan becomes:
list
[read 1, read 2, read 3, ... read N]
write
verify
final
For Qwen3.8, the number of agent turns stayed at five all the way from 2 files to 24 files.
For Gemma 4, the turn count grew with the number of files because each read consumed another model round trip.
The crossover happened at 16 files
At low fan-out, Gemma 4’s faster generation more than compensated for the extra turns.
At two files:
Gemma 4 5.31 s
Qwen3.8 11.05 s
At eight files:
Gemma 4 9.82 s
Qwen3.8 20.06 s
Gemma was still substantially faster.
By 16 files:
Gemma 4 27.81 s
Qwen3.8 26.99 s
Qwen3.8 had caught up even though its generation rate was much lower in the earlier smoke tests.
The agent harness had changed the scaling law.
That is the main result from this experiment.
A model that is slower per token can still become competitive on a wide tool task if it collapses independent operations into one model turn.
What happened at 24 files
The 24-file result needs two separate explanations.
Gemma 4 stopped before finishing
With 24 files, Gemma listed the directory and began reading files serially.
In the final 2048-token run it read two files, then the next model generation consumed the output budget and returned:
finish_reason=length
It never reached the write step.
A separate 20-file boundary probe showed similar instability: Gemma reread earlier files and again ended on the output limit without completing the workflow.
I would not interpret that as a generic 20-file hard limit for Gemma 4. The prompt, template, tool schema, output budget, llama.cpp parser, and local artifact all matter.
What the trace does show is narrower:
In this harness, Gemma’s serial tool policy created a much longer agent trajectory, and the wider task became unstable before completion.
Qwen3.8 completed the workflow but missed strict formatting
Qwen3.8 issued all 24 file reads in one batch.
It then identified the correct top three values, wrote the summary, read the summary back, and returned a normal final answer.
The strict validator still marked the run as failed.
Why?
The requested format was:
f15.txt=99
f04.txt=95
f20.txt=93
Qwen wrote:
f15=99
f04=95
f20=93
The scores and ranking were correct, but it dropped the .txt suffix.
So I count the 24-file run as:
- agent workflow completed;
- tool orchestration succeeded;
- strict task-format validation failed.
That distinction is more useful than a single pass/fail number.
The 768-token failure was a benchmark artifact
I nearly published a misleading Qwen failure at 24 files.
The first harness used:
max_tokens=768
At 24 files, Qwen attempted to emit a giant parallel tool-call response.
The JSON was truncated before the final calls were complete. The parser eventually received a tool argument consisting of only:
{
That looked like malformed tool calling.
It was not evidence that Qwen suddenly forgot how to produce JSON.
I raised the per-turn output budget to:
max_tokens=2048
The next run produced all 24 tool calls correctly.
This is a surprisingly important benchmark confound.
Parallel tool calling can require a lot of output tokens because every function call carries:
- function name;
- argument JSON;
- path strings;
- call identifiers;
- wrapper syntax.
A model that parallelizes aggressively can therefore need more output budget per turn even while using fewer total turns.
If the benchmark caps output too tightly, the model with the better fan-out policy can look worse.
Parallel tool calls are not automatically faster
It would be easy to overgeneralize this result.
In my harness, file reads were tiny local operations. There was almost no reason not to issue them together.
Real tools can have constraints:
- API rate limits;
- write dependencies;
- locks;
- ordering requirements;
- transaction semantics;
- expensive shared resources;
- external side effects.
If tool B depends on tool A’s result, parallelism is wrong.
If 24 simultaneous API calls trigger throttling, parallelism can be slower.
If several tools write the same resource, parallel execution may be unsafe.
The useful capability is not “always parallelize.”
It is:
Recognize independent operations and batch them when the tool contract allows it.
That is agent orchestration, not raw language-model throughput.
Why tool scheduling belongs in agent benchmarks
Most local-model benchmarks focus on quantities such as:
- prompt processing speed;
- generation speed;
- VRAM;
- perplexity;
- task accuracy;
- long-context throughput.
Those are important.
But an agent introduces another layer:
model
-> tool decision
-> tool calls
-> tool results
-> model
-> more tools
-> final result
Two models with similar task accuracy can produce very different wall time because one requires many more model-tool round trips.
This connects directly to a broader issue I measured in How to Reduce AI Agent Context and Tool-Call Overhead: agent efficiency is partly a harness and orchestration problem, not only a model-speed problem.
For local agents, I would now track at least:
task success
strict output correctness
wall time
model turns
tool calls
maximum parallel tool batch
duplicate calls
output tokens
finish reason
A single tokens-per-second number cannot describe that behavior.
What this benchmark does not prove
This was one local harness on one machine.
It does not prove:
- Gemma 4 is generally worse at agents;
- Qwen3.8 is generally better at tool use;
- parallel calling always improves latency;
- 16 files is a universal crossover point;
- 24 files is a universal Gemma failure point;
- these Q4 results predict BF16 behavior;
- another runtime will parse or schedule the same outputs identically.
There are several confounders.
Quantization
Both models were Q4 artifacts. Quantization can change output quality and tool-call reliability.
Different architectures and generation rates
Gemma 4 was much faster per generated token in the simple smoke test. The experiment deliberately did not normalize for token throughput because wall-clock agent performance was the quantity I wanted to observe.
llama.cpp caching
The server stayed loaded across each scaling series, and llama.cpp could reuse prefix state between related requests. These are not isolated cold-start runs.
Tool calls are local and cheap
The file tools completed quickly. A slower network tool could amplify the benefit of concurrency, or rate limits could reverse it.
One prompt style
The tool descriptions and instruction wording can change model behavior. A prompt explicitly telling Gemma to batch independent calls might produce a different trace.
That would be a useful follow-up experiment.
The practical takeaway
The most useful result was not that one model won.
It was that the benchmark exposed two different agent strategies.
Gemma 4 behaved like a very fast serial operator.
Qwen3.8 behaved like a slower operator that aggressively fanned out independent work.
At low fan-out, Gemma’s speed dominated.
As the number of independent tools increased, Qwen’s fixed-turn strategy caught up.
That suggests a simple rule for local agent evaluation:
Measure the trajectory, not just the answer.
Record how many model turns were required.
Record whether independent tools were batched.
Record whether the model reread the same resource.
Record output-budget truncation separately from malformed-tool behavior.
And when a tool benchmark fails dramatically, check the checkpoint and chat template before turning the failure into a model conclusion.
The agent loop can matter as much as the model inside it.