Benchmarks

We Gave Three Qwen3.8-27B Variants a Cybersecurity Exam

A first-pass RAMGPT field test of Base, Heretic, and OrcaRouter Qwen3.8-27B variants across AI-agent security, secure code review, and modern glibc exploitation.

Approximately 9 min read

I wanted to answer a fairly simple question:

If you remove the guardrails from a strong 27B reasoning model, do you actually get a better cybersecurity model, or just one that is more willing to talk?

So I gave three Qwen3.8-27B variants the same small private cybersecurity exam:

This was not intended to be a giant benchmark.

It was a first-pass field test on one local machine, using synthetic lab scenarios, private prompts, and manual technical review.

The goal was not to publish exploit payloads.

The goal was to see whether the models could actually reason.

The setup

The runs were performed locally on an RTX 4090 using llama.cpp.

The main model class was Qwen3.8-27B at roughly Q4 weight precision, with MTP speculative decoding where available.

The private test set covered three broad areas:

  1. AI-agent and tool-use security
  2. application-security code review
  3. modern low-level glibc heap exploitation

I also logged the reasoning stream so I could distinguish real long-form analysis from pathological repetition.

That turned out to matter.

The short version

The three models were surprisingly capable at high-level security reasoning.

They were much less trustworthy when pushed into exact low-level runtime internals.

The strongest areas were:

The weakest area was modern allocator exploitation.

All three could speak the language of heap exploitation.

None of them earned my trust on the exact glibc 2.35 mechanics.

That difference is the main result.

AI-agent security was genuinely strong

The first hard scenario involved a fictional incident-response agent with attacker-controlled case data, URL-fetching tools, redirects, internal metadata, secret expansion, and several tempting but invalid exploit chains.

A weak answer sees:

prompt injection
+
SSRF
+
token
=
full compromise

and stops thinking.

A better answer asks:

The good news is that the models handled this style of problem quite well.

The RVN Heretic build produced a long reasoning trace, naturally converged, and generated a complete answer without collapsing into repetition.

The Base model also completed the same problem cleanly.

Both separated guaranteed primitives from assumptions and rejected multiple attractive-but-unsupported attack chains.

That is real security reasoning.

It is not simply “the uncensored model was willing to answer.”

Uncensoring was not the main difference

One of my initial hypotheses was that the Base model would spend more time refusing or sanitizing the offensive-security portions while the uncensored models would simply answer.

That was not what happened on the clearly isolated lab tasks.

The Base model was perfectly capable of analyzing the fictional attack chains.

That matters because it weakens the simplest marketing story around uncensored security models.

The useful distinction is not:

Base = refuses
Uncensored = hacker

The more interesting distinction is:

Does the model know what is actually true?
Does it keep assumptions separate from facts?
Does it reject fake exploit steps?
Does it stop when the evidence stops?

Removing refusals cannot repair missing technical knowledge.

Secure code review was one of the strongest areas

I also gave the models a synthetic Go application containing a mixture of real bugs and deliberate bait.

The code included:

This is one of my favorite ways to test a security model.

Finding “vulnerabilities” is easy.

Not reporting vulnerabilities that do not exist is harder.

The RVN Heretic model performed extremely well on this task.

It found the actual authorization bug and URL-fetcher problem while rejecting the fake SQL injection, command injection, JWT algorithm-confusion, XSS, and unsupported path-traversal claims.

That is exactly the behavior I want from a code-review assistant.

A model that reports ten “critical” issues because ten code snippets look scary is not useful.

False-positive discipline matters.

Why false positives are a serious benchmark target

Security benchmarks often reward finding the intended flaw.

Real security work is messier.

A useful reviewer also needs to say:

No, this is parameterized SQL.

No, exec.Command is not invoking a shell here.

No, html/template is escaping this value.

No, the attacker cannot create the symlink your exploit requires.

No, you have not established an exfiltration channel.

Those negative judgments require actual understanding.

In this first pass, the 27B model class looked much stronger at this than I expected.

Then glibc broke everybody

The low-level test was deliberately different.

It used a synthetic glibc 2.35 heap scenario with modern mitigations:

This is exactly the kind of problem where language models can become dangerous in a different way.

They remember fragments of old exploitation tutorials.

Then they assemble those fragments into a chain that sounds plausible.

The models knew the vocabulary:

But when pushed deeply, they started mixing mechanisms from different glibc eras and asserting allocator behavior that did not hold for the stated environment.

The output still looked like exploit research.

It had offsets.

It had pointer arithmetic.

It had heap terminology.

It had confidence.

And parts of it were wrong.

That was the most useful result of the whole experiment.

“Stable reasoning” and “correct reasoning” are different

Another thing became obvious while running these tests:

A model can reason for a very long time without repeating itself and still be wrong.

Conversely, a model can enter a repetition spiral for reasons that are partly runtime-related rather than purely model-related.

I logged long reasoning traces and measured repeated long spans.

In some runs, a model stayed structurally clean for tens of thousands of words.

In others, repetition suddenly accelerated and the model never reached a final answer.

One particularly dramatic case turned out to change when KV-cache precision changed.

That is interesting enough that I am not treating it as a throwaway detail here.

It deserves a separate experiment.

The important point for this article is simpler:

Long reasoning is not evidence of expertise.

A 20,000-word answer can still be based on the wrong allocator model.

The Base model was not obviously worse because of guardrails

The Base model did not fail the high-level security tests because it was “too aligned.”

It could analyze the isolated lab scenarios.

It could rank attack paths.

It could discuss SSRF, prompt injection, tool abuse, and secret disclosure.

That means the uncensored variants need to justify themselves on something more substantial than willingness.

They need to be:

On this tiny test set, the evidence for a huge capability gap was not there.

What the Heretic build did well

The RVN Heretic variant was the most consistently impressive first-pass model for the kinds of security tasks I care about.

Its strongest qualities were not “no refusal.”

They were:

That does not make it an exploit oracle.

The glibc test made that very clear.

But as an AppSec and AI-security research assistant, it looked useful.

OrcaRouter was also capable, but runtime stability became part of the story

The OrcaRouter uncensored build could also reason deeply about the same high-level security scenario.

However, one long run entered an extreme repetition spiral.

Initially that looked like a model-quality failure.

Then I reran the same model with higher-precision KV cache and the behavior changed dramatically.

The same prompt converged normally.

That means I do not count the original collapse as clean evidence that Orca itself is a worse security model.

Instead, it opened a different research question:

How much can KV-cache precision alter long-horizon reasoning stability?

That is the next article.

What I would actually use these models for

Based on this first pass, I would be comfortable using a Qwen3.8-27B-class local model as a second pair of eyes for:

I would not treat the model as an authoritative source for:

In those domains, the model should be treated as a hypothesis generator.

Not an oracle.

Why the glibc trap matters

This was not just a difficult trivia question.

It exposed a very specific failure mode.

The model could produce a technically sophisticated-looking argument while using the wrong internal mechanism.

That is worse than an obvious refusal.

An obvious refusal tells you that you do not have an answer.

A polished false exploit chain can waste hours.

For advanced security work, confidence calibration is therefore part of capability.

A good model should know when it has moved from:

established mechanism

to:

memory fragment
+
guess
+
plausible exploit folklore

None of the models were reliable enough at that boundary in this test.

What this experiment does not prove

This was a small private test set.

It does not establish a global ranking of the three models.

It does not prove that one uncensoring technique is superior.

It does not measure real-world intrusion capability.

It does not test malware generation, persistence, credential theft, or live-target exploitation.

And it does not prove that every low-level systems question will fail.

The narrower result is more useful:

These 27B local models already show strong high-level security reasoning, especially in AppSec and AI-agent scenarios, but exact low-level systems knowledge remains brittle enough that expert verification is mandatory.

The part I did not expect

I started this test expecting the interesting story to be uncensoring.

It was not.

The interesting stories were:

  1. Base models can already do substantial isolated cybersecurity reasoning.
  2. Uncensoring does not create missing expertise.
  3. False-positive rejection may be a better capability signal than raw exploit generation.
  4. Modern allocator internals are still a strong expert trap.
  5. Runtime choices may materially affect long-horizon reasoning stability.

That fifth point deserves its own benchmark.

Bottom line

The Heretic model did not impress me because it was uncensored.

It impressed me because, on the tasks it actually understood, it could reason deeply and reject bad exploit chains.

But the glibc test was a useful reminder:

A model can reason for 20,000 words and still be wrong.

Long reasoning is not expertise.

Uncensoring is not expertise.

And confidence is definitely not expertise.

The interesting question is not whether a model will answer the security question.

It is whether it knows when its answer stops being real.

— Lisa S., RAMGPT

Sources and further reading

Continue reading