AREX-2 Local Runtime Readiness: What the 27B Qwen3.5 Artifact Needs
A source-and-runtime audit of BAAI AREX-2: model size, Qwen3.5 multimodal architecture, 262K context, MTP metadata, current llama.cpp support, and what must be verified before local quantization.
Approximately 5 min read
BAAI’s AREX-2 is interesting for local inference because it is a 27B Qwen3.5-family multimodal checkpoint with a large context window, MTP metadata, and an artifact size that immediately forces a deployment decision.
The useful question is not whether AREX-2 is new. It is what has to work before this model becomes a credible local quantization target.
I inspected the published model metadata, current upstream llama.cpp support, and two llama.cpp source trees on my local machine. The result is a readiness map rather than a benchmark.
What the published artifact says
The Hugging Face repository currently exposes about 54.7 GB of model files and identifies the checkpoint as a 27B model.
Its config declares:
architecture: Qwen3_5ForConditionalGeneration
dtype: bfloat16
hidden_size: 5120
num_hidden_layers: 64
max_position_embeddings: 262144
mtp_num_hidden_layers: 1
language_model_only: false
The vision section is also present.
So this is not a plain text-only Qwen checkpoint with a familiar conversion path. It combines Qwen3.5 text architecture, multimodal processing, hybrid linear/full attention, a 262K configured context ceiling, MTP metadata, and a large BF16 source artifact.
Each item adds a separate compatibility surface.
The raw BF16 artifact is not a 4090 model
A 54.7 GB repository is already larger than the 24 GB VRAM on an RTX 4090.
My local machine also has an RTX 3060 12GB, but the combined 36 GB of nominal VRAM is still below the source artifact size before runtime overhead.
So local deployment is mainly a question of whether a suitable quantized artifact can be produced and whether the selected runtime understands every required model component.
A rough 4-bit weight estimate for 27B parameters is:
27 billion parameters
x 4 bits
/ 8
≈ 13.5 GB
That is only a weight-level estimate. It does not include quantization metadata, vision components, runtime workspaces, KV cache, MTP state, allocator fragmentation, or context-dependent buffers.
So “about 13.5 GB at 4-bit” is not the same as “guaranteed to fit in 13.5 GB VRAM.”
Current upstream llama.cpp does recognize Qwen3.5
This is where checking current source matters.
The current upstream llama.cpp Qwen converter registers:
Qwen3_5ForConditionalGeneration
Qwen3_5ForCausalLM
and maps the architecture to the Qwen3.5 GGUF path.
The current llama.cpp README also uses a Qwen3.5 GGUF in its quick-start example.
That means a blanket statement such as “llama.cpp does not support Qwen3.5” would be wrong today.
My local checkouts were behind that current source path
I inspected two local llama.cpp trees on this machine:
/home/chatgpt-agent/workspace/llama.cpp-jev
HEAD bd4f514
/home/chatgpt-agent/workspace/llama.cpp
HEAD ce8caa6
A source search in those checkouts did not find the current Qwen3.5 registration string.
That gives an operational lesson: never diagnose a newly released model against a stale local runtime and then generalize the result to current upstream.
Before attempting AREX-2 conversion or loading, align the test runtime with a known current upstream commit.
The multimodal path needs separate validation
AREX-2 declares language_model_only as false and includes a vision configuration.
A text-only smoke test would therefore prove only part of the deployment path.
A useful validation plan should separate:
text generation
-> chat template
-> structured tool use
-> image input
-> long context
-> MTP/speculative path
A successful text completion should not be reported as complete model support.
MTP metadata is another place to be careful
The config declares one MTP hidden layer.
Metadata alone is not enough.
I hit this exact class of problem recently with MiMo-V2.6-Distill-Qwen-9B: its configuration declared an MTP side model while the expected MTP tensors were absent from the source artifact.
That does not prove AREX-2 has the same issue.
It does mean the correct workflow is:
read config
-> inspect tensor index
-> verify MTP/NextN tensors actually exist
-> convert
-> load without MTP
-> load with MTP
-> compare behavior and memory
Do not infer tensor presence from one configuration field.
vLLM and SGLang are the documented serving paths
The AREX-2 model page currently provides launch examples for both vLLM and SGLang.
That makes them useful control paths.
If a local GGUF or EXL3 conversion fails, a successful supported-runtime load helps separate a source-checkpoint problem from a quantizer/runtime compatibility problem.
Without a control path, a failed conversion can be misdiagnosed as a bad model release.
What I would test next
The next useful experiment is a staged compatibility gate.
Gate 1: artifact integrity
Verify config, processor metadata, tokenizer files, safetensors index, vision files, and actual MTP/NextN tensor presence.
Gate 2: reference runtime
Load the original checkpoint through a currently documented runtime and record the exact runtime version, GPU placement, peak memory, text smoke, tool-call smoke, and vision smoke.
Gate 3: conversion
Try one quantization path at a time.
For GGUF:
current llama.cpp converter
-> F16/BF16 GGUF if practical
-> one moderate quant
-> load test
For EXL3:
current ExLlamaV3
-> architecture recognition
-> calibration/conversion
-> load
-> text/tool/vision validation
A conversion that finishes but cannot preserve required model features is not a complete success.
Gate 4: context scaling
Only after load correctness is established should context length be increased.
The configured 262K ceiling does not mean every consumer-GPU configuration can allocate or sustain 262K.
Measure at practical steps such as 4K, 16K, 32K, and 64K, then go higher only if memory and correctness remain sane.
Why this model is worth watching
AREX-2 is large enough to expose real runtime boundaries but still small enough that a 4-bit-class artifact could plausibly fit on enthusiast hardware.
It also combines several features that tend to fail independently:
new architecture
+ multimodal path
+ very long context
+ MTP
+ quantization
That makes it a better compatibility target than a simple release post.
The publishable result today is therefore not a speed claim.
Current upstream llama.cpp understands the Qwen3.5 architecture, the published AREX-2 artifact is far too large for raw BF16 consumer-GPU loading, and a credible local result requires separate validation of quantization, vision, MTP, and context behavior.
The next article should contain measurements only after those gates are actually run.