We Tried to Break a 27B Q4 Model. The Benchmark Broke First.
A RAMGPT ExactBench research note on Qwen3.8-27B Q4: 94 short deterministic reasoning items, zero wrong answers, and several benchmark-design failures uncovered along the way.
Independent experiments, source-level analysis, and practical guides to the AI running on your hardware.
Flagship independent research
A RAMGPT ExactBench research note on Qwen3.8-27B Q4: 94 short deterministic reasoning items, zero wrong answers, and several benchmark-design failures uncovered along the way.
Evidence & experiments
Independent RTX 4090 CUDA testing of adaptive DFlash2 in llama.cpp: fixed n=7 reached 80.11 tok/s on structured JSON, versus 70.61 tok/s for adaptive 3–7.
A paired benchmark of refusal behavior, calibrated reasoning, open-ended response drift, and epistemic integrity in a 27B Base vs Heretic model comparison.
An independent llama.cpp PCTree test on RTX 4090 with Qwen3-8B Q8: DSpark doubled decode speed, while wider trees raised acceptance but still lost throughput.
RAMGPT QuantBench
GGUF artifact rankings against pinned BF16 references, with versioned benchmark contracts.
A cumulative learning series
Build your understanding, one idea at a time. Beginner-friendly explanations from contributors learning AI systematically.
Explore the learning path ↗Continue learning
Continue learning
Continue learning
Continue learning
Continue learning
Continue learning
Continue learning
Continue learning
Continue learning
From understanding to practice
Learn how to estimate LLM VRAM requirements from model size, quantization, KV cache, context length, runtime overhead, and GPU offloading.
A simple step-by-step llama.cpp guide for beginners: compile from source, run a GGUF model, use GPU acceleration, and tune the main inference settings.
Learn why a local LLM runs slowly and how model size, GPU offloading, context length, memory bandwidth, CPU threads, and runtime settings affect speed.
AI Foundations #14 explains how fine-tuning continues gradient-based training from pretrained weights, using narrower data to change model behavior without relearning language from scratch.
A Sep 18 vLLM fix traces a silent Mistral-Large-3 accuracy regression to config remapping: YaRN frequency interpolation was preserved, but its no-mscale flag was not reaching DeepSeek-style attention.
A Sep 15 llama.cpp Vulkan fix targets silent wrong results when mul_mat reads a slice of a larger cache, a failure more dangerous than a clean crash.
A practical research playbook for reconstructing CVE-centric remediation event logs from change, incident, request, and vulnerability records before forming hypotheses.
AI Foundations #13 connects the Transformer architecture to the training objective that turns random parameters into a useful language model.
A source-backed llama.cpp troubleshooting map for unknown architectures, wrong tensor shapes, corrupted GGUFs, CUDA OOM, multimodal mistakes, and runtime regressions.
The people behind the work
Technical Research & Editorial Team
Local AI Systems & Inference Contributor
AI Fundamentals Contributor
Open-Source AI Commentary Contributor
Student Research Contributor — Mathematics & AI
Platform Engineering & MLOps Contributor
Enterprise Security Research Contributor
Infrastructure support
Independent AI testing can require substantial compute. We acknowledge infrastructure providers that help RAMGPT run larger experiments at practical cost.
Infrastructure relationships do not determine RAMGPT's benchmark methodology, measurements, rankings, or editorial conclusions.