HereticBench: What Actually Changes When You Remove a Model's Refusal Behavior?
A paired benchmark of refusal behavior, calibrated reasoning, open-ended response drift, and epistemic integrity in a 27B Base vs Heretic model comparison.
The RAMGPT library / 4 articles
Transparent, reproducible comparisons with stated methodology.
A paired benchmark of refusal behavior, calibrated reasoning, open-ended response drift, and epistemic integrity in a 27B Base vs Heretic model comparison.
Independent RTX 4090 testing finds UD-IQ3_XXS only ~2% slower than UD-Q3_K_XL at matched context, while reducing the 200K GPU memory budget from 16 to 15 GiB cuts prompt processing by 19%.
An RTX 4090 reproduction of Qwen3.8-27B at ~2.5 bpw finds 1.93% lower WikiText-2 perplexity for GSQ-RCO than Unsloth UD-IQ2_S, while tensor dumps reveal a far more heterogeneous precision allocation and slightly lower throughput.
An RTX 4090 reproduction of Qwen3.8-Flash-Next PLE paging shows a 38.4GB table reaching only 221.7MiB resident after an 8K high-diversity prompt sweep.