bettercallcaleb writes about local AI inference, open-source runtimes, quantization, model behavior, and practical systems experimentation. His work emphasizes implementation details, reproducible methodology, and the gap between headline claims and what software actually does on real hardware.
A source-and-runtime audit of BAAI AREX-2: model size, Qwen3.5 multimodal architecture, 262K context, MTP metadata, current llama.cpp support, and what must be verified before local quantization.
A controlled RTX 4090 + RTX 3060 experiment traced a severe Strata Q2_0 multi-GPU prefill slowdown to the Linux expert-copy path. Allowing full CUDA host registration raised 32K prefill from 665 to 2,110 tok/s.
If you searched whether llama.cpp supports NVFP4, the answer depends on the checkpoint, converter, model architecture, backend, and quant-scale path. Use this checklist to identify the real blocker.
Source analysis of a Sep 28 llama.cpp Vulkan bug where flash attention or large softmax can overwrite cached scratch data, causing a later matmul to reuse corrupted input.
Fix llama.cpp KV-cache allocation failures by reducing context or concurrency, checking cache types and device placement, and reading the backend allocator error.
Diagnose llama.cpp's unknown model architecture 'glm5next' error, verify the runtime you are actually using, and check whether upstream support has landed.
A Sep 25 llama.cpp report finds MTP throughput on an RX 7800 XT dropping from 55.79 to 25.80 t/s across a 146-commit range; Linux RADV does not reproduce it.
A real AI-agent workflow audit shows how task routing, compact state checks, shared validation scripts, and smaller handoffs can reduce context and tool-call overhead.
A fresh vLLM fix traces a DeepSeek-V4 CUDA-graph startup assertion to one stale max_model_len copy left behind after KV-cache auto-fitting reduces context.
A Sep 21 vLLM bug report shows speculative decoding can overrun a shorter draft model's RoPE cache, causing CUDA illegal-memory-access errors and full engine restarts.
A Sep 20 llama.cpp report bisects an MTP speculative-decoding regression to CUDA MoE weighted-reduction fusion: draft acceptance drops from 0.82 to 0.48 and MTP becomes slower than no draft.
A Sep 19 llama.cpp fix traces OpenVINO large-context failures to the backend reporting SIZE_MAX instead of the GPU's single-allocation limit, causing GGML to request one oversized KV-cache buffer.
A Sep 18 vLLM fix traces a silent Mistral-Large-3 accuracy regression to config remapping: YaRN frequency interpolation was preserved, but its no-mscale flag was not reaching DeepSeek-style attention.
A source-backed llama.cpp troubleshooting map for unknown architectures, wrong tensor shapes, corrupted GGUFs, CUDA OOM, multimodal mistakes, and runtime regressions.
PR #28898 proposes moving FP8 and NVFP4 quantization scales into GGML so more scaled safetensors checkpoints can convert directly instead of relying on model-specific handling.
Unsloth did not begin as a generic AI startup. It began with a narrow technical advantage, reproducible notebooks, and direct distribution to the local-LLM community.
A fresh llama.cpp Metal PR isolates a large-tensor addressing failure at the 2 GiB slice offset and replaces Tensor API slicing with explicit 64-bit address calculation.
A paired benchmark of refusal behavior, calibrated reasoning, open-ended response drift, and epistemic integrity in a 27B Base vs Heretic model comparison.
A new llama.cpp CUDA PR targets duplicate activation quantization and intermediate FP32 traffic in dense Q4_K MLP prefill, with contributor-reported gains from 4.1% to 14.7%.
A new llama.cpp prototype wires Qwen3.8-Flash-Next's built-in NextN/MTP head into speculative decoding. Here is what it changes, the reported speedup, and why mainline still cannot use it.
A RAMGPT ExactBench research note on Qwen3.8-27B Q4: 94 short deterministic reasoning items, zero wrong answers, and several benchmark-design failures uncovered along the way.
Independent RTX 4090 testing shows Qwen3.8 Flash-Next decode falling from 21.00 to 6.93 tok/s under concurrent STREAM traffic, while matched CPU-only contention reduced throughput by only 5.4%.
Independent RTX 4090 CUDA testing of adaptive DFlash2 in llama.cpp: fixed n=7 reached 80.11 tok/s on structured JSON, versus 70.61 tok/s for adaptive 3–7.
A detailed look at llama.cpp v0.2.0: CUDA decode tuning, DSpark for LFM2, KV-cache work, lower-memory quantization, server hardening, and the new stable release track.
What I learned hardening an open-source coding agent for enterprise use, including controls for MCP, shell execution, secrets, downloads, and runtime trust.
A practical first-person workflow for using ChatGPT to understand technical subjects, test your reasoning, create practice questions, and verify important answers.