An October 3 llama.cpp change rewrites Qwen4exp indexer scoring around the lightning indexer, reducing live score-buffer memory and extending backend handling across CUDA, Metal, and Vulkan.
A source-and-runtime audit of BAAI AREX-2: model size, Qwen3.5 multimodal architecture, 262K context, MTP metadata, current llama.cpp support, and what must be verified before local quantization.
Source analysis of a Sep 28 llama.cpp Vulkan bug where flash attention or large softmax can overwrite cached scratch data, causing a later matmul to reuse corrupted input.
PR #28898 proposes moving FP8 and NVFP4 quantization scales into GGML so more scaled safetensors checkpoints can convert directly instead of relying on model-specific handling.
Unsloth did not begin as a generic AI startup. It began with a narrow technical advantage, reproducible notebooks, and direct distribution to the local-LLM community.
A source-level look at llama.cpp PR #28809, which makes server router discovery prefer a GGUF whose filename matches its model directory instead of relying on directory iteration order.
A new llama.cpp prototype wires Qwen3.8-Flash-Next's built-in NextN/MTP head into speculative decoding. Here is what it changes, the reported speedup, and why mainline still cannot use it.
llama.cpp PR #25444 adds Nemotron-3-Puzzle-75B-A9B GGUF support, including heterogeneous per-layer MoE metadata, converter fixes, and model loading — but not the model's MTP draft head yet.
A RAMGPT ExactBench research note on Qwen3.8-27B Q4: 94 short deterministic reasoning items, zero wrong answers, and several benchmark-design failures uncovered along the way.
A source-driven history of the increasingly uncomfortable relationship between Ollama and llama.cpp: upstream engineering, attribution disputes, forks, compatibility failures, community backlash, and the cost of hiding the engine.
A llama.cpp Qwen4 experimental report exposes how tiny inference-time reads can defeat fast storage, NFS caching, and otherwise capable local-AI infrastructure.
A production-inference analysis of PR #28320, which temporarily reclaims llama.cpp compute scratch so a vision projector can run on GPU without reducing context.
A source-path experiment on the first LongCat-Flash-Lite-Sparse llama.cpp fork: LSA token selection is real, but selected tokens are converted into a mask over full K/V rather than compact sparse attention execution.
Independent RTX 4090 testing shows Qwen3.8 Flash-Next decode falling from 21.00 to 6.93 tok/s under concurrent STREAM traffic, while matched CPU-only contention reduced throughput by only 5.4%.
A code-level analysis of llama.cpp's DFlash2 merge: dynamic local convolution, lattice path selection, runtime integration, benchmarks, and review tradeoffs.
Independent RTX 4090 CUDA testing of adaptive DFlash2 in llama.cpp: fixed n=7 reached 80.11 tok/s on structured JSON, versus 70.61 tok/s for adaptive 3–7.
A beginner-friendly guide to local voice cloning with Audio8 TTS Preview 0.1B, from reference audio and transcripts to generation and FFmpeg speed control.
A detailed look at llama.cpp v0.2.0: CUDA decode tuning, DSpark for LFM2, KV-cache work, lower-memory quantization, server hardening, and the new stable release track.
An independent llama.cpp PCTree test on RTX 4090 with Qwen3-8B Q8: DSpark doubled decode speed, while wider trees raised acceptance but still lost throughput.
Why llama.cpp is evolving from a GGUF model runner into a broader local inference stack spanning servers, scheduling, KV caches, and speculative decoding.
A lively look at how ik_llama.cpp split from llama.cpp, why local-AI users care, what each project does best, and why the 'fork drama' is more useful than it sounds.