AI Foundations #28 explains static batching, continuous batching, request scheduling, chunked prefill, head-of-line blocking, fairness, and why server throughput and user latency can pull in opposite directions.
AI Foundations #27 explains time to first token, time per output token, prefill, decode, batching, throughput, context length, and why one tokens-per-second number can hide the real user experience.
AI Foundations #26 explains production monitoring, data and concept drift, feedback loops, human review, and why a model that passed evaluation can still fail after deployment.
AI Foundations #25 explains alignment, preference training, guardrails, instruction hierarchy, runtime policy, and why capability and safe system behavior are different goals.
AI Foundations #24 explains tool use, function calling, agents, action loops, observations, permissions, failure modes, and why an LLM does not execute a tool by itself.
AI Foundations #23 explains retrieval-augmented generation, embeddings, chunking, retrieval, context injection, citations, failure modes, and why RAG is different from training.
AI Foundations #22 explains test sets, metrics, benchmark design, sampling variance, contamination, and why a fair model comparison must control more than one score.
AI Foundations #21 explains reasoning tokens, inference-time compute, reward signals, RLHF, PPO, and how reinforcement-learning post-training changes model behavior.
AI Foundations #20 builds a practical mental model of multimodal AI: how text, image, audio, and video representations are converted, aligned, and fed into transformer-style systems.
AI Foundations #19 explains Mixture of Experts models: routers, top-k expert selection, active versus total parameters, load balancing, and why MoE changes inference economics.
AI Foundations #18 builds quantization from number representation and rounding error, then connects 4-bit weights to model size, speed, calibration, and quality tradeoffs.
AI Foundations #17 explains context windows and KV cache: what the model can see, what keys and values are cached, why decoding gets faster, and why long chats consume more memory.
AI Foundations #16 explains how logits become generated text through greedy decoding, temperature, top-k, and top-p sampling, with simple probability intuition and worked examples.
AI Foundations #15 explains inference as the trained model's forward-only phase: tokenize a prompt, compute logits, choose a token, append it, and repeat without backpropagation or weight updates.
AI Foundations #14 explains how fine-tuning continues gradient-based training from pretrained weights, using narrower data to change model behavior without relearning language from scratch.
A practical research playbook for reconstructing CVE-centric remediation event logs from change, incident, request, and vulnerability records before forming hypotheses.
AI Foundations #12 assembles attention, residual connections, normalization, and MLP layers into the repeating Transformer block behind modern language models.
AI Foundations #11 follows tokenization and embeddings into attention: queries, keys, values, similarity scores, softmax weights, and context-dependent representations.
AI Foundations #10 connects embeddings to tokenization: how text is split into tokens, mapped to vocabulary IDs, and then looked up as learned vectors.
AI Foundations #8 explains backpropagation as repeated chain-rule bookkeeping that carries loss information backward through a neural network to every trainable parameter.
AI Foundations #7 explains gradient descent as the rule that turns loss into a direction for changing weights, using one small numerical example before backpropagation.
AI Foundations #5 follows one input through a tiny neural network to show what weights, biases, activations, and layers actually do during a forward pass.
A cybersecurity research report on the gap between fast-moving exploitation and slow enterprise remediation, and why compensating controls should be measured by survival rather than deployment alone.
AI Foundations #4 explains tensors through shapes, axes, images, batches, model weights, and memory, without assuming you already know linear algebra or PyTorch.
AI Foundations #3 explains what parameters, weights, and biases actually are, why large models have billions of them, and how precision changes model size.
Learn what LLM context length means, what counts toward the context window, how tokens, chat history, KV cache, VRAM, truncation, and long-context models work.
Learn what an LLM KV cache stores, why it speeds up token generation, how to estimate its VRAM use, and how context length and quantization affect memory.