Notes and explanations about software engineering.
7 posts
Prompt caching stores the model’s computed state, not your text — so it saves money and latency but never a single token of window space. What a breakpoint really marks, why cache entries nest instead of slicing, and the byte-level changes that silently drop your hit rate to zero.
September 10, 2026LLMA context window is not just the text you send. Tool schemas, images, thinking tokens and cached prefixes all sit in it — and the answer has to fit in whatever is left. What counts, why there are two limits and not one, what max_tokens really does, how a cut-off answer comes back as a normal 200, and how to plan the space so it doesn’t.
September 2, 2026RAGWhat chunking actually is, what makes a chunk a good one, and how a text splitter decides where to cut — including why the overlap you configure is usually not the overlap you get.
August 19, 2026RAGA beginner-friendly introduction to Retrieval-Augmented Generation: why an LLM confidently makes things up about your data, why pasting all your documents into the prompt does not fix it, and how the RAG pipeline works, one step at a time.
August 18, 2026DatabaseMVCC leaves a dead row version behind on every update. This is the cleanup story — why bloat slows queries down instead of just filling disk, what VACUUM and VACUUM FULL really do, how autovacuum decides when to run, why one forgotten transaction can block cleanup across the whole database, and how freezing stops 32-bit transaction IDs from silently eating your data.
August 15, 2026DatabaseA ground-up tour of MVCC in PostgreSQL — row versions, the hidden columns and how a tuple is really laid out on disk, snapshots, and the visibility rule — and how they let thousands of transactions read and write at once without waiting on each other.
August 5, 2026DatabaseA deep dive into pages, slotted pages, heap files, and full table scans — the physical layer underneath every SQL query we write.