The repeat bill: why caching exists, explained with arithmetic
Paste the same context into a chat five times in one day and you have paid for it five times. Caching is the fix, and the arithmetic behind it is simple enough for a child to do on paper.
A student pasted his entire About Me context - his class, his subject, his rules for how he liked answers explained - into five separate chats over the course of one afternoon, once per question. We asked him to count how many times he'd typed essentially the same paragraph. He hadn't noticed he was doing it until we asked him to add it up.
Five times the same context means five times the cost of that context, even though nothing about it had changed since the first time. That repeat bill is exactly what caching exists to prevent.
What is prompt caching?
Caching means that when the same block of text - like a standing context or a long document - gets reused across multiple questions in a short span, the system can recognise it has already "seen" that exact text and charge much less for it the second, third and fourth time, instead of paying full price for it again on every single question.
Paste the same 200-word context five times across five questions. Full cost, five separate times, for identical text.
The same context is recognised as already-seen after the first time. Later questions pay a small fraction of the cost for that repeated part.
The arithmetic, done on paper
If a 200-word context costs, say, four coins to read once, and it gets pasted fresh into five different questions without caching, that's twenty coins spent on identical text. With caching recognising the repeat, the real cost across all five looks much closer to four coins plus a small top-up each time - the difference between paying full price five times and paying full price once.
Why this closes the whole pathway
Caching rewards exactly the habits this entire ten-level pathway has been building toward - writing context down once instead of retyping it, keeping a conversation going instead of starting fresh unnecessarily, being deliberate about what actually needs to be said again. The child who understands the repeat bill has, without necessarily noticing, also understood most of what levels one through nine were teaching along the way.
One honest limit: caching only helps with text that is genuinely identical or very close to it - rephrasing the same context slightly each time defeats it entirely, because the system no longer recognises the text as already-seen. The habit that actually saves money is the same one that makes for better prompts and cleaner context in the first place: say it once, precisely, and reuse it rather than restate it.
Questions we get asked
What is prompt caching?
It is a way of reducing cost when the same block of text - like a standing context or a long document - is reused across multiple questions in a short span. Instead of charging full price for that text every single time it appears, the system recognises it has already been processed once and charges much less for the repeats.
Why does caching exist for AI prompts?
Because pasting the same context or document into multiple separate questions would otherwise mean paying full price for identical text over and over. Caching exists specifically to avoid that waste, charging close to full price only the first time and a small fraction for the recognised repeats afterward.
How can a child understand AI caching using simple arithmetic?
Count how many times the same context or document got pasted into separate chats in a day, multiply that by the per-word cost from a word-coins exercise, and the total shows how much would be wasted paying full price every time versus paying it once and reusing the recognised text - a concrete number a child can calculate on paper.
