Prompt engineeringOutput qualityLevel 2

The same AI gave these two answers, four seconds apart

A side-by-side look at what actually changes in an AI's output when a child rewrites the prompt - length, specificity, structure, and whether the answer is usable at all.

·7 min read

The claim that a better prompt gets a better answer is easy to make and easy to nod along to. It is much more convincing when you put the two answers next to each other and read them.

So that is what this is. One topic, one model, one sitting. The only thing that changed between each pair is the question.

Round one: length is a decision, and someone has to make it

Prompt A

tell me about photosynthesis

Prompt B

Explain photosynthesis in exactly four sentences for a Class 6 student.

Prompt A produced nine paragraphs. It opened with a definition, worked through the light-dependent reactions, mentioned chlorophyll three times, and ended with a summary of what it had just said. Everything in it was correct. None of it was read, because a twelve-year-old does not read nine paragraphs about photosynthesis.

Prompt B produced four sentences that a child actually finished. That is not a small difference. An answer nobody reads has a quality of zero, however accurate it is.

When you do not name a length, the model picks one, and its instinct is to be thorough because thorough is safe. Thorough is usually wrong for a child.

Round two: naming the audience changes the vocabulary

Without an audience

"Photosynthesis is the biochemical process by which photoautotrophs convert light energy into chemical energy stored in glucose…"

For a Class 6 reader

"A leaf is a tiny kitchen. It takes in sunlight, water from the roots, and a gas from the air, and cooks them into sugar. The plant eats the sugar. The leftover is the oxygen you just breathed in."

Both are accurate. Only one of them can be repeated at dinner by the person who read it, and being able to repeat something is a fairly good test of whether you understood it.

This is the point in a class where children usually go quiet for a second. They have been told to write for an audience since Class 3 and it has always been an abstraction. Here it changes the words on the screen in front of them, immediately.

Round three: asking for a shape

The third variable is format, and it is the one adults underuse too. Compare:

  • "Compare cricket and football" - gets you five paragraphs of prose, and you have to hold the comparison in your head while you read.
  • "Compare cricket and football as a table: rules, gear needed, players, how long a match lasts, and one thing that makes each one hard." - gets you a table you can scan in ten seconds and argue with.

The second version also does something quieter. By naming the five columns, the child has decided what actually matters about the comparison. The AI did not choose those categories. They did. That is most of the intellectual work, and it happened before the model ran.

The categories in a good prompt are the child's thinking. The text that comes back is just typing.

Round four: examples beat adjectives

Ask for "a funny poem" and you will get what an average of the internet thinks funny means, which is not funny. Paste in two jokes your child actually likes and ask for a third in the same style, and the output moves a long way toward their taste.

Adults call this few-shot prompting and it sounds technical. To a child it is just: show it what you mean instead of describing what you mean. They understand it faster than most adults do, probably because they have less faith in adjectives.

What did not change

Worth being straight about this. Better prompts did not make the AI correct about everything. In one of our sessions a nicely written, well-scoped, four-sentence prompt still produced a confident and completely wrong date. Specificity gets you a better-shaped answer; it does not get you a guaranteed true one.

That is why we teach prompting alongside checking, not before it. The child who writes a beautiful prompt and believes everything that comes back has learned half a skill.

The one-line version

If a child takes only one thing from all of this, make it this: the answer is a reflection of the question. When the answer is vague, the question was vague. That sentence is worth more than any list of prompt templates, and it survives every model release.

Questions we get asked

Does a longer prompt always give a better answer?

No. What helps is specificity, not length. A long prompt full of vague instructions still produces vague output, and padding a prompt with unnecessary context makes answers slower and more expensive. The useful test is whether each sentence in the prompt would change the answer if you removed it.

What are the four things most missing from a child's prompt?

How long the answer should be, who it is for, what shape it should come back in, and an example of what good looks like. Adding those four turns most weak prompts into strong ones without any technical knowledge at all.

Can a better prompt stop an AI from being wrong?

It reduces some kinds of error, particularly vagueness and wrong reading level, but it does not make an AI factually reliable. Children should be taught to check claims against a source they can see, which is why our pathway pairs prompting at level 2 with working from your own documents at level 4.

Keep reading

All posts