Drafted with AI assistance and reviewed, fact-checked and edited by hand before publishing - see our AI transparency page for how we use AI across the app.
If your child is 5-7 and practising a word like "fog" or "hop," they'll usually see more than just the word and the audio. Tap Visual hint and a small, bright illustration appears - a cartoon street lost in grey mist, a rabbit mid-jump - drawn specifically to show what that word means, with nothing written on it. It's a small thing on the screen. Behind it is a feature we think is a genuine differentiator: every one of those pictures is generated by AI, purpose-built for that exact word, reusable by every family who reaches it after you, and - this is the part most parents don't know - something you can improve yourself if the AI's first guess isn't quite right.
This post is the detail we haven't written down anywhere else: how the pictures get made, how they turn up as a reward mid-game, how to nudge a stubborn one into something better, and why one household generating a picture quietly makes the app better for every other family on Spelling.Live.
Where a child actually sees one
Picture hints are an early-years feature (ages 5-7) - the band where a concrete image genuinely helps a child hold onto what a short word means, rather than an abstract word like "because" that a picture can't really capture. Two places show them:
- The Visual hint pill, during practice. While spelling a word, a child can tap for a picture clue alongside the audio and phonics hints - a memory anchor, not the answer.
- Balloon Pop, as a reward. Pop the balloon carrying the correct spelling and, alongside the coins, that word's picture hint falls gently down the screen as a little burst of positive feedback - the image reappearing at the moment they got it right, reinforcing the word without ever having generated anything new mid-game.
Both read from the same place: a picture generated once for a word, in a given age band, and stored for reuse - never redrawn on the fly during a timed game.
How the picture actually gets made
Pictures are generated by OpenAI's gpt-image-1, prompted to produce a flat vector illustration - think Duolingo or a modern children's picture book, not a photo: bold outlines, a single clear subject, a clean white background, cheerful and safe. Every prompt carries one hard rule: the image must contain zero text - no letters, no numbers, no captions of any kind. That's not a style choice, it's a safety one - image models will happily write the word itself as a caption under the picture if you let them, which would hand a spelling app's child the answer before they've tried. The prompt describes the scene instead ("a cartoon rabbit mid-jump, cheerful, on a plain background") and never the word itself.
Some words need more care than "draw the object." A handful are hand-tuned so the AI doesn't take an odd turn on words that are short, easy to spell, but easy to illustrate badly:
- "sob" - a cartoon child with tears on their cheeks, gently. Friendly, not distressing.
- "got" - a happy child holding a wrapped gift, so the abstract idea of having received something has a concrete scene.
- "fog" - a street lost in thick grey mist, spooky in a friendly way rather than frightening.
The guardrail that matters most for a spelling app specifically: the prompt is built from the word's meaning, never the word's letters - so the picture can never accidentally spell out the answer it's meant to be a clue for.
Honestly, though: that's a strong instruction to the model, not a hard filter, and it isn't foolproof. We've tried more than one model behind this feature - the scheduled queue originally generated pictures with Google's Gemini/Imagen, and that output needed more manual cleanup than we'd have liked, occasionally including a stray word rendered right into the picture despite being told not to. We've since standardised the whole pipeline, on-demand and scheduled, on OpenAI's gpt-image-1, which respects the no-text instruction far more reliably - but "far more reliably" isn't "always." Every so often, a model still slips the word itself into a picture. We deliberately don't run a second, paid AI check over every image to catch this - that would roughly double the cost of a feature that's already the priciest thing in the app, to guard against a genuinely rare failure. So if you ever spot one, hitting Regenerate almost always produces a clean version on the next try - and if anything about a picture needs a second look from us, the Feedback button (bottom corner of the app) reaches us directly.
Two ways a picture gets generated
On-demand, in seconds. When a parent hits "Regenerate Image" in the word-list editor, or a child taps a Visual hint that doesn't exist yet, the app calls gpt-image-1 directly and the picture appears in the app within moments - no waiting.
On a schedule, for everything else. New word lists - a photographed school sheet, a pasted spelling list, a whole class assigned in one go - can add dozens of new words at once. Rather than generating all of them synchronously (slow, and expensive if a parent never actually needs half of them), those jobs join a queue. A background worker, running as a scheduled job every hour, drains that queue a batch at a time until every new word has its picture. Practically: add a fresh list today, and most pictures are ready within the hour, often sooner.
You can make the picture better
This is the part we've never written up before: the AI's first attempt at a picture isn't the end of the conversation. Open a word in the parent word-list editor and there's a field labelled "Describe the meaning for the image" sitting right under the picture, with a "Regenerate Image" button beneath it.
Why this matters: a lot of English words are ambiguous out of context, and an AI illustrator has to guess which meaning you mean. "Bark" could be a dog or a tree. "Fair" could be a funfair or a judgement. If the first picture guesses wrong, you don't have to live with it - type a short description ("a dog barking," not "tree bark") and regenerate. The description you type is saved against that word for your list, so the next regeneration (yours, or anyone else's on a shared list) starts from your steer rather than the AI's first guess.
This is also exactly where to go if you spot the leaked-word problem from the previous section - the real example that prompted us to write this post was the word "sun," where the model rendered a cheerful illustration with "SUN" lettered in underneath it, undermining the entire point of a spelling clue:
Shared across the whole Spelling.Live community
Here's the differentiator we think matters most: a picture, once generated, belongs to the word - not to your household. Every image lives on the shared word bank, keyed by the word and its age band, not by which family or class asked for it. The first time any family on Spelling.Live needs a picture for "hop" in the early band, it gets generated. Every family after that - yours, a class of 30 pupils a teacher assigns in one sitting, a household on the other side of the country - gets that same picture instantly, at no cost and no wait, because it's already sitting in the shared bank. (This is the same principle that already makes assigning a list to a whole class fast: one list assigned to 30 pupils generates its audio and images once, not once per child.)
That means the picture library isn't a fixed asset we drew once and shipped - it's a living, growing resource that gets a little more complete every time any family on the platform reaches a new word, and every parent who takes thirty seconds to refine a picture with a better description leaves it better for the next family who reaches that word too.
What this costs, honestly
Generating a picture is a genuinely more expensive AI call than a spoken clip or a line of hint text, so it draws more from the monthly AI credit allowance than most other features - 5 credits against 1 for a chat hint, a TTS clip, or a handwriting check. That's a deliberate, transparent choice: we'd rather the cost be visible than hidden, and rather charge fairly for the expensive thing than throttle it invisibly. Because images are shared globally, though, most requests never cost anything at all - you're only charged the first time a specific word needs generating for its age band, not every time your child (or anyone else's) taps to see it again. Read how credits work more broadly for the full picture.
The short version
- Picture hints are AI-generated, purpose-built illustrations for early-years words (5-7), shown as a memory clue during practice and as a falling reward when a child pops the right balloon in Balloon Pop
- Every prompt is built from the word's meaning, never its letters - and on the rare occasion a model still slips the word into the picture anyway, hitting Regenerate almost always fixes it
- Most pictures generate instantly on demand; a big new word list drains through an hourly background queue
- You can improve a picture - describe what you actually mean in the word editor and hit Regenerate; it's the same fix whether the picture guessed the wrong meaning or accidentally lettered in the word
- Once generated, a picture is shared with every family and class on Spelling.Live from then on - the library gets better as the community uses the app, not just when we sit down and draw more of it ourselves