AI-Written Kids' Content Behind a Human Gate: Picture Books, Story Tasks and Practice Tables in Plain PHP
When you build learning content for children with an LLM, the model is the easy part. The hard parts are everything around it: who is allowed to see its output, how you catch its mistakes, and what you decide not to teac
When you build learning content for children with an LLM, the model is the easy part. The hard parts are everything around it: who is allowed to see its output, how you catch its mistakes, and what you decide not to teach at all. This is a write-up of three pieces we added to kids.findnix.eu, a curated, ad-free search engine for kids: AI-written picture books, "story tasks" for ages 9 to 14, and practice tables. All of it is plain PHP and MariaDB, no framework, no accounts.
The rule everything else follows
Kids never type into an AI. Nothing is generated while a child is looking at the page. Every text and every picture is produced in advance, reviewed by a person, and only then published. There is no chat box, no "ask the AI", and what a child types into an answer field never leaves their browser.
We documented the reasons in the site's (German) wiki: a model can be confidently wrong, you cannot filter every answer to a free-form question, and a free text box invites children to type their name, school and worries into a third-party API. Pre-generated and approved content avoids all three. The Kinder KI topic cards work the same way.
Piece 1: picture books, with a translation exercise
Each book has a cover and eight pages with two to four short sentences each, for ages 6 to 9, on topics such as courage, friendship, feelings, being new in class and bedtime. They are written in German. Under every page a child can try to translate the text into English and then click "show solution"; the English text exists only as that answer key. A PDF with writing lines and solution pages is generated for printing.
The generation pipeline is deliberately boring:
- An LLM returns strict JSON: title, blurb, eight pages, tips, a note for parents, an English translation and a character sheet. Malformed or incomplete JSON is retried up to three times.
- A second pass derives the image prompts from the final German text, one per page. When the model wrote the prompts alongside the story, they were sometimes shifted by one page, and they went stale the moment the text was edited.
- An image model draws the pages in one of six style presets (flat, watercolor, colored pencil, paper collage, gouache, pastel chalk) so the books do not all look alike.
- A human reads every text and checks every picture against its page. Anything wrong gets regenerated.
- Only the admin can publish. The PDF is built by a small dependency-free PHP writer (about 180 lines: A5 pages, core Helvetica, embedded JPEGs, text stays selectable).
What the review loop actually caught: pictures showing the whole cast in every scene, a "GYM" sign painted into an illustration despite a "no text" instruction, skin tone and clothing drifting between pages, German idioms that did not belong in a children's book, and a lost toy that was visible in the scenes where it was supposed to be missing. None of that is visible in the JSON.
Piece 2: story tasks that check the model's maths
A story task is a short story and exactly three tasks that get harder from the first to the third, with a field for the child's own solution and a "show solution" button with the answer and the steps. There are seven subjects (maths, physics and IT, chemistry, biology, geography, history, German) and five levels for ages 9 to 14. The maths track climbs from money and time problems to the rule of three, percentages and simple equations.
LLMs get arithmetic wrong, and a wrong answer key is the worst possible failure in a learning product. So every numeric task carries a machine-readable expression next to the prose answer, and the server re-computes it:
// calc is a plain arithmetic expression such as "(12*4)+8"; result is the number the model claims
$v = kids_task_calc($task['calc']); // small shunting-yard evaluator, no eval()
if (abs($v - $task['result']) > 1e-6 * max(1, abs($v))) $warn[] = 'result mismatch';
if (!kids_task_text_has_number($task['answer'], $v)) $warn[] = 'answer does not state the result';
If the check fails, the draft is discarded and regenerated, up to three times. It still does not replace a human recalculating each task, but it catches most slips before a person ever sees them.
The funniest bug of the whole project: our first batch for geography, chemistry, biology and history came back as maths word problems in costume. A history task asked the child to subtract 1450 from 2026. That is not what "history task" means. Only the maths track calculates now; the other subjects use matching, ordering, explaining and reading maps. Following a spec to the letter and still missing its point is not only a human talent.
Piece 3: practice tables
For each subject there are reference tables with a learn mode (everything visible) and a practice mode (one column hidden, the child types the answer and checks it). Currently 15 are live: all 193 UN member states with capitals and flags (including a "guess the flag" mode), the periodic table with all 118 elements, formula collections for maths and physics, a unit converter, a timeline of key dates, and German grammar tables for parts of speech, cases, articles, tenses and sentence parts, plus a few biology tables.
A table is stored as plain text lines, one per row:
Europe | Germany | Berlin | 1 | de
Geometry | Area of a rectangle | A = a ¡ b ## a*b ## ab | 1
The ## part holds spellings that count as correct but are not shown. The interesting problem is deciding when a typed answer is right. Names are compared leniently (case, umlauts and accents do not matter, every capital of a country with several counts). Formulas need the opposite treatment. Our first version stripped everything except letters and digits before comparing, which made a+b a correct answer for the area of a rectangle. Formula columns now use their own normalizer in which operators and brackets count and only the spelling of multiplication, powers and Greek letters is flexible:
function normF(s) {
s = String(s).toLowerCase().replace(/\s+/g, '');
s = s.replace(/²/g, '^2').replace(/Ī/g, 'pi');
return s.replace(/[¡Ã*x]/g, '') // a*b, a x b, a¡b and ab are the same
.replace(/[Ãˇ:]/g, '/')
.replace(/[âââ]/g, '-')
.replace(/,/g, '.');
}
The periodic table also has a colored overview: a CSS grid of 118 buttons in the classic layout, ten element groups with a legend that dims everything else, and a self-test mode that hides symbols or names until you tap a cell. The flags are SVG files from the MIT-licensed flag-icons project, self-hosted. Emoji flags were the first idea, but Windows shows them as two letters. The PDF writer only takes JPEG and the converters available on the server drew several flags wrongly, so the printable sheets leave flags out rather than ship wrong ones.
One honest limitation: the flags are decorative images, so the "guess the flag" mode is not usable with a screen reader. The normal text mode is.
What we deliberately left out, and why
There are no wars, no war zones, no sex education, and no description of biological or chemical processes that could be abused for weapons, explosives, poisons or drugs. In history the timeline contains inventions, discoveries and turning points; in geography, peaceful destinations and nature. The country table lists states and capitals and says nothing about conflicts.
The reasons are practical rather than ideological. A child reading alone cannot ask anyone when something frightens them, and a three-sentence task cannot treat a war neutrally or with the context it needs. Sex education belongs first with parents and then with school, where timing, vocabulary and the maturity of the individual child can be taken into account. And anything that reads like a recipe for hazardous substances is a legal and safety risk, so chemistry stays with salt, sugar, water and sand, and electricity experiments use batteries only.
Enforcement has three layers. The rules go into every LLM prompt. A word filter scans every text for terms about war, weapons and hazardous substances, and sex education, and a draft with a hit is discarded or flagged. The publish button refuses to publish while the filter has a hit. A human reads everything before it goes live. The filter matches words, not meaning, and it is strict on purpose: an older word list once flagged the German word for noodles because it contains "nude". We would rather rephrase a harmless sentence than let a hit through.
Where it stands
As of October 11, 2026: 33 picture books exist, 13 of them published; 36 story tasks across seven subjects are live; and there are 15 practice tables. The wiki documents the decisions in German: picture books, story tasks (including the section on what is left out) and practice tables.
If there is one lesson to take away: let the model draft, make the server check what can be checked mechanically, and keep a person between the output and the child. Try it at kids.findnix.eu/aufgaben.php; the content is German, because that is who it is for.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes â full credit and traffic to the original publisher.
