Quase pronto para ir
Carregando...

Educational Videos for Kids with AI: The Two Word Pipeline Behind a 30 Episode English Cartoon Series

Educational Videos for Kids with AI: The Two Word Pipeline Behind a 30 Episode English Cartoon Series

Type "Episode 2". Get a scripted scene, a keyframe with the lesson phrase in a speech bubble, and a looping animated clip with a narrator voice. The curriculum lives inside a prompt.

Most educational videos for kids die of production overhead. You need characters, a curriculum, illustrations, animation and a voiceover — per episode. That is why so many kids learning channels post three videos and stop, and why teachers who would love custom ESL material for their class never make it.

See it live: open the English for Kids Cartoon canvas and explore every node in this article.

This case study documents the most automated Zoviz Canvas pipeline we've built so far: an episode by episode English cartoon series where the entire input for a new episode is typing two words. Literally "Episode 2". The canvas does the rest: an LLM writer with a 30 episode curriculum built into its prompt picks the topic, writes the scene, and drives an image node and a video node that output a vertical keyframe with the lesson phrase in a speech bubble and a looping animated clip with a narrator voice.

The series is "English for Kids", starring Toby the fox cub, Dede the duckling and Mimi the bunny, aimed at ages 3 to 7. Below is the full anatomy: the curriculum, the prompts, the outputs for Episode 2 ("The apple is red."), the costs, and how to make the cast and the program your own.

What you get from typing "Episode 2"

The pipeline has exactly one variable input: a tiny prompt node where you type the episode. Everything else is standing infrastructure. For Episode 2, the canvas produced a production script, a keyframe and a finished clip:

1. A production script, written by the Writer LLM.

It looked up Episode 2 in its internal program (Colors — "The apple is red."), then wrote one production prompt in three labeled parts: KEYFRAME (the scene), VIDEO MOTION (the animation) and SOUND (the narration). One generated script drives both the image and the video node.

2. A vertical keyframe with the phrase spelled exactly.

The three friends gather around a shiny red apple, with a big clean speech bubble reading "The apple is red." and a small caption at the bottom: Episode 2 of 30 - English for Kids.

3. A looping animated clip with a narrator voice.

Toby bounces with the apple, Dede taps it with her wing, Mimi nods behind her glasses — and a warm narrator says the phrase twice: once playfully, then slowly, word by word: "The… apple… is… red." The clip ends near its starting pose so it loops cleanly on Shorts, Reels and TikTok.

One reference photo. One 30-day plan. Zero filming.
See how a single canvas keeps the same AI trainer's face, body, and voice consistent across 30 straight workout reels, just by typing a day number.
Check the Fitness AI Influencer Canvas →

The one time setup: a character sheet

Before the first episode, you run a single setup node: the character sheet. It defines the cast once, with names written under each character, and every future keyframe references it so the cast stays identical across all 30 episodes.

The one time character sheet. Edit this one prompt and the whole series changes cast.
CHARACTER SHEET: three cartoon characters standing side by side in a sunny green meadow, full body, facing camera: TOBY a small round orange fox cub with big friendly eyes and a tiny green scarf; DEDE a little yellow duckling with a small red backpack; MIMI a fluffy pink bunny with round glasses; their names written in clean simple letters under each character, soft 3D cartoon style, pastel colors, rounded toy like shapes, gentle warm lighting, kids TV aesthetic, plain soft sky background.

The Writer LLM: a curriculum inside a prompt

The interesting engineering in this canvas is that the whole ESL lesson plan lives inside the Writer LLM's system prompt. It opens by defining the role and the cast:

You are a childrens cartoon writer and English teacher who makes learning feel like play for kids aged 3 to 7. Your series stars three friends: TOBY a small round orange fox cub with big friendly eyes and a tiny green scarf; DEDE a little yellow duckling with a small red backpack; MIMI a fluffy pink bunny with round glasses.

Then it hard codes the full 30 episode program, one topic and one phrase per episode. The arc runs from first words to farewell: Hello ("Hello! I'm Toby!"), Colors, Numbers 1 to 3, Animals, Family, Food, Body, Weather, Feelings, Clothes, Numbers 4 to 6, Fruits, The Farm, My House, Toys, Being Polite, The Park, Water, Night, Morning, Shapes, Big and Small, Fast and Slow, Up and Down, The Sea, Birthday, Helping, The Rain, Counting to 10, and Goodbye ("Goodbye, see you soon!"). Ask for an episode above 30 and it cycles back through the program.

The output rules encode the pedagogy and the production standards. The KEYFRAME must show the three friends acting out the topic together, "same characters as the reference image", one big clean white speech bubble "containing the episode phrase in large friendly letters spelled EXACTLY", the fixed style line, 9:16 vertical framing, and a caption reading exactly "Episode [number] of 30 - English for Kids".

The VIDEO MOTION must animate the topic with bouncy playful movements "while a warm friendly narrator voice says the episode phrase two times, first playfully, then slowly and clearly word by word, the speech bubble text staying sharp and unchanged the whole time, the clip ending near the starting pose so it loops."

The SOUND rules keep it kid safe and clean: the narrator speaking clear simple English, soft giggles and gentle sounds matching the scene, "no loud music, no other speech."

And one escape hatch: "If the user message names a topic or phrase or adds scene or styling wishes, follow those instead of the program defaults." That single line turns a fixed curriculum into a flexible one — type "Episode 5, set it at the beach" or a topic of your own and the writer adapts.

Follow a fictional beauty brand from a blank page to a broadcast-quality AI film: logo, brand kit, product photography, and a zero-cut commercial made with invisible frame transitions.
See how Lumelle was built →

How to run an episode, step by step

The HOW TO USE note pinned to the canvas is four steps:

  1. One time setup: run the character sheet image once and check your cast. Edit that one prompt to change the characters and make the series your own.
  2. Type just the episode in the small prompt node — for example: Episode 2.
  3. Run the Writer LLM, then the keyframe image node, and check the scene and the speech bubble spelling (about 7 credits).
  4. Run the video node for the animated clip with the narrator voice (about 50 credits).

An episode lands around 57 credits all in, which puts a full 30 episode season around 1,700 credits. The checking step in the middle matters more here than in any other canvas: on screen text is the hardest thing AI image models do, so the pipeline is built to verify the speech bubble spelling at the cheap image stage before spending on video.

Why this design works

Canvas of Educational English Cartoon

The text is the lesson, so the text is protected. Most AI video workflows avoid on screen text entirely. A language lesson can't — the child needs to see the phrase while hearing it. This canvas defends the phrase in three places: the writer must spell it EXACTLY in the keyframe prompt, the video prompt pins the bubble "sharp and unchanged the whole time," and the human checkpoint verifies spelling before the video renders.

Repetition is pedagogy, not filler. The narrator saying the phrase playfully first and then word by word is a standard early language teaching pattern, encoded once in the system prompt and inherited by all 30 episodes.

Loops multiply watch time. Kids rewatch; algorithms reward it. Ending each clip near its starting pose makes every episode naturally loopable on Shorts, Reels and TikTok.

Two words per episode is a real production rate. Once the character sheet exists, producing kids learning videos becomes a queue: Episode 3 tomorrow, Episode 4 the day after. That cadence is exactly what a faceless kids YouTube channel needs and almost never has.

Build your Cartoon Logo with Zoviz in minutes

Who this is for

Three audiences map cleanly onto this canvas. Creators looking for faceless YouTube channel ideas get a series format with recurring characters, a built in 30 episode roadmap and vertical, loopable output. Teachers and ESL tutors get custom lesson clips matched to their class — the off program escape hatch means "Episode 12, but with the fruits we studied this week" is a valid input. And parents get short, calm, ad free educational videos for toddlers and preschoolers with their own children's names and favorite animals, one character sheet edit away.

The same architecture also generalizes past English: swap the curriculum block in the Writer LLM prompt and the same cast can teach numbers, colors in another language, safety rules or bedtime routines.

FAQ

How do you make educational videos for kids with AI?

The pattern in this canvas: define a cast once with a character sheet image, put the curriculum and the production rules inside an LLM system prompt, then have that LLM write one production script per episode that drives an image node (the keyframe with the lesson text) and a video node (the animated clip with narration).

How much does one episode cost?

About 57 Zoviz credits: roughly 7 for the writer plus the keyframe, and about 50 for the 9:16 animated clip with the narrator voice. A full 30 episode season runs around 1,700 credits.

Can I change the characters or the curriculum?

Yes, both live in editable prompts. Rewrite the character sheet prompt to change the cast, and edit the 30 episode program inside the Writer LLM prompt to change the topics, phrases or even the language being taught.

Does the speech bubble text come out spelled correctly?

The pipeline is designed around that risk: the writer is instructed to spell the phrase exactly, and the workflow has you check the keyframe image before rendering video. If the bubble is wrong, rerun the keyframe; the check costs 7 credits, not 57.

Is this suitable for a kids YouTube channel?

The output format is built for it: vertical 9:16, looping clips, consistent recurring characters, an episode counter in every frame, and no on camera presenter; a complete answer to how to start a kids YouTube channel without filming anything.

Receba newsletters exclusivas

Inscreva-se agora para receber as últimas novidades diretamente na sua caixa de entrada. Não perca!