I asked an AI to make my career-switch video. It took 9 versions, 666 frames and no video editor
I'm a software engineer with 16 years behind me, now halfway through a double MSc in AI (EPITA × EM Normandie, the AIMS programme). In October I start looking for my end-of-studies internship. I needed one thing first: a
I'm a software engineer with 16 years behind me, now halfway through a double MSc in AI (EPITA × EM Normandie, the AIMS programme). In October I start looking for my end-of-studies internship. I needed one thing first: a short video that tells my story on LinkedIn.
I don't do motion design. I don't own After Effects. So I tried something else: I asked Claude Opus 5.5 in Cowork to make it with me, from a one-line prompt typed half in French.
This is the log of that session: what got built, how it works, where the AI pushed back, and what I'm taking away for my studies.
TL;DR
- Output: a 55-second, 1080×1350 stop-motion video in a cut-paper style, with music, sound effects and a narrator, plus a web player and a design system for my website.
-
How: every frame is drawn by JavaScript on a
<canvas>, rendered headless with Playwright, encoded with ffmpeg. The soundtrack is generated in Python. The voice comes from ElevenLabs through an MCP connector. - Effort: 9 versions of the video in one evening. My job was the brief, the feedback and the ears. The AI did the drawing, the code and the rendering.
The brief (and the questions it asked first)
My first message was roughly: "create in stop-motion / JS / Canva HTML a video telling my move into AI (leaving Prisma Media, the MSc AIMS) to improve…", and I hit Enter mid-sentence.
Instead of guessing, it asked four questions with clickable answers: what the video is for, which format, which tone and length, which language. I picked LinkedIn personal branding, HTML + MP4, punchy ~45 s, English. That took ten seconds and saved at least two wrong versions.
As a marketing-strategy student, this is the part I'd frame as a mini VSL (video sales letter): a hook, a story, proof, and one call to action. The AI didn't use that word, but the structure it proposed was exactly that.
How a "stop-motion" video is made in code
The look is cut paper on a green cutting mat. Stop-motion is mostly two tricks:
- 12 frames per second, "on twos": movement is deliberately choppy.
- Boil: every piece moves by a pixel or two on every frame, as if a hand had re-placed it.
The whole video is one pure function: draw(frameNumber). Given a frame number it paints that exact frame, with deterministic randomness. That's what makes it renderable, repeatable and cheap to fix.
// deterministic noise: same frame + same id => same jitter
function hash(a, b) {
let h = Math.imul(a | 0, 374761393) + Math.imul(b | 0, 668265263);
h = Math.imul(h ^ (h >>> 13), 1274126177);
return ((h ^ (h >>> 16)) >>> 0) / 4294967296;
}
const jit = (f, id, amount) => (hash(f, id) - 0.5) * 2 * amount;
// every paper piece goes through put(): position + rotation + a little boil
function put(f, id, x, y, rot, scale, draw) {
ctx.save();
ctx.translate(x + jit(f, id * 7 + 1, 1.7), y + jit(f, id * 7 + 2, 1.7));
ctx.rotate(rot + jit(f, id * 7 + 3, 0.006));
ctx.scale(scale, scale);
draw();
ctx.restore();
}
Paper pieces are polygons with slightly irregular edges, a drop shadow and a grain texture. Scenes are plain functions of local time: pop in, hold, slide out. The renderer is eight lines of Playwright that call draw(i) for each of the 666 frames and save a JPEG; ffmpeg turns them into an MP4.
The same file also runs live in the browser as a player with a scrubber and chapter buttons, so I could review a scene without waiting for a render.
Nine versions, one evening
| Version | What I asked for | What changed |
|---|---|---|
| v1 | The first draft | 7 scenes, paper cut-outs, 45 s, silent |
| v2 | "From software engineer to AI engineer", a little character of me, the real master's name | Character with a "Software Engineer" cap that swaps for an "AI" cap |
| v3 | Focus on GenAI, not ML. Don't say I was let go. Add sound. | A signpost scene ("Same road" → "New path: AI"), a chat window that streams an LLM answer, procedural music and sound effects |
| v4 | Shorter beard, school logos (my files), less repetitive music | The track now changes with each scene, with a build-up and a drop on the cap swap |
| v5 | Hoodie back, a laptop with the logo I supplied, a lighter beard, a glitch at frame 150 | Fixed a scene that hid its props one frame too early |
| v6–v7 | A minimalist end card with a QR code that points straight to my site, with UTM tags | QR code generated locally, verified by decoding the final MP4 frame |
| v8 | "AI" is too close to the end of the sign | Longer sign |
| v9 | The ElevenLabs narration | Scenes re-timed to the voice, music ducked under it |
Two things made this loop fast. First, it rendered contact sheets (a grid of key frames) and looked at them itself before sending me anything; it caught overlapping text and a cut-off arm that way. Second, when I asked for a new look for "me", it showed me one still frame for approval before re-rendering 666 frames.
Sound: music, effects and a narrator
This is the part that surprised me most.
Effects come from the animation code. Every time a piece pops in, slides or flips, the scene code logs a cue with its exact timestamp. The audio script reads that list and places a sample on each cue: pops, paper slaps, whooshes, a wooden knock for the signpost, pencil scratches for the checkmarks, keyboard clicks while the chat types. Change the animation and the sound follows.
const popS = (t, d, dur = .42) => { cue('pop', d); /* …scale-in easing… */ };
The music is generated in Python with numpy: plucked arpeggios, bass, a few drum sounds, a pad. My feedback after v3 was "too repetitive", so v4 got one arrangement per scene: a quiet intro, a groove, a breakdown, a build-up into the cap swap, a chorus with a lead melody, and an outro.
The narrator is a third-person, slightly self-mocking storyteller ("He talks to language models all day. They talk back."). Once I connected the ElevenLabs connector, Claude picked a voice from the library and generated the takes. Then it found the sentence boundaries by measuring silences, re-timed every scene to the voice (the new cap lands exactly on "He switched caps"), and lowered the music automatically under the speech.
Where it pushed back, and where it couldn't help
This is what I'd want to know before trying it myself.
- Logos. It wouldn't draw or "vectorise" other companies' logos (the schools, Apple), even with my OK. It used the official files I provided, unchanged. Mildly annoying at the time, and fair in hindsight.
- It can't hear. It checked loudness per section and the timing of every cue, but it told me plainly that someone needed to listen. That was my job.
- Sandboxed network. Its workspace couldn't reach Google Fonts, so it installed the fonts from npm. It couldn't download the finished ElevenLabs MP3 either, so I downloaded it and dropped it in the chat.
- My own accounts. ElevenLabs refused three of the four takes because my free account had been flagged. One take went through, and that was enough.
- The story is mine to tell. I asked it not to say I was let go and to frame the change as a choice. It rewrote the scene, and later reminded me that the brand names in the video didn't match the ones on my CV, which a recruiter would notice.
What I'm taking back to my MSc
- Clarifying questions are a feature. The best part of the first minute was the AI refusing to guess.
-
Determinism makes iteration cheap.
draw(frame)as a pure function meant every fix was a small diff and a re-render, never a re-shoot. - Close the loop with evidence. Contact sheets, a decoded QR code, measured loudness: the model checked its own output before showing it to me. That's the same discipline I want in LLM products: evaluate before you ship.
- Agents connect tools. Canvas, Playwright, ffmpeg, numpy, ElevenLabs, Vercel and a QR library sat in one session. The interesting work was the glue and the judgement, not any single tool.
- Marketing still needs a strategy. Hook, story, proof, call to action, and a UTM-tagged QR code so I can see in Google Analytics whether the video brings anyone to my site.
Bonus: the video became my website's design system
Once the video was done, I asked for the look to be exported so another coding agent can rebuild my website in the same style. I got colour and type tokens, the fonts, twelve HTML/CSS components (paper labels, sticky notes, the checklist, the chat window, the signpost…), the character in four poses, and a content deck with every section's copy. That's the next project.
I'm Hamdi, a software engineer turning AI engineer, looking for a 6-month internship as an Applied / Forward Deployed AI Engineer. If you're building with LLMs in production, I'd like to hear from you: h4md1.fr · LinkedIn.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.