AI & Automation17 min read

Claude Opus 5.5 Motion Graphics: How It Built a 4K Product Video From Code (Full Guide + Free Skill)

How Claude Opus 5.5 made a 30-second 4K60 motion graphics product film with synced original audio, entirely from code. The workflow, the maths, the before/after fixes and a free reusable Claude skill.

Claude Opus 5.5 Motion Graphics: How It Built a 4K Product Video From Code (Full Guide + Free Skill)

Table of Contents

  • 01.Watch the film
  • 02.What people think AI video is vs. what Claude actually made
  • 03.How does Claude Opus 5.5 actually work on a task like this?
  • 04.The creative concept: signal from noise
  • 05.Step-by-step: how to make a motion graphics video with Claude
  • 06.The one skill: reuse this workflow with Claude
  • 07.Limits
  • 08.FAQ
  • 09.Resources

TL;DR: I gave Claude Opus 5.5 (running in Claude Code) one long creative brief and asked for a finished 30-second launch film for SaaSearch.io. It didn't generate pixels the way text-to-video models do. It researched the live website with a scripted browser and rebuilt the real interface in HTML. It animated that interface with about 1,900 lines of deterministic code, composed and synthesised an original soundtrack from the same timeline, and captured 1,800 frames at 3840×2160 with Playwright. ffmpeg then encoded them into a 60 fps MP4. Along the way it reviewed its own renders, found real problems (a state-leak bug, an overlapping animation, a placeholder image) and fixed them before the final render. Below you'll find the film, how it works, the maths, the mistakes, and a free SKILL.md you can use to do the same.


Watch the film

30 seconds · 3840×2160 master · 60 fps · original synthesised soundtrack. The 1080p web version is shown here.

🎧 Soundtrack only: saasearch-film-soundtrack.mp3

The numbers (all measured from the final files)

Metric Value
Duration / frames 30.00 s · exactly 1,800 frames
Resolution / frame rate 3840×2160 (4K UHD) · 60 fps
Video codec H.264 High, CRF 15, BT.709 · 10.3 Mbps · 38.8 MB
Audio AAC 320 kbps, 48 kHz stereo · −13.9 LUFS integrated · −1.0 dBTP true peak · LRA 3.0 LU
Code written ~1,900 lines (renderer 466, components 215, timeline 107, easing 72, audio synth 284, capture 85, plus scripts)
Real assets pulled from the live site 49 product images and favicons, 10 font files, 16 research screenshots
Sound cues synced to picture 84 (keystrokes, clicks, whooshes, tones, bells)
Frame capture time 385 s for 1,800 4K frames (5 parallel browser pages ≈ 214 ms/frame)
Video editing software used None. No After Effects, Premiere or DaVinci; just Node.js, Chromium and ffmpeg

What people think AI video is vs. what Claude actually made

Most people hear "AI made a video" and picture a diffusion model dreaming up pixels from a prompt. That's not what happened here, and the difference matters.

What people usually assume What actually happened in this build
"Claude generated the video frames like Sora or Veo." Claude wrote software. Chromium drew every pixel from HTML and CSS; ffmpeg encoded them.
"AI video is random. You can't get the same shot twice." Every frame is a pure function of time t. I tested 15 timestamps rendered in scrambled order against fresh renders: byte-identical.
"The AI will invent a fake UI that looks like the product." It drove the live saasearch.io with Playwright, typed the real query, clicked the real tabs, and rebuilt the UI from the site's own colour tokens, fonts and data.
"AI hallucinates numbers to make it look impressive." The brief banned invented metrics. Every product name, date, tag and result count on screen was captured from the live site that day, or from the founder's own dashboard screenshots.
"One prompt, one shot." It worked like a studio: research → pre-production docs → keyframes → low-res preview → critique → fixes → final render → QC.
"AI can't do audio sync." The score and SFX were synthesised from the same timeline file as the animation, so every keystroke and click lands on its frame.
"It's flawless." It wasn't. It shipped a state-leak bug, an overlapping morph and a placeholder image, then caught all three itself. It also cannot hear, so the audio was verified with meters, not ears.
"You need a motion designer's tools." You need a good brief, a terminal, and the patience to let it iterate.

The short version: Claude doesn't make videos. It makes programs that make videos. That's why the result is precise, editable and reproducible: change one number in the timeline and re-render.


How does Claude Opus 5.5 actually work on a task like this?

Claude Opus 5.5 is a large language model. On its own it reads and writes text and code, and it can look at images. In Claude Code (I used the desktop app), it becomes an agent with tools: it can run shell commands, read and edit files, drive a browser, and open the images it produces.

That combination is the whole trick. The loop looks like this:

  1. Plan. It turns the brief into phases, a storyboard with exact timings, and a technical plan.
  2. Act. It writes code, runs it, installs what's missing (the ffmpeg bundled with Playwright is a stripped-down build, so it pip-installed a full one with H.264 and AAC).
  3. Look. It renders stills and contact sheets and views them. This is the step most people don't realise is happening.
  4. Critique. It writes down the worst problems with timestamps.
  5. Fix and repeat until the frames hold up.

When I suggested mid-task that it "take a remote session to see how the clicks and pages perform", it scripted a headless browser session on the live site. It typed the query character by character, captured the suggestions dropdown, opened the results, the details panel, Images, News, Submit and a product page, and pulled the real text and image URLs into JSON. Everything in the film is built from that capture.

Pipeline diagram: brief, research, assets and pre-production feed timeline.js, which drives both the picture renderer and the audio synth; frames are captured and encoded, with a critique loop back into the code


The creative concept: signal from noise

The brief asked for a premium launch film, not a template explainer, with three messages: discover software, get discovered, grow your presence. Claude's concept was signal emerging from noise:

  • 01 The noise (0–4 s): a drifting dark field of real product names, queries and category labels. "Too many tools. / Not enough discovery."
  • 02 The search (4–9 s): every fragment collapses onto one line, the line becomes the search bar, and the bar expands into the page. The camera pushes in while "AI tools for productivity" is typed.
  • 03 The SERP reveal (9–15 s): the real results page; push-in on the first result, a click opens its details panel, then the page steps aside: "Find the tools that move you forward."
  • 04 Beyond web results (15–20 s): Images → News → search suggestions, with the tab underline travelling between them.
  • 05 A home for founders (20–26 s): Submit → AI Autofill → a verified product page → the founder's SEO toolkit and "Write Article". "Build it. List it. Get discovered."
  • 06 The brand moment (26–30 s): everything unwinds into the opening search composition; the bar types "SaaSearch.io".

Scene 01: dark field of real product names with the headline "Too many tools. Not enough discovery."

The best idea in the film is object identity: nothing just appears. Fragments become a line, the line becomes the bar, the bar becomes the page, a click becomes a panel, and the header becomes the end card.

Sequence from 2.9 s to 5.5 s: fragments converge, a line draws, the line becomes the search pill, the pill expands into the page, the camera pushes in

Scene 02: the SaaSearch wordmark and search bar, camera pushed in while the query is typed

Scene 03: camera pushed in on the first real search result, OzBrain

Scene 04: the Images tab with real product imagery and the "More ways to discover" statement

Scene 04: the real search-suggestions dropdown

Scene 05: the founder flow from Submit to the verified product page and SEO toolkit

Scene 05: the founder's private SEO toolkit card with "Build it. List it. Get discovered."

Scene 06: end card with the SaaSearch wordmark, "Discover what's next.", a search bar reading SaaSearch.io, and "Search SaaS & AI. List your product for free."


Step-by-step: how to make a motion graphics video with Claude

This is the exact workflow, in the order it ran. The free skill at the end packages all of it.

Step 1: Write a brief that bans the usual AI failure modes

The brief was long, and that's a feature. The parts that mattered most:

  • Hard specs: 30.00 s, 3840×2160, 60 fps, H.264 + AAC, every property a deterministic function of time.
  • A six-scene storyboard with timings and approved copy.
  • Truth rules: use the real interface; don't invent metrics, testimonials or rankings; don't imply a Google affiliation; only say "free" where it's actually free.
  • An explicit list of things to avoid: purple AI gradients, random particles, glassmorphism, spinning 3D, unreadable tiny UIs.
  • Mandatory review stages: keyframes, low-res preview, a critique naming the three worst problems with timestamps, fixes, then final QC.

Step 2: Research the real product with a scripted browser

Claude wrote research.mjs (Playwright) to open the live site, type the query, wait for suggestions, press Enter, and capture the results. Then it visited Images, News, About, Submit and a product page. Two practical lessons it hit and fixed:

  • waitUntil: 'networkidle' never resolves on a site with analytics. Use 'load' plus a short delay.
  • Responsive sites often have a hidden duplicate search input, so target input[name="q"]:visible.

It also read the codebase, and that's how it discovered the wordmark isn't an image at all. It's live text: Plus Jakarta Sans, "SaaS" bold in four colours, "earch" in grey. So the logo renders as vector type, razor-sharp at 4K, instead of an upscaled PNG.

Step 3: Pre-production docs before animating

It produced a creative direction, a style guide (tokens copied from globals.css, two type families, motion principles), a storyboard with times, an asset inventory separating real material from supporting graphics, and a technical plan. Writing these down first is what kept 30 seconds of animation coherent.

Step 4: Build a deterministic renderer

The core contract is one function:

window.__render = (t) => render(t); // t in seconds; sets EVERY style from t alone

No CSS animations, no timers, no Math.random(). Randomness comes from a seeded PRNG (mulberry32), so the "noise" field and the human-feeling typing rhythm are identical every render. The UI is built once, and render(t) only mutates transforms, opacity, clip-paths and text.

Step 5: Put every time in one file

timeline.js holds scene bounds, 50+ key times, and the keystroke schedules, and it exports the audio cue sheet. The renderer imports it and so does the audio synth. That's why sync is exact by construction: there's no separate audio edit to drift.

Timeline: six scene blocks across 30 seconds with 84 audio cues for keystrokes, clicks, whooshes, tones and hits

Step 6: The maths behind the motion

Every movement in the film comes from four curves and one camera equation.

Easing (CSS-identical cubic Bézier). For control points P1 = (x1, y1) and P2 = (x2, y2):

x(u) = 3(1−u)²u·x1 + 3(1−u)u²·x2 + u³        solve x(u) = t  (Newton's method, bisection fallback)
y(u) = 3(1−u)²u·y1 + 3(1−u)u²·y2 + u³        eased progress = y(u)

ease.out   = (.16, 1, .3, 1)    arrivals: words, cards, panels
ease.inOut = (.65, 0, .35, 1)   camera moves and layout morphs
ease.in    = (.55, 0, .9, .45)  exits

Spring (closed form, deterministic at any t):

y(t) = 1 − e^(−ζωt) · ( cos(ω_d t) + (ζω/ω_d) · sin(ω_d t) )
ω = 2πf,  ω_d = ω·√(1 − ζ²)        used with ζ = 0.45, f = 3 Hz for the "Verified" badge pop

Plots of the four easing functions used in the film: ease.out, ease.inOut, ease.in and a damped spring

The camera. The real UI is built at its true 1440 px width inside one element, moved by a single transform:

canvas = T + s · page
Full-bleed:   s = 1920 / 1440 = 4/3,  T = (0, 0)
Focus page point P at canvas point C:   T = C − s·P
Search push-in:  s = 2.0667 on P = (720, 333)  →  T = (−528, −148.2)
Split layout:    s = 0.88,  T = (1920 − 1440·0.88 − 40, (1080 − 810·0.88) / 2) = (612.8, 183.6)

Three states only (full, push-in, split), interpolated with ease.inOut. Limiting the camera to three states is a big part of why it feels "directed" rather than floaty.

The noise collapse. Each fragment flies to a point on the future search bar:

target_x = 960 + clamp(0.42 · (x − 960), −380, 380),   target_y = 540
start_i  = 2.55 s + 0.45 · min(1, |x − 960| / 1100) + 0.22 · seed_i     (outer fragments arrive last)

Human typing rhythm:

t[i+1] = t[i] + b · (0.75 + 0.5 · r())  +  (char is space ? 0.35 · b : 0),   b = 58 ms, r = seeded PRNG

Resolution and throughput: a 1920×1080 CSS viewport × deviceScaleFactor 2 = 3840×2160. 30 s × 60 fps = 1,800 frames; 385 s ÷ 1,800 ≈ 214 ms per frame across 5 workers.

Step 7: Original sound design, synthesised in code

No stock music and no samples, so there are no licensing questions. audio.mjs synthesises:

  • The score: 120 BPM in D major. Dark Bm9 pads under the noise, a bright Dmaj9 bloom as the bar opens, a groove that drops on the SERP reveal at 9.0 s, a breakdown for the founder "moment of clarity", and a four-note sonic logo (D–F♯–A–D), one bell per coloured letter of the wordmark.
  • The instruments: polyBLEP saw pads, FM plucks and bells, sine/saw bass, a synthesised kick that side-chains the pads, noise claps and hats, and a Schroeder-style reverb.
  • SFX on cue: 84 events, including each keystroke (seeded variation), mouse clicks, whooshes on camera moves, and confirm tones.
  • Mastering: a two-pass loudnorm to −14 LUFS / −1 dBTP. The result: −13.9 LUFS, −1.0 dBTP, flat factor 0 (no clipping), no unintended silences.

Waveform and spectrogram of the soundtrack: sparse intro, typing clicks, groove from the SERP reveal, a breakdown, the sonic logo and a clean decay

One limit: Claude can't hear. It verified the audio with loudness meters and by looking at the waveform and spectrogram, then told me to listen before publishing, which is the right call.

Step 8: Capture with Playwright, encode with ffmpeg

await page.evaluate((t) => window.__render(t), t);
await page.evaluate(() => new Promise(r => requestAnimationFrame(() => requestAnimationFrame(r))));
await page.screenshot({ path, type: 'jpeg', quality: 95, animations: 'disabled' });
ffmpeg -framerate 60 -i f_%05d.jpg -i audio_master.wav \
  -vf "scale=out_color_matrix=bt709:out_range=tv,format=yuv420p" \
  -c:v libx264 -profile:v high -preset slow -crf 15 \
  -c:a aac -b:a 320k -movflags +faststart saasearch_launch_film_4k.mp4

A 960×540 preview renders in about a minute, so iteration is cheap. Only the approved cut goes to 4K.

Step 9: The critique loop (the part that makes it good)

After the first full preview, Claude reviewed a contact sheet with a frame every half second. It named the three worst problems with timestamps, fixed them in code, and re-rendered. These are real before/after frames from that loop:

Problem 1 (5–8 s): the typed query was too small to read on a phone. Fix: a camera push-in to 2.07× while typing, released as the results page arrives.

Before and after: the search scene at original size versus pushed in to 2.07x so the typed query reads on mobile

Problem 2 (8.5 s): the wordmark slid through the tabs during the header morph. Fix: the logo leads, the bar trails by 0.2 s, and the camera pulls back earlier.

Before and after: the wordmark overlapping the tab bar mid-transition versus a clean staggered morph

Problem 3 (any frame after 4 s): a stray white line from Scene 01 leaked into later frames. An early return skipped resetting one element, so a frame's look depended on the frame before it. Fix: set every property on every frame, plus an automated determinism test.

Before and after: a leaked white line crossing the layout versus the fixed frame

It also caught that one Images-tab card had a broken favicon on the live site and had been filled with a grey placeholder dot. The brief said no placeholders, so it swapped in the next real card from the same grid and re-rendered only that section.

Step 10: Final QC

Before delivery, Claude verified the duration (30.00 s), the exact frame count (1,800), resolution, fps, BT.709 colour tags (it found missing primaries and fixed them with a lossless remux), loudness and true peak. It also extracted both sides of every scene boundary to check for blank or broken frames.

Contact sheet of the final film with a frame every half second from 0 to 30 seconds


The one skill: reuse this workflow with Claude

I packaged the whole method as a Claude skill: SKILL.md (product-film-motion-graphics). It encodes the rules that made this work:

  • research the live product first, and use real data only
  • render(t) must be a pure function of time, with no early returns
  • one timeline.js for picture and sound
  • three camera states, four easing curves, and mask-based type
  • identity-preserving transitions, with offset morph windows so elements never collide
  • a synthesised soundtrack mastered to −14 LUFS / −1 dBTP
  • preview, critique (three worst problems with timestamps), fix, then the 4K render
  • determinism test and final QC numbers

To install it in Claude Code: copy the folder to .claude/skills/product-film-motion-graphics/SKILL.md in your project (or ~/.claude/skills/ for every project). Claude loads it automatically when you ask for a product film.

A prompt you can paste:

Make a finished 30-second product film for https://yourproduct.com. Use the product-film-motion-graphics skill. Research the live site first and use only real UI and real data, with no invented metrics. 3840×2160, 60 fps, H.264 + AAC, original synthesised soundtrack at −14 LUFS. Six scenes: problem → search/entry point → hero feature → secondary features → founder/customer value → brand end card. Show me keyframes and a low-res preview, critique the three worst problems with timestamps, fix them, then render the final and report QC numbers.


Limits

  • It's code-driven motion design, not generative footage. You won't get photoreal humans or live action. You get precise, brand-accurate interface storytelling, which is what most SaaS launch videos need.
  • Claude can't listen. Audio is checked with meters and spectrograms; you should still listen before publishing.
  • 4K60 encoding is slow on a laptop CPU. Capture took about 6.5 minutes; the H.264 encode took roughly 20 more. Plan for it, or encode a 1080p cut for social.
  • The live data is a snapshot. Search results change; the film shows them as an illustration of a search, not a guaranteed ranking. Rotating sponsored slots were deliberately left out.
  • Quality comes from the loop, not the first pass. The first cut had real problems. The critique stage is not optional.

FAQ

Can Claude Opus 5.5 make videos?

Yes, but not by generating pixels. In Claude Code it writes a program, here an HTML/CSS/JavaScript renderer, then captures frames with a headless browser and encodes them with ffmpeg into a real MP4 with audio.

Is Claude a text-to-video model like Sora or Veo?

No. Text-to-video models synthesise footage directly. Claude writes and runs code that renders footage deterministically, so the output is editable, reproducible frame by frame, and built from your real product UI and data.

What tools do you need to make a motion graphics video with Claude?

Claude Code, Node.js, Playwright (headless Chromium) and ffmpeg with libx264. Python with Pillow is optional, for contact sheets. No After Effects or video editor is required.

How long does a 30-second 4K video take to render?

In this project, capturing 1,800 frames at 3840×2160 took 385 seconds with five parallel browser pages (about 214 ms per frame). H.264 encoding at 4K on a laptop CPU took roughly another 20 minutes. A 540p preview renders in about a minute.

How does Claude keep audio in sync with the animation?

Both the animation and the audio synthesiser import the same timeline.js. Keystrokes, clicks and transitions are generated from the same timestamps that drive the visuals, so sync is exact by construction: 84 cues in this film.

Can Claude create original music for a video?

It can compose and synthesise music in code: oscillators, envelopes, filters, FM synthesis and reverb, then master it to a loudness target. It can't hear the result, so a human should listen before publishing.

Does the AI make up fake product screenshots?

Not with this workflow. The skill requires researching the live product first and using only real interface details and data. The brief explicitly banned invented metrics, testimonials and rankings.

What is a Claude skill?

A skill is a folder with a SKILL.md file that packages instructions for a specific kind of task. Claude Code loads it when a request matches its description, so a workflow like this one becomes repeatable.

Why render with HTML instead of a video tool?

HTML/CSS can reproduce a web product's real interface exactly, fonts and tokens included. Making every property a function of time turns the browser into a deterministic, scriptable motion graphics engine.

What resolution and format should a launch video be for X and LinkedIn?

Master in 3840×2160 at 60 fps, then post a 1920×1080 H.264 MP4 with AAC audio and +faststart. The 1080p cut of this film is 9.7 MB. Design it to work muted, since many people watch social video without sound.


Resources

SaaSearch.io is an independent directory, not affiliated with, endorsed by, or connected to Google LLC. Product names and images shown in the film belong to their respective owners and appear as they do in SaaSearch results.


About the author: Billal is a solo founder and the builder of SaaSearch.io. He writes about shipping indie products and using AI agents to do work that used to need a studio. Follow him on X at @billalb4u.

Billal
Billal
Founder, SaaSearch.io · 1h ago

Discussion & Reader Feedback0

Share your teardown analysis, feedback, or discuss architecture nuances with the community.

Leave a Thought or Architecture Feedback
Posting as:
Constructive technical commentary is appreciated by makers and readers.
No comments yet. Be the first to share your thoughts!