Blog Article

Which AI Model Is Best for Motion Graphics and Ads? I Blind Tested 7 of Them

I ran seven GPT and Claude models through three blind motion graphics tests and tracked time, tokens, and retail API cost. Claude won on brand accuracy, and Sonnet

I built a web app for Clearmud so I can create my own animations and graphics without using generative AI video. Then I figured, why not turn it into a benchmark?

So I ran the same prompts across seven GPT and Claude models, in a blind test, and tracked how long each one took, how many output tokens it spent, and what it would cost at retail API pricing.

Short version: the Claude models won on brand accuracy and quality, and Sonnet 5.5 came out as the budget pick. I did not see that coming.

The setup

Here are the seven models I tested:

  • Claude: Fable 5.1, Opus 5.5, Sonnet 5.5
  • GPT: Astra, GPT 6.1 Soul, Luna, Terra

I left Haiku out on purpose. For this kind of work it's just not there right now.

Every test was blind. Each output showed up as Model A through G, and I only revealed the names after picking my favorites. I'm on a subscription, but the benchmark panel calculates cost as if I were paying retail API prices, so you can compare.

I ran three tests:

  • A 15-second vertical social clip
  • A 30-second product launch demo
  • A 30-second ad promo

One more thing. I did not upload brand assets for any of these. I know that when you make a motion graphic for a brand, you'd normally attach the logo, colors, and fonts. I wanted to test each model's ability to reason, search, scrape, and then create.

Test 1: a 15-second clip for OpenClaw

The prompt was a few sentences. Create a viral USP style video for OpenClaw, and use the branding, icons, and information from their website and socials. That's it. A beginner prompt.

All seven models finished in about 6 minutes. The cheapest run cost 2 cents.

Then I watched them one by one, and the first big lesson showed up fast [03:38]. One output didn't pull the current logo. That's a failure in my eyes before the animation even starts. Another one missed the brand color completely.

After the reveal, the pattern was clear [07:50]:

  • All three Claude models used the most up-to-date OpenClaw logo, branding, and information
  • Astra was the only GPT model that pulled the latest icon and logo
  • Luna, Terra, and Soul all pulled the old one

My favorite was Opus 5.5, with Fable 5.1 as the runner-up. And my guesses about which model was which? Wrong. That's why you run it blind.

On cost, Fable and Astra were the two most expensive. Opus came in third at 81 cents. Sonnet cost 49 cents.

Sonnet's output had great mockups and visuals, it just moved through the slides too quickly. For roughly 60% of the price of Opus, that's not bad at all. With a shot list and more creative direction, I think it would have done a far better job.

Test 2: a 30-second product launch demo

For the second test, I asked each model to create a product launch demo for OpenAI's newly released personal AI agent. No attachments. I told the models to research the release features, the USPs, and the brand and creative assets, and to do a better job than OpenAI did with their own promo.

Again, a beginner prompt. I don't have a creative background. I'm not a director. I'm not a producer. This is how I would naturally explain things, without first asking a model to prep me a shot list.

The longest run took close to 11 minutes.

The results were all over the place. One felt like a corporate PowerPoint. One had the best flow and pacing of the bunch, but I wasn't sure the details were accurate. Two were weak enough that I didn't bother finishing one of them.

My two favorites were Model E, my favorite by far, and Model F, which had this paintbrush effect to it. The reveal [15:21]: Sonnet and Opus.

Here's what the 30-second clip cost for the models I called out:

| Model | Retail API cost | |-------|-----------------| | Fable 5.1 | $2.48 | | Opus 5.5 | $1.84 | | Astra | $1.21 | | GPT 6.1 Soul | $0.68 |

Soul was cheap and arguably one of my favorites. I just wish it were more accurate.

Why I cut Terra and Luna

After two rounds, Terra and Luna had a pattern. They always finished first, and they always had the worst quality.

Let me be clear, this is not me talking negatively about them in general. Luna is great on Macs for day-to-day tasks. For motion graphics specifically, Terra and Luna just don't have it, so I dropped them from the final round. I kept Astra and Soul so at least two GPT models stayed in.

Test 3: a 30-second ad promo with a master prompt

The first two tests were vibed. For the last one, I used the master prompt feature in my app. It generates a production-grade prompt with a script and a shot list, and I can toggle which model preps it. I picked Opus.

The ad was for Puddle, something I'm teasing for Clearmud. The prompt walked through three interactions, including "Who paid for dinner again?" and "What skill MDs do I have system-wide?", with three callouts:

  • All your skills in one location
  • Easy to use agents
  • Never forget anything

I still didn't upload a brand kit. No typography, no color schemes, no logos. I wanted to see what each model does with the gap.

Five models ran. The longest took close to 13 minutes.

And here's what I noticed [21:38]. The outputs looked a lot alike. Similar flow, similar movement, similar scene transitions. It works like a PRD, a product requirements document. Give different models the same detailed PRD and you get similar builds, with slight variations unique to each model. Same thing happens with a detailed shot list.

The differences were in the motion. One output was too basic and cut off some text. One had lots of movement but glitched partway through. One did a phenomenal job with the screen moving while everything else animated around it.

The reveal: Sonnet 5.5

Final round costs:

  • Fable 5.1: $2.34, the most expensive
  • Astra: $0.86, third
  • Sonnet 5.5: $0.75, second cheapest

I went back and watched Sonnet's output again, and it was my favorite. Great movement, and the movement matched what was being explained. There's a stretch in the middle that felt a bit stale, and that's where I'd vibe in improvements.

Sonnet seems to be this undercover motion graphics expert. Who would have thought?

Opus 5.5 and Fable 5.1 did a great job in a lot of scenarios too. Opus also took the longest in the final round. If you're trying to do this on a budget, though, Sonnet 5.5 is the one I'd reach for.

What I took away from this

  • Check the brand first. If a model pulls the wrong logo, nothing else matters. Every Claude model got OpenClaw right. Only one of four GPT models did.
  • Fastest and cheapest was the worst here. Terra and Luna finished first every time and lost every time.
  • Detail makes outputs converge. The more specific you get, the more consistent the output. That's also when you start to see which model is best.
  • Test blind. I guessed wrong in the first round.

What's next for the app

The next layer I'm adding is the ability to go in and vibe individual scenes, plus audio, music, and a text-to-voice layer, so I can record my own voiceover or let AI handle it.

I see a lot of use cases here for digital marketing agencies and anybody who manages and creates ads. This could take a lot of the heavy lifting off your plate so you can focus on ideation.

I've only tested three formats so far. If there's one you want me to try, drop a comment on the video and tell me who you are, what you do, who your target audience is, and what your question is. I'll add it to my queue.

Watch the full test, including every output side by side: https://www.youtube.com/watch?v=fle-lW0csUs

Watch the full walkthrough on YouTube.

Watch on YouTube