Blog Article

Claude Opus 5.5 Effort Levels Compared: Same Prompt, Six Results

Six Claude Opus 5.5 effort levels, one identical prompt, one attempt to clone Sonic the Hedgehog in 3D. Timings, bugs, and why the top tier was not the pick.

I do not think anyone has tested all of the Claude Opus 5.5 effort levels side by side. So I did. Same prompt, six sessions, one goal: clone my all-time favorite nostalgic childhood video game, Sonic the Hedgehog, in 3D with a first-person racing view.

This is not a lab benchmark. It is a build-in-public test with real timings, real bugs, and one metric I regret not capturing. Here is exactly what happened.

The prompt every session got

Fairness was the whole point, so every effort level received this exact prompt:

> Create a new folder relevant to the effort level this session is configured for. Then inside, build me a 1:1 clone of Sonic the Hedgehog. Include a new view that allows us to toggle from its traditional side view to a first-person over-the-shoulder view where we're actually racing the course. I want you to get as creative as possible and let's make it 3D and ultra realistic.

Each build ran on its own localhost port, 9000 through 9006, labeled one through six in my browser tabs. That let me flip between them and compare the results directly.

How long each effort level took

Before we get to gameplay, here is the wall-clock time for each run. Keep in mind I was doing other work on the side during the longer runs, so these numbers are not clean lab measurements.

| Effort level | Time to complete | |---|---| | Low | Under 2 minutes | | Medium | Just under 18 minutes | | High | The better part of 30 minutes | | XHigh | 1 hour and 1 minute | | Max | About 1 hour and 53 minutes, interrupted by a usage limit | | Ultra Code | Force-stopped |

Two things stand out. First, Low finishing in under two minutes was kind of wild. Anthropic models have been notoriously slow, so I expected every tier to take a significant amount of time. Second, the Max run consumed close to two hours and then hit a usage limit. I had to wait for the reset, after which it ran for another two minutes and finished. My side projects contributed to that limit, so these tests did not cause it alone. Sadly, I will never know whether they would have all completed inside a single five-hour session.

Ultra Code was a mistake. After the reset, same exact prompt, it immediately started spinning up sub agent after sub agent. Seven were running at one point. It only completed because I forced the stop. I honestly was not going to include it in the video, but I did.

Low effort: playable in two minutes

The low-effort clone had no sound. Pressing down triggered a spin dash, and the C key toggled from side view to the racing view.

Here is the first bug. When I switched to the racing view and pressed forward, nothing changed. I had to keep pressing the controls as if I was still looking at the side view. That is kind of annoying. Climbing the first hill was painfully slow, and there was no endgame boss.

Still, for two minutes of work, not bad. Super nostalgic, and it was playable.

Medium effort: the surprise of the test

Medium added sound right away. There was a weird glitch on load, but it resolved itself after a hard refresh.

The important difference: when I switched into the racing view, the controls switched with it. I could go left and right. Some of my other tests with other models never allowed that. The ocean had nice shaders, and the level of detail was genuinely impressive for a mid-tier run. Down plus space let me power through everything.

Eighteen minutes for this result. Not bad. Honestly, wow.

High effort: more views, but an armadillo

High added a forward control that Low missed, and it offered a lot more camera options. Where Medium gave two views, High let me toggle into a true first-person view.

It also gave me an armadillo instead of Sonic. This is where the IP protection started showing up, and it kept showing up from here on. Still no endgame boss.

XHigh effort: shocking regression

XHigh rendered the same armadillo and had the same weird startup bug as Medium. Then the real surprise: the controls did not change when I switched views. I could not go left or right in the racing view.

That is shocking coming from an extra high output. An hour of reasoning, and it missed a detail that Medium handled in 18 minutes. Funny looking character too. Not bad overall, but not the step up I expected.

Max effort: the best build

Max was obviously the best one yet. The first-person side view was very cool, and the world it created was impressive. I love its recreation of Sonic in this build.

It was not perfect. One bridge was not complete, so I could not go back at one point. And to be clear, none of the six tiers produced a true one-to-one clone. Copyright and IP protection guardrails mean a "clone X" prompt lands as "inspired by X" every time.

Ultra Code: gorgeous, laggy, and one obvious miss

Ultra Code took a while to load, which concerned me. It had the best splash screen of the bunch and a stunning background with water fountains off to the left. Another armadillo.

There was noticeable lag. And the miss that bothered me most: after switching views, it did not remap W to move forward. You would think Ultra Code would nail that. It did not. The one thing I did like was being able to run backwards through the course.

The Ultra Code environment is undeniably pretty sick. The output on this task did not justify seven sub agents and a forced stop.

What I actually learned

Lower effort is faster than you think. I expected every tier to grind. Low and Medium were done before I finished setting up the tabs for the others.

Higher effort clearly produces a better clone. The worlds, the detail, and the camera options improved as effort climbed. Max was the most complete build.

More effort does not mean fewer bugs. The control-remapping miss appeared in Low, XHigh, and Ultra Code. Medium and High got it right. Small UX details broke at both the bottom and the top of the ladder.

The top tier is not the automatic pick. Honestly, I am liking Max and XHigh a little more than Ultra Code for this kind of task. Ultra Code looked the best but burned the most resources and still missed a basic key binding.

Usage limits are a real constraint. A near two-hour Max run plus side work pushed me into a limit. If you are running high-effort jobs, do it when nothing else is competing for your session window.

The one thing I regret

I did not capture cost in equivalent API tokens. Time tells part of the story, but without cost per tier the comparison is incomplete. If you run this test yourself, log both.

How to run this yourself

The setup was simple, and you can copy it for your own use case:

  • Write one prompt and do not change a word of it between sessions.
  • Ask each session to create its own folder named for its effort level.
  • Serve each build on its own localhost port so you can compare them side by side.
  • Record wall-clock time and, unlike me, token cost.
  • Play-test the small stuff: controls, sound, view toggles. That is where the tiers diverged.

I am not an AI expert. I am building in public and sharing what actually works. If you want a video on a specific topic, drop a comment on the video with who you are, what you do, who your target audience is, and what your question is. I will add it to my queue.

Watch the full side-by-side test here: https://www.youtube.com/watch?v=ufqJZuh6Y48

Watch the full walkthrough on YouTube.

Watch on YouTube