Blog Article
GPT 5.6 Sol Built a Website Factory in 79 Minutes
I tested GPT 5.6 Sol on a Sonic clone and an AI website factory. One failed, while the other researched, built, validated, and deployed a real workflow.
I gave GPT 5.6 Sol two very different tests. First, I asked it to recreate the feel of the original Sonic the Hedgehog with modern web technology. Then I asked it to turn a loose business idea into an automated website factory.
The Sonic result was underwhelming. The website factory exceeded my expectations.
That contrast taught me more than a clean benchmark ever could. A model can struggle with camera controls and game physics, then show impressive judgment across research, lead validation, design, deployment, and approval boundaries.
The Sonic Test Was Detailed, but Still Missed
I started with a fairly demanding Sonic-style platformer prompt. The build had to run in a single HTML file with Three.js. I asked for a 3D environment, character controls, enemy AI, collision logic, and distinct movement states.
Physics mattered most. I explicitly requested acceleration, friction, inertia, and sliding before stopping. Most quick clones move a character at one constant speed, so this was a useful test of real-time logic and state management.
The model produced a working prototype, but the navigation felt strange. Left and right appeared reversed, the camera moved away from the classic side view, and the character behaved unpredictably. Adding the original 1991 game as a reference improved it, but not enough.
I also tried a much simpler prompt. The graphics changed, but the experience still wasn't impressive. It was a fun nostalgic experiment, not a result I would ship.
Then I Described a Website Factory
The second test began without a master prompt or prepared product requirements document. I wanted a beginner-first experiment. Most people with an early business idea won't stop to write a technical specification before they know whether the idea is worth pursuing.
I described the goal in plain language: find small or midsize businesses with missing or outdated websites, research each company, build a stronger site, prepare outreach, and let me take over when a lead becomes qualified.
We narrowed the first market to Las Vegas HVAC companies. Each prospect would receive a complete website concept in a sandboxed preview. The business could pay a flat website fee, with deployment handled separately if needed.
Initial outreach would stay behind human approval. That boundary matters because a working automation is not the same thing as a trustworthy production system.
What the Website Factory Actually Did
After a short round of questions, GPT 5.6 Sol created a seven-phase workflow and got to work. I stopped recording, hid the desktop, and recorded another video. One hour and 19 minutes later, it had completed an unexpected amount of work.
The workflow included:
- Defining the business model and target market
- Finding primary and backup prospects
- Checking Nevada's official contractor records
- Inspecting existing websites visually
- Building a brand-specific replacement site
- Creating an internal operations dashboard
- Deploying resources through Vercel and Convex
The contractor check was the strongest part. It found candidates, checked whether their licenses were active, and avoided treating an inactive company as a worthwhile lead. I didn't ask it to do that. It understood that verification needed to happen before spending time on a custom website.
That is the kind of judgment I want from an agent. Writing code is useful. Knowing which work should happen first is more valuable.
Autonomy Still Needs Approval Gates
The system eventually asked for Gmail access so it could prepare email automation. I denied the request.
If I continue building this website factory, it will use a separate sandboxed Google Workspace. I don't want an experimental agent working inside my personal or primary business email. Credentials, outreach, pricing, and production launch all need clear boundaries.
I would also test every site locally before approving a deployment. The first run exposed Vercel preview issues, pop-up behavior, and demo credentials that would need to change. The work was impressive, but it was still a first draft.
A safe version of this workflow looks like this:
- Let the agent research and rank prospects.
- Require official-source validation.
- Build in a sandbox with separate credentials.
- Review the website, claims, and contact details.
- Approve outreach one message at a time during the MVP.
- Expand autonomy only after repeated, verified results.
What This GPT 5.6 Sol Test Really Showed
The biggest lesson wasn't that GPT 5.6 Sol can build a website. Many models can generate a decent landing page.
The real result was sustained work across tools. It planned, researched, verified, built, and deployed without asking for constant direction. Compared with configuring several separate agents, schedules, and tool permissions, the Codex workflow handled much of the coordination itself.
That doesn't mean the factory is ready to launch. I still need to improve the pricing, outreach copy, component system, security, and quality checks. But I want to move toward one finished website per day, then see whether the process can grow from there.
Watch the full video above to see both tests, including the failed Sonic experiment and the website factory reveal. For more practical AI experiments, subscribe to the Clearmud newsletter and follow the build in public.
What business workflow should I test with GPT 5.6 Sol next?
Watch the full walkthrough on YouTube.
Watch on YouTube