Blog Article

Is Grok 4.7 Good for Coding? Three Hours of Live Testing Say Yes for Backend, No for Front End

First-impression results from three hours of live testing with Grok 4.7: pricing, a working NFL data app, a browser-OS benchmark where Grok 4.6 beat 4.7, and why the

Grok 4.7 dropped about ten minutes before I went live. I had never prompted it before the stream started, so everything below is a first impression from roughly three hours of building with it in real time.

Short version: the backend and research work held up, the pricing is the real headline, and the front end fell short on every single test. The most surprising result of the day was that the older Grok 4.6 beat 4.7 on my signature benchmark.

Here is what I ran, what came back, and where I landed.

The pitch and the price

xAI's announcement called Grok 4.7 its most powerful model for coding and knowledge work, twice as fast and half the price of comparable models. Their chart put Grok 4.7 on extra high effort almost equivalent to Fable 5.

The number that actually got me was pricing.

| Model | Input (per 1M tokens) | Output (per 1M tokens) | |---|---|---| | Grok 4.7 | $2 | $6 | | Comparable frontier model A | $10 | $50 | | Comparable frontier model B | $4 | $20 |

They kept it at the same price as 4.6. I love that. I keep hearing creators on X say to enjoy the cheap prices while they last, and that has always confused me. My belief is that with more time, more progress, and more competition, things only get cheaper. I am planting my flag on that. Elon entering at a fraction of the cost is exactly the pressure that keeps prices down.

One usage note worth sharing. After nearly three hours of parallel builds on a SuperGrok Heavy plan, the usage meter read 6%. I was honestly shocked it was not higher.

How I tested it

I wanted to hit Grok 4.7 from three entry points: the Grok Build CLI, T3 Code, and grok.com. That plan lasted about five minutes. On launch day, grok.com chat was still running 4.6 with no way to toggle to 4.7. So the CLI and T3 Code did all the work.

Quick note on T3 Code. I installed it early when Theo first started promoting it and it was not polished enough for me. Over the weekend I gave it some real Clearmud projects and it is growing on me. Switching models per thread, running off existing subscriptions with no API keys, pinning threads, and an in-app browser all made the parallel testing easier. It still needs side-by-side threads and env file cloning.

My approach to prompting new models is simple prompts first. No skills, no PRD, no design system. If you hand a model an insanely detailed PRD, it does not matter much which model you use. Most of them will produce a similar result because the prompt is so defined. A basic prompt forces the model to pick the stack and make the design calls, which is what actually reveals its reasoning.

Test 1: The Pro Football Reference app

This is the build I cared about most. Pro Football Reference has pretty much all the data most people pay for, for free, but you have to know how to navigate it. Every fantasy season I do that manually. Last year I noticed a rookie whose hometown was near Seattle was about to play his first road game at the Seahawks' stadium. I stacked him with Baker Mayfield and JSN and won a lot of money. That research was all clicking and Googling.

So I asked Grok 4.7 to build a filterable web app on top of that data, with the plan to add Jev by Typesafe AI on the backend to classify the questions I ask it. Things like, what is this player's record when he visits his hometown team's stadium.

Grok pushed back that the site's terms prohibit automated collection, then built a working v1 anyway. I ran a hometown stadium question and it came back with specific games. I went and verified the game logs by hand. It was correct.

I am impressed with the fact that it works. Let's put that up top. It works.

Then I looked at the UI. This felt like a website built in the early 2000s, not a website built in 2026. I gave it more direction after that: shadcnblocks, glassmorphism, a lime-green brand color, DraftKings' dark palette. It started to come together, but it never got there on its own.

Once I added the Typesafe AI API key, Jev correctly classified hometown stadium questions and routed anything else as "something else." That is both the value and the tax. Jev only classifies what you configure, so you have to build the catalog of question types yourself.

For comparison I sent the same prompt to Astra. Astra built a much better UI and even pulled in nflverse on its own, but the data link did not work. Grok's version is the one I am keeping.

Test 2: The Clearmud browser-OS benchmark

Every new model gets this prompt: analyze our website and rebuild the brand's public site as an interactive operating system that runs entirely in the browser. It is inspired by PostHog's site, which is amazing from a techie perspective. Every prior model I have tried gave me something like macOS or Linux in a browser tab with a magnetic dock.

Grok 4.7 got the brand color and nothing else. Compared to what Astra and Fable did with the same prompt, I wanted to puke. I did not even dig into it.

Then Mike in chat suggested running the same prompt on Grok 4.6. I made a new directory, switched the model, pasted the prompt, and waited.

Grok 4.6 produced an actual OS in the browser. A dock, widgets, snappable windows, and a terminal that answers "who am I." Which one looks more like an OS in the browser? 4.6 did the better job. The older model beat the newer one. What world do we live in where that happens?

Test 3: The HVAC landing page

A realistic use case: find an outdated HVAC company website in Las Vegas and rebuild it for 2026 and beyond, making every creative and stack decision yourself.

The result was a single-page scroller that looked worse than the original WordPress site. Front-end design is not a responsibility I would give to Grok 4.7 right now.

Test 4: Sonic the Hedgehog, twice

The vibe-prompted version was playable and had a working over-the-shoulder view toggle. The spin animation was anchored wrong and the end boss was a savage.

Then I tried the PRD-first approach. I had Grok 4.7 write its own PRD in plan mode, opened a fresh session, and built from that. I deliberately did not let Fable or Astra write the plan because that is kind of cheating. The PRD version fixed the spin but made the level confusing, with a breaking bridge and unexplained purple circles. It felt kind of worse. Structure alone did not fix the weak spot.

Test 5: Recoil, Grokbot office, and the strengths test

A viewer asked for a clone of Recoil, the 1999 tank game. First pass was drivable with weapon perks and world edges, but low detail. Asking for Three.js shaders, normal maps, and PBR improved the world and left the tank blocky.

I also had it generate a Three.js office visualizing Grokbot agents, inspired by a Grokbot team member who built an island for hers. Fun experiment, robots with backwards arms, not pursued.

The one that changed my read came late. Matt Palmer from xAI posted that Grok 4.7 is especially good at software engineering, long-running knowledge work, electrical engineering, and legal tasks. Nobody mentioned front end. That matched what I was seeing. So I asked Grok to read the thread and design its own showcase. It built a fictional legal and QC case around a kiln feeder at a ceramics company in Palo Alto with a sticky checklist UI. I could not fully judge the legal content, but that became my favorite test of the day.

The one-shot myth

One thing came up in chat that I want to address directly. When somebody says they one-shotted an app, the amount of context in that prompt is insane. They are not writing one sentence. They are an expert at their craft feeding the model everything it needs.

I did not prompt for a one shot. If I wanted a one shot, I would have prepped a PRD first. So judge the model on the prompt you actually gave it, and give it the benefit of the doubt before blaming it.

Where I landed

After almost three hours, here is my honest read.

  • Backend, research, and functional builds: it works, and it is cheap. The NFL app pulled real data correctly on the first pass.
  • Front end and visual design: it fell short on every test. Sure, things work, so maybe you cannot call it a failure, but I would not use Grok 4.7 for anything front end. Period.
  • Versus 4.6: on my browser-OS benchmark, 4.6 won. Run your old model on the same prompt before you assume newer is better.
  • Versus Astra and Fable: I do not see it as something I would reach for over them yet as a daily driver.

It has its strengths. It clearly has weaknesses from a UI perspective. Where I want to use it is from an agent perspective. I run Astra interactively but route routines and automated agent work to a cheaper model, and Grok 4.7 is the next candidate for that slot.

Out of everything we built, the Pro Football Reference codebase is the only one I am moving forward with. I am dumping it into another project as a feature.

Caveats

These are launch-day impressions. My tests leaned heavily visual, and the legal, electrical engineering, and long-running knowledge-work strengths only got one prompt. Some slowness and errors late in the stream could have been T3 Code or launch-day load rather than the model.

You experienced this with me. I had not used 4.7 before the stream. Watch the full session here:

https://www.youtube.com/watch?v=bTR6R3E-7XM

Watch the full walkthrough on YouTube.

Watch on YouTube