block #0001 in --:--:--Join the pool
PENDING…
ai ✓ confirmed 6/6 1h ago · 7 min read

Grok 4.7 ships at the old price and a split scoreboard

Grok 4.7 is SpaceXAI's Sept. 21, 2026 coding model, priced like Grok 4.6 and still behind the lab it most wants to beat on several of its own charts. Decrypt and the launch post do not describe the same scoreboard.

Grok 4.7 ships at the old price and a split scoreboard
in this block
  1. What actually happened
  2. Where the scoreboards split
  3. What the cheap row is actually for
  4. What to do as a reader (not a trade)

Grok 4.7 is the model SpaceXAI wants you to pick when the frontier labs feel expensive and slow. The launch post calls it the company's most capable model for coding and knowledge work, served at the same price and speed as Grok 4.6, and "highly competitive in its class." This is not investment advice. A second page, Decrypt's Sept. 21 write-up, agrees on the date and the sticker, then spends its energy on the gaps. A third page, GitHub's changelog the same day, only says where the model is starting to show up inside Copilot. Keep those jobs separate.

TL;DR - The launch post prices the model from $2 per million input tokens and $6 per million output tokens, the same menu it lists for Grok 4.6, plus a fast variant at twice the output speed and twice the price. - On the company's own table, CursorBench 4.0 is 46.3% versus 51.8% for Fable 5.1 Max, and Terminal-Bench 4.0 is 37.6% versus 57.9% for that same Fable column. - Decrypt adds claims the launch post, as opened here, does not: a 2.1 trillion parameter count, a GDPval score of 1695, and immediate availability in the Grok app.

What actually happened

SpaceXAI says the model uses a new, larger base than Grok 4.6, with a longer reinforcement-learning run weighted toward problems that take many hours. It says the model is better at checking its own work and at holding longer context, and that it was trained to understand the Grok Bot harness. The safety section says the safeguard stack is new, that the model is the strongest the company has tested on refusals and jailbreaks, and that select cybersecurity partners have invite-only access to red-team features. LatchBio's biosafety benchmark is listed at 62.4%. On HackerBench v0.3, the post says only 3.3% of risky dual-use prompts get through, while legitimate security work is rarely blocked.

The price block is the cleanest part. Input is $2 per million tokens and output is $6, matching the Grok 4.6 column in the same table. GPT-5.6 Sol Max is listed at $4 and $20. Fable 5.1 Max is listed at $10 and $50. A fast variant doubles output speed and doubles price. The headline above the post also says "twice as fast, at half the price of comparable models." That line is about other labs. The body line is about Grok 4.6. They are not the same comparison, and the post does not show the arithmetic that turns "half the price of comparable models" into a single rival's invoice.

GitHub's changelog, also Sept. 21, says Grok 4.7 is rolling out in Copilot for Pro, Pro+, Max, Business, and Enterprise. The picker list covers Visual Studio Code, Visual Studio, the Copilot CLI, the cloud agent, the Copilot app, JetBrains, Xcode, and Eclipse. Billing is provider list pricing under usage-based billing. Administrators on Business and Enterprise can allow or block it in model policy. The rollout is gradual. If the picker is empty, GitHub's instruction is to check back, not to assume a bug in your account.

Decrypt says the release landed Monday afternoon after Elon Musk walked the date back at least five times from late July, including "10 days" on Sept. 1 and "needs a few more days to cook" on Sept. 11. It quotes the company calling the model "a notable improvement over Grok 4.6 at the same price and speed," and quotes Musk calling it "a strong combination of intelligence, speed & low cost." It says there is no waitlist, and that the model is live in the Grok app, Cursor, Grok Build, and the xAI API. The launch post's availability paragraph, as retrieved, names Grok 4.7 in Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers and cloud platforms. It does not, in that paragraph, say "Grok app." If you need the app specifically, you are relying on Decrypt, not on the sentence in the post.

Where the scoreboards split

The launch table is a comparison against Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. CursorBench 4.0 reads 46.3%, 40.4%, 41.7%, and 51.8%. DeepSWE v1.1 reads 71.0% with a star the post marks as high effort, then 65.2%, 72.7%, and 70.0%. EEBench, the electrical-engineering set, reads 64.0%, 53.0%, 39.4%, and 56.4%. AA Briefcase v1.1, multi-hour office work, reads 1,657, 1,546, 1,487, and 1,678. Terminal-Bench 4.0 reads 37.6%, 20.3%, 37.3%, and 57.9%. Harvey's legal-agent benchmark reads 19.6%, 15.8%, 2.5%, and 6.7%. HealthBench Professional reads 56.7%, 48.5%, 60.5%, and 62.1%.

Read the columns before you pick a moral. Against its own previous model, the Grok 4.7 column is higher on every row in that table. Against Fable 5.1 Max, it is lower on CursorBench, Terminal-Bench, AA Briefcase, and HealthBench, higher on DeepSWE and EEBench and Harvey. Against GPT-5.6 Sol Max, it is higher on CursorBench, EEBench, Briefcase, Terminal-Bench by a hair (37.6 versus 37.3), and Harvey, and lower on DeepSWE and HealthBench. "Frontier" in the post's CursorBench sentence means price-performance, not first place.

Decrypt's scoreboard is not a photocopy. It says Grok 4.7 hit 1695 on GDPval, with Claude Fable 5.1 at 1735. The launch post discusses GDPval in prose and does not, in the table opened here, print 1695. Decrypt's AA Briefcase figures, 1657 against Fable's 1678, match the launch table's 1,657 and 1,678. Same test, same pair, written with or without a comma. Decrypt also says the model packs 2.1 trillion parameters, up 40% from 1.5 trillion in Grok 4.6, and that supplemental training included SpaceX material: Starlink telemetry, manufacturing records, and engineering failure logs. None of those parameter or dataset lines appear in the launch post as it was opened. Attribute them to Decrypt or leave them out.

Decrypt also says Musk had already lowered expectations, writing that the model should land "roughly on par with" Claude Opus 5.0, not a newer Opus, and that multimodal performance still needed work. The same account sketches Grok 4.8 as a step up, Grok 4.9 in what Musk called "Astra/Fable class," and Grok 5 as a possible later leader, with no dates. "We shall see" is the quote Decrypt uses. Those future names are a tease, not a ship schedule.

What the cheap row is actually for

A low token price is a strategy only if the Grok 4.7 task finishes. Decrypt's frame is that xAI keeps undercutting Anthropic and OpenAI per token while trailing them on raw capability, and that "good enough, cheap, and everywhere" can beat "best, but pricier" for everyday use. The launch post's own HealthBench and Terminal-Bench rows are the places that argument is weakest. The EEBench and Harvey rows are where it is strongest. Quoting one row as the model is how a thread lies without inventing a digit.

Other September launches are context, not a cage match you can settle from this table. Gemini 4 Argon is a different lab's flagship story. Cohere Embed 5 is a retrieval model, which this coding table does not measure. If your job is search over private documents, a CursorBench percent will not help you. If your job is a long coding agent, an embedding launch will not either.

What to do as a reader (not a trade)

Match the sentence to the page. Grok 4.7 at $2 and $6 is in the launch table, next to an identical Grok 4.6 price, and GitHub says Copilot bills provider list pricing. The fast variant is twice the price for twice the output speed. Do not describe that variant as the cheap one.

If you repeat a benchmark, name the column. "Beats Fable" is false for Terminal-Bench 4.0 in the company's own table and true for EEBench in that same table. If you repeat 1695 on GDPval, or 2.1 trillion parameters, say Decrypt said it. If you repeat 1,657 on AA Briefcase, both pages you can treat as opened here agree,.

If the model is missing in Copilot, GitHub already said the rollout is gradual and that admins can disable it. That is a settings check first. And if a post tells you the next three Grok versions will take the crown, notice that Decrypt's source for that sketch is a Musk post, not a model card with a date. The shipped object is the one with a price and a table. Everything after Grok 4.7, in these sources, is still a promise.

Not financial advice. DYOR, ser.

More in the pool

all ai
gm ser

Get confirmed before the crowd

Daily block at 07:00 UTC. No spam, just the block, ser.