September 30, 2026

xAI Launches Grok 4.7, Betting on Price and Speed Over Benchmark Wins

1

xAI released Grok 4.7 on September 21, a bigger model for coding and knowledge work priced the same as Grok 4.6, but still trailing Claude Fable 5.1 on several benchmarks.

Smartphone screen displaying an AI chatbot interface asking what can I help with, representing conversational AI assistants like Grok

Photo by <a href="https://unsplash.com/@timwitzdam?utm_source=WP+Agent&utm_medium=referral">Tim Witzdam</a> on <a href="https://unsplash.com/?utm_source=WP+Agent&utm_medium=referral">Unsplash</a>

Elon Musk’s xAI released Grok 4.7 on Monday, September 21, ending a rollout that had slipped past its own announced dates at least five times since late July. The company is positioning the model as its strongest yet for coding and professional knowledge work, and it is doing so without raising prices. Independent benchmark comparisons published alongside the launch put Grok 4.7 ahead of its own predecessor on nearly every measure, but still behind Anthropic’s Claude Fable 5.1 on several of the same tests — a pattern that has followed the last few Grok releases.


Smartphone screen displaying an AI chatbot interface asking what can I help with, representing conversational AI assistants like Grok
Photo by Tim Witzdam on Unsplash

What Happened With Elon Musk’s xAI?

xAI, which now operates under the SpaceXAI brand after SpaceX’s acquisition of the company earlier in 2026, published its official announcement for Grok 4.7 on its news page on September 21. The company describes it as “SpaceXAI’s most powerful model for coding and knowledge work,” available immediately with no waitlist in the Grok app, the Cursor coding editor, xAI’s own Grok Build agent, and the xAI API, with GitHub Copilot integration rolling out as well.

The release had been delayed repeatedly. According to reporting from Decrypt, Musk pushed the expected launch date back at least five times: from “four weeks out” in late July, to “3 to 4 weeks” in mid-August, to a roughly 10-day countdown on September 1, to “needs a few more days to cook” on September 11. xAI’s official account posted on X that Grok 4.7 was “a notable improvement over Grok 4.6 at the same price and speed,” and Musk separately described it as “a strong combination of intelligence, speed & low cost.”

What Is New in the Latest Grok Update?

A Larger Base Model, Trained for Longer Tasks

According to xAI’s own release materials, Grok 4.7 uses a larger base model than Grok 4.6, trained with a longer reinforcement-learning run focused on harder tasks that take many hours to complete rather than quick single-turn answers. The company says the model checks its own work more carefully, manages longer context better, and was trained to natively understand the harness behind Grok Bot, xAI’s agent product, which the company says improves its performance on conversational and general knowledge tasks.

Multiple outlets, including Decrypt and Yahoo Tech, report that Musk has described Grok 4.7 as built on roughly 2.1 trillion parameters, up from 1.5 trillion in Grok 4.6 — a 40% increase. It’s worth noting that xAI’s own launch page does not publish a parameter count; that figure comes from Musk’s public statements rather than an official technical specification, so it should be treated as a company claim rather than an independently verified spec.

SpaceX Engineering Data in the Training Mix

xAI also says it added supplemental training data drawn from SpaceX, including Starlink satellite telemetry, manufacturing records, and engineering failure logs, with the stated goal of improving the model’s reasoning about hardware and physical engineering problems — an area general-purpose models trained mostly on internet text have historically struggled with.

The Technology Behind the Update

Pricing is unchanged from Grok 4.6: $2 per million input tokens and $6 per million output tokens, with a faster variant available at twice the output speed for twice the price. xAI also introduced a new safeguard system alongside the model. The company says Grok 4.7 is its strongest version yet on refusing jailbreak attempts, and reports a score of 62.4% on LatchBio’s third-party biosecurity benchmark, along with a 3.3% pass-through rate for risky prompts on HackerBench v0.3, its internal test for dual-use cybersecurity requests. xAI says it has begun giving a small number of cybersecurity partners invite-only access to the model’s red-team capabilities for defense research. These figures come from xAI’s own model card and launch materials; they have not yet been independently reproduced by outside researchers.


Laptop screen displaying colorful programming code, representing coding and software engineering benchmarks used to test AI models
Photo by Mohammad Rahmani on Unsplash

How It Compares With Previous Grok Models

Gains Over Grok 4.6

xAI’s published benchmark table shows gains over Grok 4.6 across the board. On CursorBench 4.0, a Cursor-built test of real coding tasks weighed against cost, Grok 4.7 scored 46.3% at its highest reasoning setting versus 40.4% for Grok 4.6. On EEBench, an electrical-engineering benchmark, it improved to 64.0% from 53.0%. On Terminal-Bench 4.0, a test of multi-hour terminal work, the jump was from 20.3% to 38.0%. On the Harvey Legal Agent Benchmark and HealthBench Professional, gains were smaller but still positive, moving from 15.8% to 19.6% and from 48.5% to 56.7%, respectively.

Where It Stands Against Rivals

Against outside competitors, the picture is more mixed. On GDPval, a benchmark that scores models on real professional work such as legal memos and spreadsheets using an Elo-style rating, Decrypt reported Grok 4.7 at 1695, ahead of Grok 4.6’s 1605 but behind Claude Fable 5.1’s 1735. On AA-Briefcase, a multi-hour office-work benchmark from Artificial Analysis, Grok 4.7 scored 1657 against Fable 5.1’s 1678. It’s a pattern that has held across the last three Grok releases: incremental gains over its own prior version, while still trailing Anthropic’s top model on several head-to-head professional-work tests. Where Grok 4.7 does distinguish itself is price: at $2 input and $6 output per million tokens, it costs a fraction of Fable 5.1’s published $10 input and $50 output rate, making its price-to-performance position considerably stronger than its raw benchmark rank alone would suggest.

Why This Matters for the AI Industry

Grok 4.7 arrives at a moment when several major AI labs are releasing new frontier models within weeks of each other, intensifying competition on both capability and cost. xAI’s approach — consistently pricing below Anthropic and OpenAI even while trailing them on top-line benchmark scores — reflects a strategy built around volume and accessibility rather than claiming the single best model on the market. That strategy carries real weight given Grok’s distribution: the model reaches users not just through the standalone Grok app and X, but through Tesla vehicle dashboards, Cursor’s coding environment, and now GitHub Copilot, giving it a considerably larger default audience than a chatbot app alone would have.

The push toward coding and agentic, multi-hour task completion also reflects where the broader industry has moved this year: benchmark categories like Terminal-Bench and AA-Briefcase did not exist in their current form a year ago, and their emergence signals that vendors are now competing less on single-turn chat quality and more on whether a model can be trusted to work independently on a task for hours without supervision.

What It Means for AI Users and Businesses

For developers already using Grok inside Cursor or Grok Build, the update is available immediately with no action required and no price change, making it a low-friction upgrade to evaluate. For businesses comparing frontier models — a decision covered in more detail in our comparison of ChatGPT, Gemini and Claude — Grok 4.7’s combination of mid-tier benchmark performance and low per-token pricing makes it a reasonable candidate for high-volume, cost-sensitive coding or document-generation workloads, though teams doing legal or clinical work should note that Grok 4.7’s Harvey and HealthBench scores, while improved, remain well behind the leading models on both benchmarks.

xAI’s emphasis on its new safeguard system is also relevant to organizations weighing AI security exposure, a topic that has drawn scrutiny across the industry following recent incidents involving AI-assisted security research. xAI’s claimed improvements on jailbreak resistance and dual-use refusal rates are company-reported figures rather than independently audited results, so businesses with strict compliance requirements should treat them as a starting point for their own evaluation rather than a substitute for it.

What Elon Musk and xAI Are Building Toward

Musk has been explicit that Grok 4.7 was not intended to top the leaderboard. According to Decrypt, he said before launch that the model should land “roughly on par with” Anthropic’s Claude Opus 5.0 — not the newer Opus 5.1 — with multimodal performance still needing further work. In the same post, he reportedly sketched out a roadmap: Grok 4.8 as a meaningful step up, Grok 4.9 aiming for what he called “Astra/Fable class” performance, and Grok 5 as a potential frontier leader, though none of those releases have a confirmed date. This roadmap is Musk’s own stated expectation, not a commitment from xAI, and should be read as forward-looking guidance rather than fact.

The release also sits inside a larger corporate shift. SpaceX acquired xAI in February 2026 and folded the AI company into its own brand as SpaceXAI following SpaceX’s IPO in June. Analysts have pointed to the growing overlap between xAI’s engineering-focused training data and SpaceX’s own rocket and satellite operations as a sign that Musk intends the model line to become as useful for hardware and physical-systems engineering as it is for software. Gene Munster of Deepwater Asset Management, an independent analyst, wrote that Musk “effectively shipped a product on time today” despite the delays, and separately predicted that a future Grok 5 release could become an unusually capable engineering tool if it draws on SpaceX’s full body of operational data — a forward-looking opinion, not a confirmed product plan.

What to Watch Next

  • Independent verification: All benchmark figures released so far come from xAI’s own materials. Third-party testing over the coming weeks will show whether Grok 4.7 holds up outside company-controlled evaluations.
  • Multimodal performance: Musk has already flagged this as a weaker area; whether xAI addresses it in a near-term update is worth tracking.
  • Grok 4.8 and 4.9: Musk’s own roadmap points to two more incremental releases before Grok 5, though neither has a confirmed date.
  • Enterprise and coding-tool adoption: How quickly Cursor, GitHub Copilot and other platforms shift default usage toward Grok 4.7 will be a practical signal of real-world uptake, separate from benchmark scores.

Conclusion

Grok 4.7 is a confirmed, shipped product: it is live today, priced the same as its predecessor, and backed by a published set of company benchmarks showing improvement over Grok 4.6. What remains less settled is how it holds up against Claude Fable 5.1 and other frontier models once independent testers get access, and whether xAI’s low pricing is enough to offset a benchmark position that, by the company’s own published numbers, still sits behind the category leader on several professional-work tests. For now, Grok 4.7 reinforces xAI’s consistent bet: undercut rivals on cost and distribute widely, rather than chase the top spot on every leaderboard.

Sources

1 thought on “xAI Launches Grok 4.7, Betting on Price and Speed Over Benchmark Wins”

Leave a Reply

Your email address will not be published. Required fields are marked *