Grok 4.7 keeps the price and grows the model
SpaceXAI has released Grok 4.7, a larger model aimed at long coding and knowledge-work tasks, at the same token prices as Grok 4.6 and well below its closest rivals.

What actually changed in Grok 4.7
SpaceXAI has released Grok 4.7, describing it as the company's most capable model for coding and knowledge work and claiming it runs about twice as fast as comparable models at roughly half their price. The model went live on 21 September, in the same week that the company's Grok Voice Transcribe 2.0 speech model arrived.
Under the bonnet the release is a rebuild rather than a tune-up. SpaceXAI says Grok 4.7 uses a new and larger base model than Grok 4.6, trained with a longer reinforcement-learning run on a harder mix of tasks weighted towards jobs that take many hours to finish. The company also says the model verifies its own work more carefully, handles longer context, and was trained to understand the Grok Bot harness natively.
The numbers, and whose numbers they are
The headline figures come from the vendor's own comparisons, so treat them as claims rather than independent results. On CursorBench 4.0, which is built around longer-running coding tasks, Grok 4.7 scores 46.3% against Grok 4.6's 40.4%. On DeepSWE v1.1 it reaches 71.0% at high effort against 65.2%, on the multi-hour terminal benchmark Terminal-Bench 4.0 it climbs from 20.3% to 38.0%, and on EEBench, an electrical-engineering set, it moves from 53.0% to 64.0%. On the Harvey Legal Agent Benchmark it goes from 15.8% to 19.6%, and on HealthBench Professional from 48.5% to 56.7%.
Two rival frontier models sit in the same tables. Fable 5.1 (max) still beats Grok 4.7 on CursorBench at 51.8% and on Terminal-Bench 4.0 at 57.9%, and edges it on AA Briefcase at 1,678 against 1,657; GPT-5.6 Sol (max) leads on DeepSWE at 72.7%. What separates the release is not the top score but the meter running behind it: Grok 4.7 is priced at $2 per million input tokens and $6 per million output tokens, identical to Grok 4.6, against $4 and $20 for GPT-5.6 Sol and $10 and $50 for Fable 5.1.
Safety claims, also self-reported
Grok 4.7 ships with what the company calls an entirely new safeguard stack, and SpaceXAI says it is the strongest model it has tested on refusals and jailbreak resistance. The supporting numbers are again its own: 62.4% on LatchBio's biosafety benchmark, and a 3.3% pass-through rate for risky dual-use prompts on the company's HackerBench v0.3. SpaceXAI adds that select cybersecurity partners now get invite-only access to the model's red-team capabilities for defensive work.
Who can actually use it today
GitHub made Grok 4.7 available in Copilot on the same day, described as a gradual roll-out for Copilot Pro, Pro+, Max, Business and Enterprise subscribers, selectable in Visual Studio Code, Visual Studio, the Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains, Xcode and Eclipse. GitHub bills it at provider list pricing under usage-based billing, and administrators on business plans can switch it off from the Copilot model policy.
Our opinion
There is a real story in the price sheet rather than in the leaderboard. Grok 4.7 does not win the benchmark tables outright and does not pretend to; it lands second or third on most of them at a fifth of Fable 5.1's output price, which is precisely the trade an agent running for hours on a refactor will take. The numbers that should be read with the most caution are the safety ones, because a refusal rate measured on the vendor's own benchmark and a jailbreak-resistance claim tested by the vendor's own team are marketing until someone outside the building reproduces them - and the biosafety score in particular says nothing about whether the model will help a determined person past its own guardrails. Watch what independent evaluators publish over the next fortnight, and watch whether the lower token price survives the quarter, because the last two years of this market have been a race to cut prices to win developers and then quietly reprice once the habit is set.