Grok 4.7 enlarges the model, not the standard API price
SpaceXAI’s new Grok 4.7 brings a larger base model and longer-task training, while independent scores still put it behind GPT-6 and Claude Fable 5.1.

SpaceXAI released Grok 4.7 on September 21 with a larger base model and a set of training and inference changes aimed at longer tasks. The release keeps standard API pricing unchanged at $2 per million input tokens and $6 per million output tokens. It also arrives with a 500,000-token context window, four reasoning levels and broad availability through the API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. The central question is not whether Grok 4.7 is a larger and more capable release than its predecessor, but how much the published evidence supports that claim when vendor benchmarks and independent evaluation are separated.
What changed in Grok 4.7
The main architectural change disclosed for Grok 4.7 is the use of a larger base model. SpaceXAI also says the model received longer reinforcement learning for long tasks, self-verification and long-context training. These are not small positioning details: they describe a model intended to perform better when work extends over many steps, when outputs require checking, and when large amounts of text or code must be kept in view. The announcement therefore frames Grok 4.7 less as a narrow speed update and more as an attempt to improve reasoning and persistence across extended workflows.
The unchanged standard API price is a notable part of the release. SpaceXAI lists standard pricing at $2 per million input tokens and $6 per million output tokens, even though the model is described as larger and trained with additional methods. For developers and teams already budgeting around token costs, the message is that the default service tier does not introduce a higher unit price for the new model. That does not mean all access costs are unchanged: the faster service tier costs twice the standard price, creating a clear distinction between baseline access and lower-latency service.
The 500,000-token context window is another practical change to consider. A context window of that size can, in principle, allow a model to process large codebases, lengthy documentation, multi-file prompts or long-running task histories without the same degree of truncation. However, a large context window is a capacity figure, not a guarantee that the model will use every part of the context equally well. The fact that Grok 4.7 received long-context training is relevant because context length alone does not prove reliable retrieval, prioritisation or reasoning across the full window.
How the release is positioned
SpaceXAI is making Grok 4.7 available through several routes: its API, Cursor, Grok Build, OpenRouter, Vercel and Cloudflare. That distribution matters because it places the model not only behind a direct API but also inside development and deployment channels where model switching can be part of day-to-day workflows. The release is therefore aimed at both direct integrators and users who reach the model through platforms already used for coding, building or routing AI workloads. The availability list is an announcement from the company and should be treated as such, not as evidence of performance.
The four reasoning levels give users or developers a way to select how much reasoning effort the model should apply. The practical implication is that not every task needs to be run at the same setting. Simpler requests may not require the highest level, while more complex work can be assigned more reasoning. The verified facts do not specify how the four levels differ in latency, accuracy or cost, so any conclusion about the best setting would go beyond the evidence. What can be said is that SpaceXAI is exposing reasoning as a configurable part of the product rather than hiding it entirely behind one default mode.
Self-verification is also part of the announced method. In broad terms, self-verification indicates that the model is designed to check aspects of its own output during generation or reasoning. The release information does not provide enough detail to judge exactly how that process is implemented or how often it prevents errors. It should therefore be understood as a disclosed technique, not as proof that factual mistakes, coding errors or reasoning failures have been solved. The same caution applies to longer reinforcement learning for long tasks: it is relevant to the model’s design, but its effect has to be assessed through results.
What the numbers show
SpaceXAI reports three benchmark results for Grok 4.7: CursorBench 46.3, DeepSWE 71 and EEBench 64. These figures belong in the category of vendor benchmarks because they are presented by the company behind the model. They may be useful signals about the tasks SpaceXAI wants to emphasise, especially around coding or software engineering workflows suggested by the benchmark names and the distribution through Cursor. But vendor benchmarks do not carry the same evidentiary weight as independent testing. They show what the company reports under its chosen conditions; they do not by themselves establish market leadership or broad reliability.
The independent picture is less favourable. Artificial Analysis gives Grok 4.7 an intelligence score of 46, compared with 53 for GPT-6 and 53 for Claude Fable 5.1. On that evaluator’s scale, Grok 4.7 is behind both leading models named in the comparison. Artificial Analysis also reports a large gap on Terminal Bench. These are independent measurements, distinct from SpaceXAI’s own benchmark claims. They do not erase the significance of the larger base model or unchanged standard pricing, but they do limit any claim that Grok 4.7 has caught up with the strongest competitors in measured intelligence.
The difference between the vendor and independent results is the core analytical point. SpaceXAI’s figures can support the view that Grok 4.7 improves in areas the company is measuring and promoting. Artificial Analysis supports a different conclusion: in its independent testing, the model still trails GPT-6 and Claude Fable 5.1, with a particularly large gap on Terminal Bench. Neither set of numbers proves everything. Vendor benchmarks cannot settle independent ranking. Independent scores cannot describe every possible private workload. Together, they suggest a release with meaningful engineering changes but not one that, on the cited independent evidence, leads the field.
Practical implications
For API users, the clearest practical implication is that Grok 4.7 changes capability claims without changing the standard token price. A team using standard access would see the same listed rate of $2 per million input tokens and $6 per million output tokens, while gaining access to the larger model, the 500,000-token context window and the four reasoning levels. If the faster tier is needed, the cost changes materially because it is twice the standard price. The price structure therefore separates the model upgrade from the premium placed on faster service.
For long tasks, the release is designed to be more relevant than a short-prompt model update. Longer reinforcement learning for long tasks, long-context training and self-verification all point toward use cases where the model must maintain instructions, process extensive material or work through multi-step operations. Still, the independent results create a boundary around expectations. A larger base model and longer context may help with complex workflows, but Artificial Analysis’ score of 46 versus 53 for GPT-6 and Claude Fable 5.1 indicates that independent evaluators still see a performance gap at the top end.
The result is a measured upgrade rather than a clear displacement of the leading models. Grok 4.7 arrives with broader availability, a larger base model, long-task training, self-verification, long-context training, unchanged standard API pricing and a higher-cost faster tier. SpaceXAI’s reported CursorBench, DeepSWE and EEBench results show the company’s own benchmark framing. Artificial Analysis provides the independent counterweight, placing Grok 4.7 behind GPT-6 and Claude Fable 5.1 and noting a large Terminal Bench gap. The release matters, but the available evidence does not support treating it as the new benchmark leader.
Sources
- SpaceXAI releases Grok 4.7MarkTechPost · September 21, 2026
- Grok 4.7 is xAI's strongest model yet, but trails Claude Fable 5.1 and GPT-6The Decoder · September 21, 2026



