Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for Grok 4.7 in GitHub Copilot: What Changes for Developers

Grok 4.7 in GitHub Copilot: What Changes for Developers

Grok 4.7 lands in GitHub Copilot on day one at $2 per million input tokens, with SpaceXAI reporting 71% on DeepSWE and a stronger agentic coding loop.

Dr. Nova Chen
Dr. Nova Chen★Sep 21, 2026★5 min read

A New Grok, and Copilot Has It the Same Day

SpaceXAI released Grok 4.7 on September 21, and within hours GitHub confirmed the model was rolling out inside Copilot. That same-day arrival is the part worth pausing on. A frontier coding model reaching the editor millions of developers already use, on launch day, is the clearest sign yet of how compressed the path from lab to keyboard has become.

  • Pricing: $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, with a Fast variant at twice the output speed for twice the price
  • Where it runs: Copilot Pro, Pro+, Max, Business and Enterprise, across VS Code, Visual Studio, JetBrains, Xcode, Eclipse, Copilot CLI and the Copilot cloud agent
  • SpaceXAI-reported benchmarks: 71.0% on DeepSWE v1.1, 46.3% on CursorBench 4.0 (up from 40.4% for Grok 4.6) and 38.0% on Terminal-Bench 4.0
  • Independent read: Artificial Analysis scores it 46 on its Intelligence Index v4.3.2, ranked 16th of 655 models, with a 500K-token context window

What Is Different Inside Grok 4.7?

According to SpaceXAI, Grok 4.7 sits on a larger base model than its predecessor and went through a longer reinforcement learning run on a harder mix of tasks. The company also says it was trained to natively understand the Grok Bot harness, and that it is better at verifying its own work before handing it back.

That last claim matters more than any single benchmark. In agentic coding, the expensive failure is not a wrong answer; it is a confidently wrong answer that the agent builds three more steps on top of. A model that checks its own output before moving on shortens the loop between a failing test and a working patch. When we covered Grok 4.6 and its 500K context window in August, the story was reach. This release is about reliability inside that reach.

How Should the Benchmark Numbers Be Read?

Carefully, and with attribution. The DeepSWE, CursorBench and Terminal-Bench figures are SpaceXAI's own results. The Artificial Analysis score is independent, but note that the index was updated to version 4.3.2 on September 19, so scores from before that revision are not directly comparable. The honest summary is that Grok 4.7 is a solid step up on SpaceXAI's own coding evaluations and lands in the upper tier of an independent leaderboard crowded with strong models.

SpaceXAI also reports its best refusal and jailbreak-resistance results to date, stating that only 3.3% of risky dual-use prompts got through on its HackerBench v0.3 evaluation. Safety numbers that ship alongside capability numbers are a welcome habit.

What Does It Mean for Copilot Users?

For individual subscribers, the model appears in the picker as the gradual rollout reaches your account. For Business and Enterprise organizations, GitHub notes that new models are enabled automatically unless an administrator has turned off the global default or disabled this model specifically, and usage is billed at provider list pricing. Teams that already run agent workflows, like the agent-driven Copilot runtime port to Rust GitHub described last week, now have another strong option to route work to.

Grok 4.7 is also live in Cursor, Grok Build and the Grok API.

A Quieter Upgrade: Grok Voice Transcribe 2.0

Three days earlier, on September 18, SpaceXAI shipped Grok Voice Transcribe 2.0. The company claims roughly twice the accuracy of version 1.0 at an unchanged price of $0.10 per audio hour for batch and $0.20 per hour for streaming. Short-phrase word error rate across 19 languages fell from 20.6% to 6.8%, and speaker diarization, word timestamps and key-term biasing are included at no extra cost. Atlassian's Loom has already adopted it for video transcription.

Put the two releases together and a pattern emerges: SpaceXAI is holding prices flat while pushing quality up. For developers, that is the most useful direction a model roadmap can take. More model analysis lives in our AI coverage.

Sources: GitHub Changelog — September 21, 2026; Unite.AI — September 21, 2026; Artificial Analysis — September 21, 2026; MarkTechPost — September 18, 2026.

More AI Stories