SAN FRANCISCO — SpaceXAI on Monday released Grok 4.7, the newest and most capable version of its flagship large language model, capping a remarkable stretch in which Elon Musk's AI team shipped two models in a single week.
Priced from $2 per million input tokens and $6 per million output tokens — identical to the outgoing Grok 4.6 — the model is aimed squarely at coding and knowledge work, the fast-growing arena where companies are handing more of their daily workflow to autonomous software agents. SpaceXAI says Grok 4.7 works longer on hard problems, checks its own output more carefully, and manages longer context than any Grok before it. The launch comes just days after the company's Grok Voice Transcribe 2.0 speech model, underscoring a release cadence few rivals can match.
A bigger base model and a longer training run
According to the release, Grok 4.7 is built on a new, larger base model than Grok 4.6 and was trained with a longer reinforcement-learning run weighted toward problems that take many hours to complete. The result, the company says, is a model that is markedly better at verifying its own work and holding a long thread of context without losing the plot.
SpaceXAI also trained Grok 4.7 to natively understand its Grok Bot harness — the team of always-on agents that run on their own cloud computer, work inside real apps, and keep going around the clock. That tighter integration makes the model stronger at conversational tasks, document creation and presentations, not just raw code.
Benchmarks put it near the frontier
The numbers back up the pitch. On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 scored 46.3%, up from 40.4% for Grok 4.6 and ahead of GPT-5.6 Sol at 41.7% — all at an average cost of just $4.69 per task, which SpaceXAI says puts it at the frontier for price-performance. On Terminal-Bench 4.0, covering multi-hour terminal work, the model nearly doubled its predecessor's score to 38.0%. It also posted category-leading results of 19.6% on the Harvey Legal Agent Benchmark and 64.0% on EEBench for electrical engineering. Full figures are laid out in SpaceXAI's release notes.
Safety built in from the ground up
Grok 4.7 ships with what the company calls an entirely new safeguard stack, and SpaceXAI describes it as its strongest model yet on refusals and jailbreak resistance. It topped LatchBio's biosafety benchmark at 62.4% and, on the company's own HackerBench v0.3, allowed only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. SpaceXAI has also begun giving select cybersecurity partners invite-only access to the model's red-team capabilities for defensive research.
A relentless cadence
The release continues a blistering pace that has become a hallmark of Musk's AI operation since it was folded into SpaceX. Grok 4.7 arrives weeks after the company put Grok Bot to work across enterprises at its Galaxy event in San Francisco, and it is available immediately in Cursor, the Grok API, Grok Build and major cloud platforms. With Grok 5 still promised before year-end, the gap between SpaceXAI and its better-funded competitors looks smaller with every shipment.