Grok 4.7 beats Fable 5.1 and GPT-5.6 Sol on Harvey’s legal agent test at a fifth of Fable’s input price

- Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark, ahead of Fable 5.1 at 6.7% and GPT-5.6 Sol at 2.5%.
- xAI prices Grok 4.7 at $2 per million input tokens, a fifth of Fable 5.1’s $10.
- Fable 5.1 still beats Grok 4.7 on Terminal-Bench 4.0, 57.9% to 38.0%.
Grok 4.7 was launched by xAI on Monday. It beat Anthropic’s Fable 5.1 and OpenAI’s GPT-5.6 Sol on a legal-agent benchmark, while charging a fifth of Fable’s price per input token.
Output runs $6 per million tokens against $50 at Anthropic
xAI’s release said that Grok 4.7 scored 19.6% on the Harvey Legal Agent Benchmark. Fable 5.1 got 6.7% on the same test and GPT-5.6 Sol got 2.5%.
That puts Grok 4.7 at about three times Anthropic’s score and close to eight times Sol’s. Cryptopolitan reported last month that Grok 4.6 already led that duo at 15.8%.
Legal work is one of the few domains where Grok 4.7 outperforms both competitors outright. It also clears Fable 5.1 on EEBench electrical engineering and edges past it on the DeepSWE coding test.
Grok 4.7 pricing is $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6. That input rate is 1/2 of GPT-5.6 Sol’s $4 and 1/5 of Fable 5.1’s $10.
Output runs $6 vs $20 for Sol, and $50 for Fable, about an eighth of Anthropic’s number.
xAI plotted the CursorBench 4.0 scores against the average cost per completed task and claims the model sits on the price-performance frontier. Grok 4.7 gets around 46% on that chart at ~$6 per task, while Claude Opus 5 needs almost double the spend for a similar outcome.
Fable 5.1 performs better at higher budgets, going up to 51.8% at around $17 per task. The release was “a strong combination of intelligence, speed & low cost,” Musk said on X.

Terminal-Bench 4.0 hands Anthropic a 57.9% to 38.0% lead
Grok 4.7 was runner-up to Fable 5.1 on GDPval and the AA Briefcase office-work test. Anthropic’s model keeps a clear lead on longer coding and terminal benchmarks.
The largest margin is on Terminal-Bench 4.0, where Fable 5.1 had 57.9% versus Grok 4.7’s 38.0%, a difference of some 20 points.
Fable also leads on CursorBench and on HealthBench Professional clinical reasoning, where GPT-5.6 Sol also beats Grok.
Grok 4.7 scored 1,695 Elo on GDPval, which scores models on tasks done by lawyers, nurses and financial analysts, up from Grok 4.6’s 1,605. That’s behind Fable 5.1’s tally of 1,735, but ahead of the 1,542 OpenAI’s newer GPT-6 Astra got on the same chart.
Grok 4.7 is ahead of GPT-5.6 Sol on five of seven benchmarks. It only loses DeepSWE and the clinical test.
Grok 4.7 runs on 2.1 trillion parameters, 40% more than the 1.5 trillion powering Grok 4.6. xAI has included supplemental SpaceX training data, including Starlink satellite telemetry and manufacturing records.
The company says the model is more likely to spend extra time on arduous problems and double-check its own answers than Grok 4.6.
The model launched on the Grok app, Cursor, Grok Build and the xAI API, with no waitlist.
All numbers here are from xAI’s own testing. CursorBench, the test xAI leads with, is made by Cursor, which SpaceX finished acquiring last month.
xAI benchmarked against GPT-5.6 Sol, and OpenAI’s newer GPT-6 Astra appears only in the GDPval, AA Briefcase and EEBench charts.
The smartest crypto minds already read our newsletter. Want in? Join them.
FAQs
How much does Grok 4.7 cost compared to Fable 5.1 and GPT-5.6 Sol?
Grok 4.7 costs $2 per million input tokens and $6 per million output tokens, against $10 and $50 for Fable 5.1.
Which benchmarks does Grok 4.7 win and lose?
Grok 4.7 wins the Harvey legal benchmark at 19.6%, while Fable 5.1 leads on CursorBench, Terminal-Bench and HealthBench Professional.
How big is Grok 4.7 and what was it trained on?
Grok 4.7 runs on 2.1 trillion parameters and includes supplemental SpaceX training data such as Starlink satellite telemetry.
Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses
Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.
















