Grok 4.7 Bets on Longer Coding Jobs

Most AI models are sprinters. Ask a quick question, get a quick answer. But a lot of real work, the kind lawyers, engineers, and software developers actually get paid for, is a marathon. It takes hours, requires backtracking, and punishes sloppy shortcuts. xAI's new Grok 4.7 is pitched squarely at that marathon.

What it is

Grok 4.7 is xAI's latest large language model, the software that reads and writes text and code. The headline claim is stamina. The company says it "works longer on difficult tasks" and "checks its own work more carefully." Under the hood, it uses a bigger base model than its predecessor, Grok 4.6, and went through a longer round of reinforcement learning, a training method where the model is rewarded for good answers and nudged away from bad ones. This time the practice problems leaned toward tasks that take many hours to finish.

xAI also trained the model to natively understand its own "Grok Bot harness," the software wrapper that lets a model chat and use tools. In plain terms, that is meant to make it smoother at conversation and general knowledge work.

Why it matters

The interesting part is the price tag, or rather the lack of a new one. Grok 4.7 costs the same as Grok 4.6: $2 per million input tokens and $6 per million output tokens. Tokens are the chunks of text a model reads and writes, so this is essentially the metered rate for using it. Keeping the price flat while claiming better results is the whole pitch here.

The benchmark numbers back up a steady step forward rather than a leap. On CursorBench 4.0, a test of longer-running coding tasks, Grok 4.7 scores 46.3%, up from 40.4% for Grok 4.6, and ahead of GPT-5.6 Sol Max at 41.7%. Rival Fable 5.1 Max still leads that test at 51.8%, but it also costs far more, at $10 input and $50 output per million tokens. xAI frames its edge as price-performance, and on that measure the argument holds.

Elsewhere the picture is mixed, which is refreshingly honest to read in the numbers. Grok 4.7 tops the pack on electrical engineering (EEBench, 64%) and legal work (Harvey's benchmark, 19.6%, though a low ceiling across all models there tells you these tasks remain hard). On clinical reasoning it trails, scoring 56.7% against 62.1% for Fable 5.1. On terminal work, the command-line tasks that power a lot of real software engineering, it improves sharply over 4.6 but sits well behind Fable's 57.9%.

A caveat worth flagging: these are the maker's own reported figures, not independent audits, and one strong software-engineering score carries an asterisk for "high effort," meaning the model was allowed to work harder than the default setting.

The safety angle

Grok 4.7 ships with what xAI calls an entirely new safeguard stack, and this is where the company sounds most confident. It claims the model is its strongest yet at refusing genuinely dangerous requests while still helping with legitimate ones. In cybersecurity, it says the model lets through only 3.3% of risky dual-use prompts on HackerBench v0.3 while rarely blocking honest security work. It also tops LatchBio's biosafety benchmark at 62.4%. The tricky balance in these fields is being useful to defenders without handing tools to attackers, and xAI is now giving select cybersecurity partners invite-only access to the model's red-team, or attack-simulation, abilities for defense research.

What's next

Grok 4.7 is available today in Cursor and Grok Build, through xAI's API, and via various coding tools and cloud platforms. A faster variant runs at double the output speed for double the price.

The real test will not be the launch-day charts but whether the endurance framing holds up in messy, hours-long real work. If AI is going to move from answering questions to actually finishing jobs, stamina and self-checking are the right things to compete on. Whether Grok 4.7 delivers them outside the benchmark suite is what the next few months should reveal.