Grok 4.5 brings coding, agentic tasks, and office work into one model

SpaceXAI launched Grok 4.5 on July 16, 2026, positioning it for coding, agentic tasks, and knowledge work across Grok Build, Cursor, and the API.

SpaceXAI launched Grok 4.5 on July 16, 2026, positioning it as a model for coding, agentic tasks, and knowledge work. The important part of the release is not only the model, but the execution environment around it: Grok 4.5 is designed to complete multi-step work, use tools, and produce working files and documents.

SpaceXAI says Grok 4.5 was trained on coding, science, engineering, and mathematics data, then reinforced with hundreds of thousands of tasks focused on multi-step software engineering and other technical work. The company also says agentic rollouts can run for hours inside an asynchronous training process using tens of thousands of NVIDIA GB300 GPUs.

The launch page reports official benchmark results including 83.3% on Terminal Bench 2.1 and 64.7% on SWE-bench Pro. SpaceXAI also says Grok 4.5 averages 15,954 output tokens per SWE-bench Pro task, compared with about 67,020 for Opus 4.8 on the page, or roughly 4.2 times fewer tokens. These figures come from SpaceXAI's announcement, and some comparisons cite other developers' system cards or benchmark leaderboards. They should be read as vendor-reported results rather than independent replication.

Grok 4.5 is now the default model in Grok Build. SpaceXAI shows examples of building a complete application from one prompt, researching the web before producing an Excel model, and using native PowerPoint and Word shapes for documents and presentations. The model is also available in Cursor and through the SpaceXAI API, with the launch page listing prices of $2 per million input tokens and $6 per million output tokens.

For agent workflows, the release shows how model capability and execution environment are becoming inseparable. Writing code is only the first layer. Tool permissions, file isolation, test execution, rollback, and human approval determine whether the output can safely become a deliverable. Once an agent can create files, change a repository, or produce office documents, auditability cannot be an afterthought.

Grok 4.5's benchmark and token-efficiency claims are worth tracking, but they should not be treated as a guarantee that every team will get the same result. Teams still need to evaluate the model against their own codebase, documents, tools, and acceptance criteria, while keeping high-risk actions inside an approval boundary. An agent completing more steps does not mean that it can skip verification.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.