Elon Musk’s SpaceXAI (formerly xAI) officially unveiled its next-generation flagship model, Grok 4.6, on Wednesday (the 12th), focusing major upgrades on long-running AI agents, complex programming, knowledge work, and interactive and visual tasks. SpaceXAI stated that Grok 4.6 has entered the top tier across multiple frontier model benchmarks, entering the market with an aggressive pricing of $2 per million input tokens and $6 per million output tokens—approximately half the cost of some competitors—aiming to capture developer interest through a 'high-performance + low-cost' strategy.

In its official announcement, SpaceXAI explained that Grok 4.6 builds upon Grok 4.5 with enhanced capabilities specifically tailored for long-duration agents and more challenging interactive and visual workflows. The new model can handle multi-step research tasks, codebase analysis, and application development, autonomously testing and validating outputs during extended operations. This enables the model not just to answer questions but to continuously complete entire workflows.

Grok 4.6 underwent extended supplementary training beyond Grok 4.5, incorporating filtered model-generated reasoning data and high-quality engineering datasets, along with improvements to optimizers and training methodologies. The company leveraged Grok 4.5 to generate SFT trajectories across varying reasoning intensities, agent frameworks, and domains such as STEM, software engineering, and knowledge work, then applied reinforcement learning to train Grok 4.6 on general-purpose programming, knowledge tasks, core code optimization, web development, and computer-aided design.

One of the most notable achievements is Grok 4.6 scoring 61 points on the Artificial Analysis Intelligence Index, tying with OpenAI’s GPT-5.6 Sol. This index aggregates multiple evaluations to assess overall performance in mathematics, science, programming, and reasoning. As such, Grok 4.6 significantly narrows the performance gap with leading models.

In the GDPVal-AA v2 benchmark, Grok 4.6 achieved 1,753 Elo, surpassing GPT-5.6 Sol’s 1,728 and Claude Fable 5’s 1,741 from Anthropic. This test emphasizes economically valuable professional work—including knowledge-intensive tasks across various professions and industries—closely mirroring real-world enterprise scenarios where AI agents are deployed.

However, Grok 4.6 did not lead in all evaluations. Other tests published by SpaceXAI show improvements: CursorBench v3.2 increased to 69.9% (from 66.7%); FrontierCode v1.1 rose to 61.3% (from 56.6%); APEX-Agents improved to 57.5% (from 47.1%); APEX-SWE reached 56.4% (from 53.6%); AA-Briefcase climbed from 1,313 to 1,577; and Harvey LAB increased from 12.9% to 15.8%. While competitive, it does not dominate every category, meaning Grok 4.6 has joined the elite group in agent, programming, and knowledge work—but hasn’t surpassed all rivals universally.

The ranking landscape remains dynamic, with Artificial Analysis updating models and metrics regularly. Thus, SpaceXAI’s reported standings reflect a snapshot comparison at launch time.

Beyond performance, Grok 4.6’s pricing strategy may become a key battleground in the next phase of AI competition. The standard API is priced at $2 per million input tokens and $6 per million output tokens, with a faster version available at approximately double the price. SpaceXAI claims this standard rate is about half that of other leading models.

This pricing directly targets developers and enterprise AI markets. The industry’s focus is shifting from pure benchmark scores to inference cost per task, token efficiency, speed, and practical task completion. As model capabilities converge, the computational cost required to achieve the same outcome may become a decisive factor for enterprise adoption.

Still, lower API unit prices do not guarantee proportional reductions in total operational costs. Factors such as token consumption rates, inference intensity, execution duration, and the number of tool calls made by agents influence actual expenses. Therefore, the 'half-price' claim refers primarily to list pricing, not necessarily a 50% reduction in total cost across all use cases.

To expand accessibility, SpaceXAI has broadened Grok 4.6’s developer channels. The model is now available via Cursor, Grok Build, and API, and is also offered through platforms like OpenRouter, Vercel, and Cloudflare (NET-US). For the first week, Cursor and Grok Build users receive double usage allowances, lowering the barrier for developers to test the new model.

This continues SpaceXAI’s recent push into the AI programming market. In July, the company launched the open-source version of Grok Build, releasing the core framework of its coding agent—including context assembly, tool calling, code reading and editing, command execution, skills, plugins, and sub-agents—and enabling local deployment.

Grok Build was originally positioned as a coding agent for professional software engineering and complex programming tasks. Users can have the model plan work, review and approve the execution plan, and view each code change as a diff file. With Grok 4.6 further enhancing agent capabilities, it effectively powers this development toolkit with a stronger model engine.

SpaceXAI also emphasized Grok 4.6’s improved ability to “turn ideas into products.” Given a vague product concept, the model can research unfamiliar domains, design application architectures, build core features, and iterate based on test results and user feedback. In visual and interactive projects, it can rapidly produce a complete first version.

This shift reflects how AI competition is evolving from “answering” to “executing.” Previously, large models demonstrated capabilities through Q&A, text generation, or code output. Now, companies like OpenAI, Anthropic, Google, and SpaceXAI are treating agent capabilities as the next frontier—models that can autonomously use tools, decompose tasks, revise outputs, and sustain long-term operations.

Grok 4.6’s performance on GDPVal-AA v2 exemplifies this trend. The 1,753 Elo score isn’t from simple Q&A—it comes from completing real professional tasks within an agent environment. SpaceXAI uses this to demonstrate that the model goes beyond chat proficiency and can handle workflows closer to actual enterprise processes.

Elon Musk quickly promoted the new model on social platform X, sharing posts about its 1,753 Elo score and leadership in several benchmarks, praising the model highly. This marks another instance of Musk personally championing Grok products recently.

The timing of Grok 4.6’s release is also significant. Grok 4.5 was only launched on July 16, meaning SpaceXAI released the next version in under a month. Grok 4.5 already prioritized programming, agent tasks, and knowledge work, highlighting joint training with Cursor—indicating SpaceXAI is rapidly shortening its model iteration cycle.

Recently, SpaceXAI has been expanding Grok beyond a chatbot into a full AI work platform. The July launch of Automations allows users to set up one-time tasks that Grok executes automatically based on schedules or triggers like email receipt, then reports back. The May release of Connectors enables Grok to directly connect to enterprise and personal applications.

FACT BOX

  • Source: PR Times
  • Category: New Product
  • Organizations: OpenAI / Anthropic / Google
  • Products / services: Grok 4.6 / Grok Build