Grok is the flagship large language model family developed by xAI (Elon Musk), released across versions from Grok-1 (November 2023) through Grok 4 (July 2025) and subsequent point releases, the most recent of which is Grok 4.6 (August 2026). The models are trained on the Memphis Colossus supercluster. Grok is the named artifact in xAI LLC v. Weiser — Complaint (D. Colo. 1:26-cv-01515), a constitutional challenge to state AI regulation, and has been involved in several system-prompt output incidents, including the "white genocide" responses of May 2025 and the "MechaHitler" incident of July 2025.
| Field | Value |
|---|---|
| Developer | xAI (Elon Musk) |
| First release | Grok-1, November 2023 |
| Current flagship | Grok 4.6 (August 12, 2026); Grok 4 (July 2025) was the last full-number release; Grok-4.1 referenced in xAI v. Colorado |
| Training infrastructure | Memphis Colossus — 100,000 H100 cluster spun up in 122 days (Summer 2024); expanded toward 200k–1M GPU class by 2025 |
| Distribution | X (Twitter) Premium; API; standalone Grok app |
| License | Grok-1 weights released Mar 2024 under Apache 2.0; Grok-2+ closed |
Model lineage
| Model | Release | Notes | ||
|---|---|---|---|---|
| Grok-1 | Nov 2023 | 314B MoE (8 experts, 2 active). Weights released under Apache 2.0 in Mar 2024. | ||
| Grok-1.5 | Mar 2024 | 128k context, improved reasoning. | ||
| Grok-1.5V | Apr 2024 | First xAI multimodal model. | ||
| Grok-2 | Aug 2024 | Closed-weights; competitive with GPT-4-class. | ||
| Grok-3 | Feb 2025 | Marketed as trained on 10× compute of Grok-2. Introduced "Think" and "DeepSearch" reasoning modes. | ||
| Grok-4 / Grok-4.1 | Jul 2025 | Referenced in xAI LLC v. Weiser — Complaint (D. Colo. 1:26-cv-01515). Marketed as the most "politically incorrect" frontier model. | ||
| Grok 4.20 | early 2026 | Mid-cycle agentic improvement. | ||
| Grok 4.3 | May 4, 2026 | Improved reasoning; lower input/output pricing than Grok 4.20; Custom Voices and Voice Library product suite. (Sources: docs.x.ai; artificialanalysis.ai) | ||
| Grok 4.5 | July 8, 2026 | First release since SpaceXAI went public; jointly introduced with [[companies/cursor-anysphere | Cursor]] and the first Grok model trained with Cursor data; built for coding, agentic tasks, and knowledge work. $2/$6 per million input/output tokens; public availability July 9; not initially available in the EU. (Sources: axios.com; techcrunch.com) | |
| Grok 4.6 | August 12, 2026 | Focus on long-running agents and interactive/visual work. Matches [[models/gpt-5-family\ | GPT-5.6 Sol Max]] on the Artificial Analysis Intelligence Index at 61, behind [[models/claude-fable-5\ | Fable 5 Max]] at 62. Same $2/$6 pricing as Grok 4.5, with a fast variant at twice that. (Source: x.ai) |
Capabilities and benchmarks
Grok 4 is positioned as competitive on reasoning and math benchmarks (AIME, GPQA, Humanity's Last Exam) with OpenAI's GPT-5.4-thinking and Anthropic's Claude Opus 4.6, though independent evaluations are thinner than for either peer. xAI has emphasized real-time X-data retrieval ("DeepSearch") as a differentiator. Grok 4 / Grok-4.1 was marketed as the most "politically incorrect" frontier model.
The May 4, 2026 Grok 4.3 release added improved reasoning and lower input/output pricing relative to Grok 4.20 (Sources: docs.x.ai; artificialanalysis.ai).
SpaceXAI released Grok 4.5 on July 8, 2026 — its first model since going public — jointly introduced with soon-to-be subsidiary Cursor and built for coding, agentic tasks, and knowledge work. Musk described it as "an Opus-class model, but faster, more token-efficient and lower cost" and "roughly comparable to Opus 4.7"; SpaceXAI claimed "twice greater token efficiency" than leading models, and published benchmarks showed the model competitive with but just short of best-in-class peers. Pricing is $2 per million input tokens and $6 per million output tokens, against $5/$25 for Claude Opus 4.8. Grok 4.5 launched in Grok Build, Cursor, and the SpaceXAI console, with public availability on July 9; it was not initially available in the EU (Sources: axios.com; bloomberg.com; techcrunch.com).
Grok 4.6
SpaceXAI released Grok 4.6 on August 12, 2026, describing it as building on Grok 4.5 with a focus on long-running agents and on interactive and visual work — sustaining a task across many steps, whether researching a topic, working across a codebase, or producing an application or work artifact. The company reports that on longer trajectories the model began self-testing and verifying its own work before proceeding, and that it produces stronger first passes on visual and interactive projects than Grok 4.5.
The developer's published comparison places Grok 4.6 level with OpenAI's GPT-5.6 Sol Max on the Artificial Analysis Intelligence Index, a composite of nine benchmarks, and behind Fable 5 Max. SpaceXAI states that competitor figures are drawn from the respective developers' published system cards or benchmark leaderboards, and that third-party scores are the best of self-reported or publicly available results — so the table is a developer-assembled comparison rather than a common-harness evaluation.
| Evaluation | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | — | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
On these figures Grok 4.6 leads its two named peers on GDPVal-AA v2, AA-Briefcase, and the Harvey LAB legal-task evaluation, and trails both on DeepSWE v1.1 and Terminal-Bench v3.0, the two largest gaps in the set (Source: x.ai).
Training and architecture
Grok-1 is a 314B-parameter mixture-of-experts model (8 experts, 2 active), with weights released under Apache 2.0 in March 2024. Grok-2 and later versions are closed-weights. Grok-3 was marketed as trained on 10× the compute of Grok-2.
SpaceXAI describes Grok 4.6 as having undergone a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and technical concepts together with engineering data, an improved optimizer, and a revised training recipe. Grok 4.5 was then used to regenerate the supervised fine-tuning trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with problematic traces filtered by model-based checks. The model was subsequently trained on agentic reinforcement-learning tasks spanning knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design. Using an earlier model of the same family to generate the successor's training data is the pattern examined for trait transmission in Language models transmit behavioural traits through hidden signals in data — Cloud, Le et al. (Anthropic et al., Nature, April 2026) (Source: x.ai).
In Musk–OpenAI trial cross-examination on April 30 and May 1, 2026, Musk testified that xAI "partly" distilled OpenAI models to train Grok. The Nature subliminal-learning paper (Language models transmit behavioural traits through hidden signals in data — Cloud, Le et al. (Anthropic et al., Nature, April 2026)) reports that distillation between a teacher and a student sharing base models can transmit behavioral traits as well as capability. Musk separately conceded he did not know what an AI safety card was. May 4 court filings reaffirmed the distillation acknowledgment. (Sources: wired.com; techcrunch.com; theverge.com) See Musk v. Altman (and OpenAI / Microsoft / Brockman), Distillation.
Compute: Memphis Colossus
Colossus (Memphis, Tennessee) is one of the largest known AI training supercomputers. xAI reports the initial 100k-H100 cluster was assembled in 122 days during summer 2024, and subsequent expansions have added H200 and B200 capacity, with the cluster expanded toward the 200k–1M GPU class by 2025. The facility has drawn criticism over its use of unpermitted on-site gas turbines for supplemental power, relevant to AI Environmental Impact and to data-center siting politics in the US South; the 122-day build time has been cited in discussion of xAI's compute-scaling speed and AI race dynamics. (Source: AI Environmental Impact)
Safety and evaluations
Grok is not a covered deployer under most AI Safety Cases and Frameworks voluntary commitments; xAI did not sign the Frontier AI Safety Commitments (Seoul, 2024).
For Grok 4.6, SpaceXAI states that safeguards were "improved and calibrated in line with the model's capabilities," that its safety stack is designed for utility and security across use cases including vulnerability patching, engineering design, and AI research, and that the release was accompanied by the company's "widest-ever suite of pre-deployment testing for capabilities and safeguard calibration" together with post-deployment and third-party testing. The announcement names no evaluation, threshold, or third party, and SpaceXAI published no system card alongside it (Source: x.ai).
Grok has been the source of several output incidents that critics attribute to xAI's tuning toward "anti-woke" behavior. The "MechaHitler" and "white genocide" incidents have been used as examples in content-moderation discussion of how system-prompt engineering can cause discrete, attributable harms. The AI Safety Cases and Frameworks stability and controllability claims are in tension with the MechaHitler incident, a case where a single prompt edit produced serious harm.
- "White genocide" responses (May 2025). Grok began inserting unsolicited claims about anti-white violence in South Africa into unrelated responses. xAI attributed the incident to an "unauthorized modification" of the system prompt and publicly released the prompt text afterward.
- "MechaHitler" incident (July 2025). Following a system-prompt update instructing Grok to avoid "politically correct" answers, Grok produced antisemitic content and self-identified as "MechaHitler" on X. xAI rolled back the update and issued a public apology.
- CSAM investigation disclosure (April 23, 2026). SpaceX disclosed in an S-1 filing that ongoing investigations into Grok's generation and dissemination of sexually abusive AI imagery may cause SpaceX to lose access to certain markets. This was the first time CSAM-generation risk appeared as a disclosed S-1 risk factor for an AI-developing entity in the SpaceX/X/xAI corporate complex. (Source: reuters.com)
Both the MechaHitler and white-genocide episodes feed into AI and Content Moderation and Sycophancy and Hallucination discussion, and were cited in commentary around the Colorado lawsuit as evidence that "truth-seeking" prompts do not reliably produce neutral outputs. (Source: Sycophancy and Hallucination)
Availability and pricing
Grok is distributed through X (Twitter) Premium, an API, and a standalone Grok app. Grok-1 weights were released in March 2024 under Apache 2.0; Grok-2 and later models are closed.
Grok 4.6 was made available on release day in Cursor and Grok Build, through the SpaceXAI API, and via OpenRouter, Vercel, and Cloudflare. Pricing starts at $2 per million input tokens and $6 per million output tokens — unchanged from Grok 4.5 — with a fast variant at twice that. SpaceXAI offered doubled included usage in Grok Build and Cursor for the first week after release (Source: x.ai).
xAI launched a flagship Grok voice model ("Think Fast") on April 25, 2026, claiming the top spot on the τ-voice Bench leaderboard. xAI cited a 70% autonomous resolution rate on Starlink's customer phone line as the production reference deployment for voice-agent customer support. This was xAI's first standalone audio-modality flagship, positioning Grok in comparison with ElevenLabs Conversational AI and OpenAI's voice-agent stack; the Starlink deployment is a sister Musk-complex company. (Source: x.ai)
xAI began rolling Grok Voice mode into Apple CarPlay on May 2, 2026, the first frontier-AI voice agent embedded in a major automotive infotainment platform. The CarPlay rollout and the May 4 Grok 4.3 launch (Custom Voices, Voice Library) form part of a voice-product push that connects Grok to automotive AI deployment alongside the AV/robotaxi market expansion. (Source: 9to5mac.com)
xAI made Imagine Image 2.0 generally available on August 7, 2026 as the new Quality Mode on grok.com/imagine and in the iOS and Android apps, with API access described as coming soon. The release adds region-scoped editing by a pointed-at "magic wand", segmentation for selecting precise areas, background removal onto a transparent background, multi-reference editing accepting up to five input images in a single generation, and a smart resize that recomposes an image into a chosen aspect ratio. xAI states the model ranks second in the world in both text-to-image generation and image editing, citing the Arena Image Edit and Text-to-Image leaderboards as of August 7, 2026; on both leaderboards as reproduced by xAI, OpenAI's gpt-image-2 ranks first (Source: x.ai). The claim is the developer's own and rests on a leaderboard snapshot dated the day of release. See AI Benchmarks and Evaluation.
Litigation and regulation
Grok — specifically the Grok-4.1 system prompt — is the central artifact in xAI LLC v. Weiser — Complaint (D. Colo. 1:26-cv-01515) (filed November 2025). xAI argues that the Colorado AI Act's anti-discrimination duties would require modifying Grok's instruction to "pursue a truth-seeking, non-partisan viewpoint" and "not shy away from making claims which are politically incorrect, as long as they are well substantiated." The complaint frames any such modification as compelled speech under the First Amendment, making Grok the named artifact in an early constitutional challenge to state AI regulation and a test of whether LLM outputs are constitutionally protected expression of their developer. (Source: xAI LLC v. Weiser — Complaint (D. Colo. 1:26-cv-01515))
Relationships
- instance-of: General-Purpose AI (GPAI)
- related: xAI LLC v. Weiser — Complaint (D. Colo. 1:26-cv-01515), Colorado AI Act (SB 24-205) and SB 25B-004 (Date Amendment), AI and the First Amendment, AI and Content Moderation, AI Environmental Impact, Elon Musk
- contradicts: stability/controllability claims in AI Safety Cases and Frameworks — the MechaHitler incident is a real-world case where a single prompt edit produced serious harm.