In August 2024, xAI released Grok-2 and Grok-2 mini, and within days an early test version of the model climbed to third place on one of the AI industry’s most closely watched public benchmarks — a genuine, independently verified result. At the same time, the company shipped a new image generator with almost no content restrictions, and it produced fabricated images of political figures and copyrighted characters almost immediately. The launch captured both halves of xAI’s approach in one week: real technical progress, paired with safety decisions that drew swift criticism.
Table of Contents
What xAI Announced
xAI introduced Grok-2 and Grok-2 mini on August 13, 2024, describing Grok-2 as an early preview and “a significant step forward” from its previous model, Grok-1.5, with improvements in chat, coding, and reasoning. Grok-2 mini was positioned as a smaller, faster sibling built for cases where speed matters more than squeezing out every last point of quality.
Both models rolled out first to X Premium and Premium+ subscribers, replacing the older Grok-1.5 as the default assistant on the platform. xAI also said the models would reach developers through an enterprise API “later this month” — a promise that, as detailed below, took longer to materialize than the announcement suggested.
Notably, xAI did not publish Grok-2’s exact parameter count, architecture, or training data at launch. Independent estimates circulating since then have put the model somewhere between 270 billion and 314 billion parameters, but these figures come from outside analysis, not from xAI itself, and should be read as approximate.
The Benchmark Result That Got Independently Verified
The most concrete evidence of Grok-2’s capability didn’t come from xAI’s own marketing — it came from a public benchmark that had no reason to flatter the company.
Weeks before the official announcement, xAI quietly entered an unnamed test version of Grok-2, labeled “sus-column-r,” into the LMSYS Chatbot Arena (now known as LMArena), a site where users compare two anonymous models side by side and vote on which response is better. Once identities were revealed, LMArena’s own results showed that “sus-column-r” had climbed to third place on the overall leaderboard, based on more than 12,000 community votes, putting it roughly level with OpenAI’s GPT-4o. It also ranked second in coding, second in math, and fourth on especially difficult prompts.
That ranking matters because it’s independent: a third-party benchmark, run and published by LMArena rather than by xAI. It’s a different kind of evidence from the academic benchmark table xAI published in its own announcement, which showed Grok-2 competitive with GPT-4-Turbo, GPT-4o, Claude 3 Opus, and Claude 3.5 Sonnet on tests like MMLU, GPQA, and MathVista. Those numbers are xAI’s own reporting, not independently reproduced, and should be read with that distinction in mind — strong results, but self-reported ones.
A Repeated Funding Announcement, Not New Money
Coverage of the Grok-2 launch sometimes mentioned, in the same breath, that xAI had recently raised $6 billion. That’s true, but it’s not new information tied to Grok-2 — it’s a repost of older news.
xAI actually announced its $6 billion Series B round on May 26, 2024, nearly three months before Grok-2 shipped. The round valued the company at $24 billion, up from an $18 billion pre-money valuation, with backing from Andreessen Horowitz, Sequoia Capital, Fidelity Management & Research, Valor Equity Partners, Vy Capital, and Prince Alwaleed Bin Talal’s Kingdom Holding, among others. The funding helped xAI build out its infrastructure and get its first products to market, but it wasn’t a new capital event coinciding with the Grok-2 launch itself.
What Actually Trained Grok-2
It’s tempting to assume Grok-2 was trained on xAI’s now-famous Memphis supercomputer, Colossus — the facility that later made headlines for scaling to 100,000, then 200,000, Nvidia H100 GPUs in a matter of months. But that assumption gets the timeline backward.
Grok-2 was trained on roughly 8,000 GPUs, a fraction of Colossus’s eventual size. Colossus itself wasn’t built until after Grok-2 shipped; xAI constructed it specifically to train Grok 3 and the models that followed. In other words, the supercomputer most people associate with xAI’s rapid scaling didn’t exist yet when Grok-2 was being developed.
The Image Generator Controversy
The most consequential part of the Grok-2 launch wasn’t the language model at all — it was the image generator that came with it.
xAI partnered with Black Forest Labs, an AI image and video model startup, to integrate its Flux.1 model into Grok on X. Unlike rival tools such as OpenAI’s DALL-E or Google’s Gemini, which use layered filters to block depictions of real public figures, violence, or copyrighted characters, Grok’s image generator launched with almost none of those restrictions.
Users found this out almost immediately. Within days, people had generated fabricated images placing Donald Trump and Kamala Harris in fictional and offensive scenarios just ahead of the 2024 presidential election, sexualized depictions of public figures including Harris and Rep. Alexandria Ocasio-Cortez, and copyrighted characters like Mickey Mouse rendered in ways that drew swift criticism from commentators and, reportedly, from Disney. Grok itself, when asked about its own limits, reportedly gave users a generic description of safety guardrails that didn’t match what it was actually willing to produce.
No confirmed lawsuits or formal regulatory penalties tied specifically to this August 2024 episode have been reported. What is well documented is the volume and consistency of the criticism, and the fact that the underlying tension — minimal restrictions colliding with real-world misuse — did not go away. It resurfaced, more seriously, over a year later.
Enterprise API — Promised vs. Delivered
xAI’s original announcement said the enterprise API would be available “later this month” — meaning August 2024. That didn’t happen on schedule.
The public API beta, offering the models grok-beta and grok-vision-beta, didn’t actually launch until November 4, 2024, roughly three months after the original promise. Updated versions — grok-2-1212 and grok-2-vision-1212, with better instruction-following and multilingual support — arrived even later, on December 12, 2024, when xAI also confirmed pricing for the first time: $2 per million input tokens and $10 per million output tokens.
The gap between “later this month” and what actually shipped is a useful reminder that early xAI announcements described intentions, not guaranteed timelines.
Where Grok-2 Stands Today
Grok-2’s commercial life was relatively short. In December 2024, xAI made it free for all X users as part of a broader pricing shake-up that also raised the X Premium+ subscription from $16 to $22 a month. Grok 3 followed in February 2025, and Grok 4 in July 2025, each superseding Grok-2 as xAI’s flagship model.
In August 2025, xAI open-sourced Grok 2.5 — its 2024 flagship — publishing the model’s weights on Hugging Face under a license that allows free commercial and non-commercial use but prohibits using the model to train other AI systems. As of mid-2026, xAI’s active frontier models are Grok 4.5 and Grok 4.6, and Grok-2 survives mainly as a legacy, openly available research artifact rather than a current product.
Grok-2: Key Dates
| Date | Event |
|---|---|
| May 26, 2024 | xAI announces $6B Series B, $24B valuation |
| Aug 13, 2024 | Grok-2 and Grok-2 mini launch in early preview |
| Nov 4, 2024 | Public enterprise API beta launches |
| Dec 12, 2024 | Updated grok-2-1212/vision models and confirmed pricing |
| Aug 2025 | Grok 2.5 open-sourced on Hugging Face |
The Guardrail Pattern That Followed
The tension exposed in August 2024 — fast releases outpacing content safeguards — didn’t stay in the past. In January 2026, X restricted image generation and editing in its newer Grok Imagine tool to paying subscribers, following renewed backlash over non-consensual sexualized deepfakes, including images of women and, in some cases, children. The European Commission called the circulating images “unlawful” and “appalling,” and the UK’s data regulator pressed X for details on how it was complying with data protection law. The UK Prime Minister’s office went further, calling the subscription-only restriction “insulting” to victims.
That episode involved Grok Imagine, a separate and later product, not the original Flux-based generator bundled with Grok-2. But the underlying pattern — shipping generative image tools with minimal built-in restraint — is the same one that first became visible during Grok-2’s launch.
Conclusion
The verified evidence from Grok-2’s launch tells a mixed but clear story. On the technical side, an independent, crowd-voted benchmark placed an early Grok-2 checkpoint in the same tier as GPT-4o — a real, third-party-confirmed result, not just a company claim. On the safety side, the accompanying image generator shipped with few of the protections competitors had already built in, and the consequences showed up in public within days.
What remains uncertain is narrower than it might seem: xAI never disclosed Grok-2’s exact size or training data, and no confirmed legal action tied to the 2024 image controversy has surfaced. What’s worth watching is less about Grok-2 itself, which is now a legacy open-source model, and more about whether the guardrail problems that surfaced in 2024 — and resurfaced with Grok Imagine in 2026 — get addressed before the next release cycle, or simply repeat again.
Frequently Asked Questions
Did Grok-2 actually beat GPT-4o and Claude 3.5 Sonnet? An early Grok-2 test version reached third place on the independent LMSYS Chatbot Arena leaderboard, roughly matching GPT-4o based on community voting. On xAI’s own academic benchmarks, Grok-2 was reported as competitive with, rather than clearly superior to, GPT-4o and Claude 3.5 Sonnet — those specific numbers are self-reported by xAI.
Was the xAI $6 billion funding announced at the same time as Grok-2? No. The $6 billion Series B round was announced on May 26, 2024, about three months before Grok-2’s August 2024 launch. Mentions of the funding alongside Grok-2 coverage were a repost of earlier news, not a new financing event.
Was Grok-2 trained on xAI’s Colossus supercomputer? No. Grok-2 was trained on approximately 8,000 GPUs. Colossus, the Memphis supercomputer that later scaled to over 100,000 GPUs, was built afterward specifically to train Grok 3 and subsequent models.
Why did people criticize Grok-2’s image generator? The generator, built on Black Forest Labs’ Flux.1 model, launched with few restrictions compared to competitors. Users quickly produced fabricated images of political figures in offensive scenarios, sexualized depictions of public figures, and copyrighted characters used without permission.
When did the Grok-2 enterprise API actually become available? xAI initially said the API would launch “later” in August 2024, but the public beta didn’t arrive until November 4, 2024. Updated models with confirmed pricing followed in December 2024.
Is Grok-2 still xAI’s current AI model? No. Grok-2 was superseded by Grok 3 (February 2025), Grok 4 (July 2025), and later models. xAI open-sourced a version of it, Grok 2.5, in August 2025, and by 2026 its active flagship models were Grok 4.5 and Grok 4.6.
What happened to Grok-2’s image-generation guardrail problems later on? A similar controversy resurfaced with xAI’s newer Grok Imagine tool. In January 2026, following backlash over non-consensual sexualized deepfakes, X restricted image generation to paying subscribers — a move regulators in the EU and UK criticized as insufficient.

