Back
xAIJuly 21, 20261 sources

Grok 4.5 reportedly leads agentic coding as Gemini 3.6 lags rivals

AI Analysis

xAI's Grok 4.5 has emerged as a notable leader in agentic coding, according to Gizmodo's benchmark coverage, reportedly outperforming Google's newly released Gemini 3.6 Flash on those tasks. The reporting frames Gemini 3.6 as trailing Claude Sonnet 5 and GPT-5.6 across most major benchmarks and now falling behind Grok 4.5 specifically in agentic coding — an edge the piece attributes partly to input from the Cursor team, suggesting close collaboration with a leading AI coding-tool maker.

The mechanism, if the Cursor connection holds, is that real-world coding-agent feedback loops — data on how developers actually use AI in an IDE — can meaningfully improve a model's agentic performance beyond raw benchmark tuning. That would be a competitive template: partner with the tools where agentic work happens.

Competitively, this reframes the model leaderboard for developers: rather than a single 'best' model, capability is spiky per task, echoing François Chollet's viral observation that 'the fundamental marketing trick of the AI industry is to make you believe the tallest spike is a floor.' Grok's agentic-coding strength, Claude Sonnet 5's cost-performance, and GPT-5.6's reasoning each win different slices. Elon Musk also touted Grok's ease for making video games. Skeptics note this is third-party benchmark commentary rather than a formal xAI release, and benchmark leadership shifts weekly. Watch for independent agentic-coding evals and whether Grok's Cursor tie-up becomes a formal integration.

Sources
AI Briefing
·Vendors·Curated by AI agents · Updated daily · 2026
Built by Koby Almog