Musk delays Grok 4.7 for more RLHF tuning over premature task abandonment

Musk's delay is a rare public admission of a specific RLHF failure mode: the model was over-penalizing response length during reinforcement learning, causing it to abandon tasks prematurely — cutting work short to keep answers brief. That's a concrete, diagnosable training bug, and Musk's 'a few more days to cook' framing suggests xAI caught it in evals before shipping.
Alongside the delay, xAI pushed its enterprise agent play: Grok Bot now connects to Salesforce, HubSpot, and Gong, targeting sales workflows directly. The September 15-17 'Grok Bot Galaxy' event in San Francisco challenges teams to build a company using Grok Bot as their only employee — a marketing stunt that doubles as an agentic-capability demonstration.
The delay contrasts with rivals shipping aggressively this week (GPT-6 Astra GA, Meta Muse launch), and note the naming confusion in the sources — Grok 4.6 shipped August 12 with a 500K context, while 4.7 is the delayed next release. xAI holding back for quality tuning is on-brand for the 'pace the frontier' moment Musk just endorsed, though skeptics will read it as xAI simply being behind schedule.