Alibaba says Qwen3.8-Max ran 33 recursive self-improvement cycles, as FBI copying allegations hang over it

HPCwire reports that Alibaba ran Qwen3.8-Max through more than 33 recursive self-improvement cycles over roughly a month. Its Artificial Analysis Intelligence Index score rose from 40 to 45 over that period. If the result holds up, it is one of the most concrete public claims of a frontier lab using a model's own outputs to improve itself iteratively at production scale.
Alibaba has not detailed the mechanism in the sources. In general, recursive self-improvement here means the model generates training data, critiques or reward signals, or code changes to its own pipeline, and each cycle's model seeds the next. A five-point gain on a composite index in a month is significant but not a runaway. It looks more like efficient automated post-training than an intelligence explosion.
The political backdrop complicates the story. The FBI recently accused Alibaba of 'malicious' copying from Anthropic. Days later, the Times of India reported that the US Federal Register website was running an Alibaba model. That combination is likely to prompt procurement scrutiny in Washington. It also echoes OpenAI's October 1 disclosure of a Moonshot-linked campaign to extract reasoning traces. Distillation from US models is becoming the default accusation against Chinese labs, which makes any self-improvement claim harder to evaluate in isolation.
The broader Qwen family is moving fast on other fronts: Qwen3.8-27B is now on Nebius, Qwen Intelligence launched on HONOR phones, and r/LocalLLaMA users are running Qwen3.8 Flash Next at 50 tokens per second on 12GB laptops. Watch for a technical report on the cycles and any US government response.