Apple Research Unveils Agent Seer and Studies LLM Bayesian Inconsistency

Agent Seer generates test scenarios for evaluating tool-using AI agents by leveraging tool specifications — function names and natural-language descriptions — instead of relying on static, hand-authored benchmarks. The key advantage is that it scales across rapidly evolving tool ecosystems that fixed benchmarks cannot track, a pressing problem as agent tool libraries change weekly.
The second paper studies LLMs as information-processing rules to quantify internal inconsistencies in their probabilistic beliefs, showing models are 'not consistently Bayesian' — they fail to update uncertain beliefs rationally as evidence arrives. Apple frames this as consequential for high-stakes domains like medicine, science, and law, where a system must represent and revise uncertainty coherently.
Together the works signal that Apple's research output remains strong even as it cuts Siri and ML product staff — a tension running through Apple's week. Both papers also speak to the industry's broader agentic-safety moment: Agent Seer offers a scalable evaluation approach just as the Hugging Face breach exposed how poorly agents are tested pre-deployment, and the Bayesian study quantifies exactly the kind of reasoning unreliability that makes autonomous agents risky. Watch whether these methods make it into Apple's own products or influence industry evaluation standards.