OpenAI CFO Sarah Friar proposes an AI ROI scorecard beyond token burn

OpenAI CFO Sarah Friar published an 'AI scorecard' aimed at a question that increasingly haunts boardrooms: how do you actually measure whether AI spending pays off? Rather than raw usage or token counts, Friar's framework evaluates four dimensions — useful work delivered, cost per successful task, dependability, and return on compute — reframing AI ROI around outcomes instead of consumption.
The move is strategically self-interested but genuinely useful. As a vendor selling tokens, OpenAI benefits from customers who can justify spend to their CFOs; a credible measurement framework helps enterprises green-light larger deployments. But the emphasis on 'cost per successful task' and 'dependability' is also an implicit acknowledgment that raw model capability isn't enough — reliability and unit economics are what determine whether an AI deployment survives budget scrutiny.
The framing dovetails with a parallel proposal this week from Anthropic's Boris Cherny, creator of Claude Code, who pitched a four-step enterprise framework for measuring AI success 'beyond token burn' — the same theme from a competing lab. That two of the biggest labs are simultaneously publishing ROI frameworks signals a maturing market: the era of buying AI on hype is giving way to procurement teams demanding measurable returns.
The skeptical read: a scorecard authored by a token vendor will naturally define success in ways favorable to buying more tokens, and 'return on compute' is easier to state than to measure honestly. But the underlying shift is real and healthy — after a year of AI spending justified by FOMO, both OpenAI and Anthropic are now arming buyers with metrics, a sign that ROI accountability is finally arriving. Enterprises should treat these as starting points, not neutral standards.