All AI News
    Latent SpaceSaturday, July 25, 2026 11 min read
    AI

    Claude Opus 5: Fable-level performance at Opus price (half of Fable)

    Claude Opus 5 matches or beats Fable on most benchmarks at half the price, but formal evals understate real-world gains.

    Koko brief

    Claude Opus 5 matches or beats Fable on most benchmarks at half the price, but formal evals understate real-world gains.

    Anthropic's Friday drop of Opus 5 lands atop Artificial Analysis's intelligence index and ties Fable 5 on software-engineering benchmarks—while costing half as much. Epoch's composite score shows a razor-thin gap versus Fable, yet practitioners report clear practical superiority, especially in agentic browser control and coding tasks. The disconnect exposes a structural problem: current frontier evals can't capture qualitative capability jumps that both users and Anthropic itself acknowledge exist.

    Watch: whether Chatbot Arena's forthcoming human-preference scores finally close the gap between benchmark rankings and practitioner reality for frontier models.

    In a rare Friday release, Opus 5 took the headlines today. Athrough most of its official benchmarks have it technically beating Fable , the official messaging still says it “ comes close ”. This mostly reflects the difficulty of Evals - today’s AIE track drop - not reflecting “big model smell” that Anthropic obviously knows Fable retains but can’t measure. Fortunately, independent evaluations of Opus confirm the outperformance: @AnthropicAI has released Claude Opus 5, the new leader on the Artificial Analysis Intelligence Index, and ","username":"ArtificialAnlys","name":"Artificial Analysis","profile_image_url":"https://pbs.substack.com/profile_images/2042402069320290304/A8C1lP07_normal.jpg","date":"2026-07-24T22:10:41.000Z","photos":[{"img_url":"https://pbs.substack.com/media/HOBjK6cbIAA2Yph.jpg","link_url":"https://t.co/SFuDwqY6XE"}],"quoted_tweet":{},"reply_count":16,"retweet_count":45,"like_count":451,"impression_count":34514,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":false}" data-component-name="Twitter2ToDOM"> And the improved efficiency story, beyond just pricing, is also important… although it only just matches GPT 5.6 Sol: AI News for 7/23/2026-7/24/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap Top Story: Claude Opus 5 model launch What happened Anthropic’s Claude Opus 5 launch triggered a mix of benchmark scrutiny, strong anecdotal coding-agent praise, and renewed debate about frontier model evaluation. Multiple tweets explicitly discuss Claude Opus 5 as a newly launched model and compare it to other frontier systems on coding and general capability metrics, including Epoch’s ECI assessment , a FrontierCode anomaly discussion , and early user reactions from tool-use workflows like browser automation @abacaj , @abacaj . Epoch reported that Claude Opus 5 achieves an ECI of 159 , “slightly below Fable 5’s value of 161,” while matching Fable 5 on SWE-ECI at 161 on software engineering benchmarks @EpochAIResearch . The ECI result immediately drew criticism from users who felt the score understated Opus 5’s practical improvements; one response called it “incredibly underrated,” noting it appears only 1 point better than Opus 4.8 despite seeming “much better at everything” in practice @scaling01 . The same user argued for harder public benchmarks @scaling01 . A separate thread highlighted an apparent benchmark irregularity: Opus 5 scored better on FrontierCode at medium effort than at higher effort , even though more effort improved performance on other evals @jerhadf . That suggests either task-specific search/effort tradeoffs or evaluation instability rather than monotonic gains from extra inference-time compute. Several technically literate users praised Opus 5’s coding performance. Mikhail Parakhin @MParakhin —said “Best-of-n rules” and reported a clear head-to-head win against Fable “for math and everything, really,” while wishing it were available in Codex. Arena promoted first impressions of Opus 5 and said leaderboard scores based on real-world use were coming soon @arena , indicating community evals were still catching up at posting time. Nous Research’s portal added access to the model, with a tweet saying users could directly use Opus 5 through Nous Portal and that a 20% discount applied to all models including Opus 5 @witcheer . This is distribution/availability rather than a capability claim. User anecdotes emphasized browser control / agentic tool use . One post said Opus 5 opened the browser and canceled a ChatGPT Pro subscription @abacaj , followed by “This thing can really drive a browser wow” @abacaj . These are isolated demos, not systematic evals, but they align with broader market interest in computer-use agents. Other early reactions were more memetic than technical, including “Opus 5 subway FPS result” @bijanbowen , “On Claude bro” @andrew_n_carr , and “They’re terrified of Anthropic” @teortaxesTex . These reflect sentiment but not evidence. Technical details Epoch Capabilities Index (ECI): Claude Opus 5 ECI = 159 Fable 5 ECI = 161 Claude Opus 5 SWE-ECI = 161 , matching Fable 5 on software engineering @EpochAIResearch Community response noted the model appears only +1 ECI point vs Opus 4.8 , which some readers considered too small relative to qualitative gains @scaling01 , @scaling01 . FrontierCode behavior: one evaluator noted medium-effort > high-effort on FrontierCode for Opus 5 despite the usual pattern of improvement with more effort elsewhere @jerhadf . The tweet does not provide raw numbers in this excerpt, but the central technical point is that increased effort was not uniformly beneficial. Anecdotal comparative claims: A clear head-to-head win vs Fable in one user’s testing, especially with best-of-n sampling @MParakhin Matching “mythos” in one ecosystem summary post, though without attached numbers @eliebakouch Facts vs opinions More factual / measurement-oriented claims Epoch’s benchmark statement that Opus 5 scored 159 ECI and 161 SWE-ECI is the clearest empirical claim in the set @EpochAIResearch . Arena’s statement that first impressions are available and real-world leaderboard scores are forthcoming is factual but incomplete @arena . Nous Portal offering access to Opus 5 with a 20% discount is a product-availability fact @witcheer . Interpretations / opinions “ECI is underrated” and “we need harder public benchmarks” are opinions about benchmark validity and sensitivity @scaling01 , @scaling01 . “How to shake faith in any benchmark: show Anthropic doing meh on it” is rhetorical skepticism about benchmark discourse and community bias @teortaxesTex . “Best-of-n rules” and Opus being a “very clear winner” over Fable are informal practitioner judgments, useful but nonstandardized @MParakhin . “They’re terrified of Anthropic” and AGI-timeline speculation tied to Anthropic are pure opinion/speculation rather than launch evidence @teortaxesTex , @teortaxesTex . Different opinions Supportive views The strongest positive interpretation is that Opus 5 is materially stronger in real use than public aggregate benchmarks currently show , especially for coding and tool-use tasks. @MParakhin reports it beats Fable in his own testing and says best-of-n improves outcomes. @abacaj , @abacaj highlight effective browser automation, suggesting practical agentic competence. @bijanbowen calling the “subway FPS result” the best one yet implies visual/computer-use demo quality impressed viewers. @eliebakouch places Opus 5 among top closed-model releases and says it is “matching mythos,” framing it as a top-tier frontier entrant. Skeptical / critical views The main criticism is not that Opus 5 is weak, but that benchmarking around it is unstable, underspecified, or misaligned with user impressions . @jerhadf points to a puzzling effort scaling inconsistency on FrontierCode. @scaling01 argues the ECI result seems too low relative to observed improvements and uses that to call for harder public benchmarks @scaling01 . @teortaxesTex implies some benchmark trust is contingent and anthropic-specific results provoke benchmark criticism, i.e. social interpretation may be contaminating technical assessment. Neutral / analytic views Epoch’s framing is restrai

    Key takeaways
    • 01Anthropic's Friday drop of Opus 5 lands atop Artificial Analysis's intelligence index and ties Fable 5 on software-engineering benchmarks—while costing half as much.
    • 02Epoch's composite score shows a razor-thin gap versus Fable, yet practitioners report clear practical superiority, especially in agentic browser control and coding tasks.
    • 03The disconnect exposes a structural problem: current frontier evals can't capture qualitative capability jumps that both users and Anthropic itself acknowledge exist.
    Keep going — across the app