Skip to main content
    All AI News
    Latent SpaceWednesday, September 30, 2026 44 min read
    AI

    Why Dwarkesh Is Wrong About Computer Use + How OpenAI Shipped Its Jev Competitor in 1 Week

    Computer use agents now outpace average humans on some tasks—OpenAI's Sky acquisition is driving the shift.

    Key takeaways
    • 01OpenAI's acquisition of Sky Software has quietly reshaped what computer-use agents can do.
    • 02Ari Weinstein, formerly Sky's co-founder, argues the field is "180 degrees different" from months ago—combining accessibility trees, DOM data, Playwright, and generated code to close the loop between software writing and testing.
    • 03Meanwhile, OpenAI's API team shipped a Decisions API in roughly one week, already powering internal classification workflows and GPT Live.
    • 04Superhuman computer use is now the stated near-term target.
    Koko brief

    Computer use agents now outpace average humans on some tasks—OpenAI's Sky acquisition is driving the shift.

    OpenAI's acquisition of Sky Software has quietly reshaped what computer-use agents can do. Ari Weinstein, formerly Sky's co-founder, argues the field is "180 degrees different" from months ago—combining accessibility trees, DOM data, Playwright, and generated code to close the loop between software writing and testing. Meanwhile, OpenAI's API team shipped a Decisions API in roughly one week, already powering internal classification workflows and GPT Live. Superhuman computer use is now the stated near-term target.

    Watch: How OpenAI's Decisions API evolves beyond its current wrapper status—the team's willingness to clone competitors' patterns signals an unusually fast iteration cadence.

    In brief · from latent.space

    Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform.

    Read the full article at latent.space
    Show the full text · 44 min read

    Three months ago Dwarkesh, who has been posting incredible blogs and episodes about RL, posted a framing question for his video essay on RLVR which upset a lot of Computer Use folks: We are no strangers to learning in public and are no strangers to the stress of getting things wrong when you have a big platform. However, we were at Anthropic for the Computer Use launch , there for Claude Cowork with the first big podcast on it, organized the first Computer Use track at AIE presenting the state of the art, and were close to the OpenAI-Sky Software acquisition that now powers the complete domination of computer use that Codex enjoys today. This is why we’re excited to bring you today’s first guest, Ari Weinstein , cofounder of Sky and now leading all the amazing CUA progress that casuals might miss: Ari explains why Computer Use is now “180 degrees different” from where it was months ago, how agents are learning to debug and recover from failures, why combining screenshots with accessibility data, the DOM, Playwright, and generated code changes the speed equation, and why the next frontier is making agents literally superhuman at using software. OpenAI clones Jev In the second half, Nikunj Handa from OpenAI’s API team breaks down the new developer stack: async tool calling, mid-turn steering, WebSockets, UltraFast inference, the Decisions API, prompt caching, pre-warming, compaction, and the Agents API . Given that we were the first Jev podcast , we particularly focus on the unusually fast sprint on the Decisions API: And why it is just a Luna wrapper for now but the team is motivated and egoless enough to clone what they consider to be good patterns. We discuss: Why OpenAI thinks Computer Use has changed dramatically in just the last few months Dots and what changes when every agent gets its own Linux computer Why Computer Use can now complete some tasks faster than the average human The path from human-level to “literally superhuman” computer use Why modern agents are much better at debugging and recovering from failure How screenshots, accessibility trees, the DOM, Playwright, and generated JavaScript work together App Shots and why they give models much richer context than ordinary screenshots Why Computer Use can close the loop between writing software and testing it Trust, permissions, and safety when agents can make payments and operate websites Async function calling and why models no longer need to stop reasoning while tools run Mid-turn steering, WebSockets, and the architecture behind more responsive agents UltraFast inference and how OpenAI is pushing frontier models toward much lower latency The rapid internal story behind the Decisions API Why Decisions API is more than structured outputs at low latency GPT Live, fast tool calling, and real-time computer control How OpenAI is already using Decisions API for support classification and internal workflows Longer prompt caching, cache pre-warming , and cache-aware applications Server-side compaction vs manual compaction for long-running agent threads What should live inside an Agents API versus a developer’s own harness OpenAI as an “AI cloud” and the search for higher-level primitives beyond raw model APIs Ari Weinstein Product & Engineering, Computer Use at OpenAI X: x.com LinkedIn: linkedin.com Nikunj Handa Product, API at OpenAI X: x.com LinkedIn: linkedin.com Timestamps 00:00:00 OpenAI DevDay: Dots, GPT-6.1, Agents API, and Decisions API 00:02:52 Dots and Personal Cloud Computers 00:04:59 Why Computer Use Is “180 Degrees Different” 00:06:04 From Sky to Self-Debugging Computer Use Agents 00:09:24 How Computer Use Sees and Operates Software 00:12:09 From Faster Than Humans to Superhuman Computer Use 00:16:03 Agents API: Trust, Permissions, and Safety 00:17:31 Computer Use for Coding, Testing, and QA 00:19:14 GPT-6 APIs, Async Tool Calling, and UltraFast Inference 00:23:21 The Rapid Story Behind Decisions API 00:25:32 What Decisions API Is and How It Works 00:30:24 What OpenAI Is Building With the New APIs 00:32:23 Prompt Caching, Pre-Warming, and API Performance 00:35:20 Context Compaction for Long-Running Agents 00:37:13 Memory, Higher-Level APIs, and the AI Cloud Transcript Introduction: OpenAI DevDay and the New Agent Stack Vibhu [00:00:00]: Okay. We’re very excited to be here. Today is OpenAI DevDay. Special podcast Swyx [00:00:08]: We’re the first podcast after your livestream. Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today? Ari Weinstein [00:00:24]: Yeah. yeah, it was a super exciting day. we just got out of the keynote. It was really sick. there were a bunch of Computer Use announcements that I think are worth thinking about. We have, Dots, which is the new, sort of personal assistant product, and, that has some really exciting Computer Use features. There’s GPT-6.1 Sol, which is this amazing new model, that I think is particularly great for Computer Use ‘cause of, sort of the cost and speed, advantages. I think, I think we shared that it’s, a fifth of the cost of Astra and a seventh of the cost if you’re looking at Computer Use specifically, which is really amazing. sorry, there were so many things. I’m trying to sort through it. Swyx [00:01:02]: And the API. Ari Weinstein [00:01:03]: Agents API, which now has Computer Use in it, which is really cool, ‘cause now developers can build on the same Computer Use, that is part of Codex, and ChatGPT. and then there were some demos of our existing Computer Use features, like app shots, where you can take the context of something you’re doing on your computer and bring it into Codex and ChatGPT really fast. And then, like, native Computer Use on your Mac, where Roman had it taking screenshots of his app, automatically, and he could do other things on his computer while Computer Use was using his applications. so yeah, really exciting keynote. Swyx [00:01:35]: And not to mention the Decisions API. Ari Weinstein [00:01:37]: Decisions API. Swyx [00:01:38]: Off the bat, are they all the same model? Like, this is. Or the same dataset distilled to different models? Swyx [00:01:44]: Like, basically, like, is Computer Use using Decisions API, or are they, like, kinda separate? Ari Weinstein [00:01:49]: So what’s really cool about the Decisions API is it, you know, it has all these new capabilities. It does inference in parallel. it doesn’t have reasoning. It’s a smaller model, than the ones we use for Computer Use. and so those capabilities make it really fast. Dots and Delegating Work to a Cloud Computer Swyx [00:02:07]: Yeah. Ari Weinstein [00:02:07]: They also make it a little bit less good at doing, like, long horizon, sort of sophisticated tasks. And so I think I would say it’s still an open area of research for how we, like, bring those approaches together. But, yeah, I’m really excited to see what people build with the Decisions API. Vibhu [00:02:24]: One of the interesting things is Dots now have attached personal computers. Ari Weinstein [00:02:28]: Yeah. Vibhu [00:02:28]: So it seems like they’re very much more persistent. You’ve been using them for a while. How should people push the bounds? Like, what should people aim for? What should they try? Personally, right now I use it for a lot of customer service. Like Ari Weinstein [00:02:41]: Cool Vibhu [00:02:41]: “Oh, this was wrong. I don’t wanna sign in. I don’t wanna authenticate.” Find whatever and just get it fixed. Ari Weinstein [00:02:45]: Yeah. Vibhu [00:02:46]: How sho

    Don't miss tomorrow's

    The Daily Pulse in your inbox each morning — sourced and linked.

    How often
    Keep going — across the app