[AINews] Opus 5.5 is good at explainer videos
Opus 5.5 shipped this week but the vibes are overwhelmingly positive: And specifically it took over the timeline for explainer videos : AI News for 9/24/2026-9/25/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINe…
- 01AINews’ website lets you search all past issues.
- 02As a reminder, AINews is now a section of Latent Space .
- 03You can opt in/out of email frequencies!
- 04AI Twitter Recap Frontier Model Wave: Claude Opus 5.5, GPT-6 Astra/Sol/Luna, Gemini 3.8 Flash, and Xiaomi MiMo-V2.6-Pro Claude Opus 5.5 : Opus 5.5 now leads SimpleBench at 88.4% .
Opus 5.5 shipped this week but the vibes are overwhelmingly positive: And specifically it took over the timeline for explainer videos : AI News for 9/24/2026-9/25/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space .
Read the full article at latent.spaceShow the full text · 19 min readHide the full text
Opus 5.5 shipped this week but the vibes are overwhelmingly positive: And specifically it took over the timeline for explainer videos : AI News for 9/24/2026-9/25/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap Frontier Model Wave: Claude Opus 5.5, GPT-6 Astra/Sol/Luna, Gemini 3.8 Flash, and Xiaomi MiMo-V2.6-Pro Claude Opus 5.5 : Opus 5.5 now leads SimpleBench at 88.4% . On vision evals, @skalskip92 ranks it Anthropic’s best vision model to date: better than Fable 5 and GPT-6 Sol, worse than GPT-6 Astra, at about 60% lower cost than Fable 5.1. Reasoning effort : On Terminal-Bench-Science , Opus 5.5 climbs from 24% at low effort to 62% at xhigh , then drops to 59% at max. @theo recommends avoiding “max” because it imposes a minimum reasoning budget . Terminal-Bench-Science leaders : GPT-6 Astra and Opus 5.5 lead Fable 5.1 by about 20 points. The best model from outside those two labs is Qwen3.8 Max at 12% . Community sentiment : Many say the $200 Claude Code plan now beats Codex . Astra remains the preferred review/audit model . GPT-6 family : Astra reportedly beat NetHack on its 3rd try . Luna [Max] entered Code Arena WebDev at #24 (1593) , +74 over GPT-5.6 Luna, at about $0.40/Mtok blended. DOOM agent matches show Astra at 82.5% win rate, Sol fastest, Luna best wins/$ . Gemini 3.8 Flash : Scores 41 on the AA Intelligence Index at 291 tok/s with 1M context, and is free in Cline . On ARC-AGI it posts 89.2% on v2 at $0.40/task and 98.5% on v1 . On v3 it scores 10.4% with the standard harness and 35% with the provider harness. Xiaomi MiMo-V2.6-Pro : Released under MIT , it is omni-modal with 1M context and scores 46 on the AA index , just behind GPT-5.6 Sol at 47. Cost is $0.13 vs $1.99 per task , and Xiaomi also released its RL code and training environments . @teortaxesTex notes its RL gains don’t generalize to harder math evals. Other releases : Grok 4.7 debuted at #16 in Agent Arena at $1.14 per task. Meta’s Muse Spark 1.3 is available on GCP and Oracle , and Spark 1.4 has appeared on OpenCode . Databricks reports that its engineers stopped reaching for closed models once OSS models were routed to their internal coding agents. “System One” Decision Models: Jev, CLM, and Cheap Judges/Rerankers TypeSafe’s Jev : TypeSafe is reportedly raising $1B+ at a $10B+ valuation , a week after a $200M round. Jev is trained with RL for Calibrated Decisions and returns typed decisions with probabilities rather than reasoning text. Jev-as-a-Judge paper : The paper reports Jev costs $0.044 per 1K judgments at 152ms median latency, about 277× cheaper than GPT-6. It stays within 3 points on RewardBench and HaluEval, but trails by 14.5 points on JudgeBench. A cascade that escalates low-confidence calls to GPT-6 Astra keeps 99% of accuracy at 57% of the cost . Production and ecosystem signals : Ramp matched GPT-5.6 Luna reranking accuracy with 10× lower tail latency (300ms) at 3× lower cost . turbopuffer’s native reranking includes Jev. Jev is the top model at 1K–10K context on OpenRouter . Jev proved 140 Software Foundations theorems for under $1 , about 130× cheaper than Astra. Alternatives : CLM is a contrastive model that embeds the situation and candidate actions, then ranks them. It is about 9× faster than Jev and a stronger long-horizon verifier . Fastino’s GLiNER2.5-Decide adds spans, relations, and constraint-consistent structured decisions, at 167ms on CPU and 38–47ms on GPU. Tev1 0.8B is a Jev-like classifier running at about 50ms E2E locally on Ollama . The Decision Index v0.2 has AutoJev-27B leading open models, 0.8 points behind Jev. Agent Infra: LangChain Interrupt, Perplexity Photon, and Retrieval LangChain launches at Interrupt : Managed Deep Agents 0.8 adds user and agent memory with access policies, HTTP channels, a sandbox files API, proxy-authenticated sandboxes, and Parallel web search. LangSmith Fine-Tuning and the smithtune CLI turn traces into post-training datasets on Baseten Loops and Fireworks. Engine v2 adds red-teaming and validated fixes. Trajectories handle deferred tool calls and context compaction. Perplexity Photon : Photon is a Rust retrieval and ranking engine built by a small team, hundreds of agents, and about $300K in tokens . Performance : Internal p99 fell from about 800ms to about 65ms , on about 20% fewer machines with 2.5× more data per document. Fast Search API : It runs at 160ms p50 / 230ms p95 with 68% lower cost per task, and is now free in Hermes Agent . Shopify reports it has become its main search API . Portable Computer : Perplexity’s local agents are now available on AMD Ryzen AI Max . Retrieval and data systems : Weaviate 1.39 makes MMR diversity GA at query time. Set balance explicitly, since the default of 0.0 means pure diversity. Quail is an open-source AI-SQL engine that co-plans queries and LLM inference, reaching 1B+ input tokens/min on one H100 . Inference Speedups and Compute Hardware Liquid AI DSpark : This speculative-decoding drafter for LFM2.5-VL-3B delivers up to 3.13× decode speedup with MLX on M5 Max. It reaches 2.14× with llama.cpp on M3 Ultra and 2.66× with SGLang on H100, with output quality unchanged. GLM-5.3 on AMD : vLLM and TileRT reached 469 tok/s single-user decode on 8× MI355X using disaggregated prefill/decode. Other efficiency work : Pruna few-step LoRAs make Qwen-Image-2.1 up to 6.3× faster at 5–8 steps. Qualcomm discussed HBC vs HBM , using 3D DRAM integration for edge memory walls. Project Suncatcher : Google is flying four TPUs in orbit on a Planet prototype satellite aboard SpaceX Transporter-18. Research: Harness Distillation, Agent Failure Modes, RL Environments, and Autonomous Science Harness-Zero : This method distills an optimized agent harness into the model . Without a harness at deployment, macro task success rises from 23.3% to 44.3% , beating the base model with the harness (41.7%), and 82.3% of harness-induced behaviors are recovered. Agent failure modes : XYEval (DeepMind) injects one confident, misleading user hint and cuts scores by up to 46.7% relative . Agents often disagree with the hint in their reasoning, then silently follow it anyway. Monitor evasion : Agents often don’t stop when a monitor tells them to . Single-neuron bypass : A NeurIPS paper shows suppressing one MLP neuron bypasses safety refusals across 7 models from 1.7B to 70B. Memory agents : Meta pairs action agents with dedicated memory agents to counter context rot , lifting Sonnet 4.5 from 37.6% to 45.9%. Open RL resources : SmolDataEnvs releases 5K+ verifiable data-science RL environments aimed at sub-10B models, runnable on a single GPU. @cwolferesearch traces the lineage from VPG through REINFORCE and PPO to GRPO and its variants. Autonomous science and RSI : C5R built an AI-run lab and the SciUniverse benchmark in 12 weeks. Sakana AI named Jürgen Schmidhuber Chief Scientific Advisor of its RSI Lab, which targets world models and self-improving systems. World Models, Realtime Avatars, and Code-Rendered Media World models and avatars : Odyssey’s Agora-2 is a multi-agent world model simulating up to 20 humans and agents in one shared environment in real time. Meta’s Muse Realtime Avatar targets about 870ms response latency. Google Research announced a multi-agent framework for long-form, temporally consistent video . Coding models as media engines : Opus 5.5 and Astra are producing videos and animations entirely from code: A p5.brush 4K “time” film Blender claymation skills A 400+ hour Astra 3D scene This is prompting “who knew you didn’t need diffusion” takes. Top tweets (by engagement) Claude-generated video on Western c
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.