Skip to main content
    All AI News
    Latent SpaceThursday, October 1, 2026 12 min read
    AI

    Gemini 4 Argon: Google DeepMind's Answer to Astra/Fable, with 1M Output

    Gemini 4 Argon claims 13/19 benchmark wins and an industry-first 1M-token output limit—but ships gated to cybersecurity partners first.

    Illustration: Gemini 4 Argon: Google DeepMind's Answer to Astra/Fable, with 1M Output
    AI illustration generated from this story — not a photograph of the event.
    Key takeaways
    • 01Google DeepMind's Gemini 4 Argon re-enters the frontier tier after eight months, topping 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5.
    • 02Its headline feature, Long Decode Continuation, extends output to 1M tokens across chained API calls.
    • 03Pricing opens at $4/$20 per million tokens with a 50% launch discount.
    • 04Internal deployments include migrating 800K lines of C++ to Rust and—per the team—advancing a formal math conjecture.
    Koko brief

    Gemini 4 Argon claims 13/19 benchmark wins and an industry-first 1M-token output limit—but ships gated to cybersecurity partners first.

    Google DeepMind's Gemini 4 Argon re-enters the frontier tier after eight months, topping 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5. Its headline feature, Long Decode Continuation, extends output to 1M tokens across chained API calls. Pricing opens at $4/$20 per million tokens with a 50% launch discount. Internal deployments include migrating 800K lines of C++ to Rust and—per the team—advancing a formal math conjecture. Hallucination rates beat Astra's, though raw accuracy trails it. - Watch: Whether gated rollout signals genuine safety caution or competitive benchmark-gaming before broader scrutiny arrives.

    Watch: Whether gated rollout signals genuine safety caution or competitive benchmark-gaming before broader scrutiny arrives.

    In brief · from latent.space

    GDM last shipped a larger-than-Flash model in February ( 3.1 Pro ), and after successive incremental 3.x Flash versions and the big GDM management shakeup last month, the largest question for GDM was when they would catch up to peers who have in the meantime launched Fable and Astra class models.

    Read the full article at latent.space
    Show the full text · 12 min read

    GDM last shipped a larger-than-Flash model in February ( 3.1 Pro ), and after successive incremental 3.x Flash versions and the big GDM management shakeup last month, the largest question for GDM was when they would catch up to peers who have in the meantime launched Fable and Astra class models. Well, Argon’s here , with VERY respectable benchmarks (SOTA in 13 of 19 credible benchmarks)… but only accessible in limited cybersecurity preview, though access is promised “ as soon as possible ”: We like the experimental Long Decode Continuation , which increases output tokens up to 1M as an industry first. AI News for 9/29/2026-9/30/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap Gemini 4 Argon: Google Returns to the Frontier Launch : Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense ( @GoogleDeepMind , @sundarpichai ). Availability : Access starts with government users and trusted cyber defenders in the Fairwind Program. Google says it will refine guardrails before opening access to developers, enterprises and consumers ( @Google , @demishassabis ). Output limit : Google cites an industry-leading 1M-token output limit, up from 64K ( @GoogleAI , @TheRundownAI ). Measurement note : Vals lists 262K max output. Artificial Analysis reached 1M output tokens through Long Decode Continuation, a new API feature that pauses long responses and resumes them across calls ( @ValsAI , @ArtificialAnlys ). Pricing : Standard pricing is $4/$20 per 1M input/output tokens. A 50% introductory discount brings it to $2/$10, with no end date announced. Cached input gets a 95% discount ( @_philschmid , @ArtificialAnlys ). Google’s claimed results : Argon takes first place on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5. On DeepSWE it scores 77.9%, versus 74.2% for Opus 5.5 and 74.1% for Astra ( @TheRundownAI ). Internal deployments : Google reports that Argon agents freed more than 300 TiB of data-center memory and are migrating more than 800K lines of C/C++ kernel code to Rust ( @kimmonismus ). Video decoder : Agents replaced 32K lines of SIMD code with safe Rust, making the existing Rust port 2.7x faster with identical output. Research use : The team says internal agent loops built on Argon helped complete the CK conjecture ( @mirrokni ). Artificial Analysis evaluation : Argon scores 53 on the Intelligence Index, matching GPT-6 Astra (53) and edging GPT-6.1 Sol (52) ( @ArtificialAnlys ). Cost per task : At discounted pricing it costs $1.99 per task versus $3.26 for Astra; standard pricing would raise this to $3.98. Token use : The savings come from price, not efficiency. Argon averages 62K output tokens per task against Astra’s 27K. Agentic work : It ranks #1 on AutomationBench-AA at 77.5% and scores 57% on Terminal Bench 4, behind Sonnet 5.5, Opus 5.5 and Astra. Hallucination : Its 15% rate on AA-Omniscience compares with 51% for Astra. The tradeoff is lower accuracy: 50% versus Astra’s 63% ( @aipulseda1ly ). Vals evaluation : Argon is #1 on the Vals Index at 68.9%, at an average $15.68 per task ( @ValsAI , @ValsAI ). Coding : It built 30 Vibe Code Bench apps perfectly, against 25 for Opus 5 and 24 for Astra ( @ValsAI ). Terminal and security : Terminal-Bench 4.0 rose from 19.0% to 57.6%. It scores 70% on CyberBench proof-of-concept tasks and 100% on IOI 2024–2026 ( @ValsAI ). Efficiency : It uses about a quarter of Sonnet 5.5’s output tokens on Vals Index tasks ( @ValsAI ). Arena and other evals : Argon is #1 in Text Arena at 1525 and #8 in Code Arena WebDev at 1679 ( @arena ). Agent Arena : It ranks #8 overall and #1 for steerability on a preliminary 3K sessions ( @arena ). PostTrainBench : It scores 45.3%, up from 21.99% for Gemini 3.1 Pro ( @karinanguyen ). Skepticism : Some observers questioned the published numbers. Legal benchmark : Argon’s reported 19.6% on Harvey’s legal benchmark trails Muse Spark 1.2’s listed 25.42% ( @BlackHC ). Other critiques : Commentators raised possible preference-data benchmaxxing and objected to some figures, including DeepSWE ( @teortaxesTex , @teortaxesTex ). GPT-6.1 Sol and OpenAI’s DevDay Agent Stack Independent evals : GPT-6.1 Sol is the new #1 on MathArena ( @j_dekoninck ). Code Arena : It ranks #3 on WebDev at 1759, 70 points above GPT-6 Sol for the same $2/$10 pricing ( @arena ). Cost per task : Artificial Analysis measures $0.72 per task at max effort, versus $3.26 for Astra and $1.04 for GPT-6 Sol ( @ArtificialAnlys ). Source of savings : Sol uses fewer turns and has a lower cache-read price ( @ArtificialAnlys ). Luna bug fix : OpenAI fixed an image-encoding bug, adding 1 Intelligence Index point to GPT-6 Luna. Ultrafast inference : OpenAI quotes up to 300 tok/s. SemiAnalysis reports it runs on NVIDIA GPUs at low batch sizes, not on Cerebras ( @kimmonismus ). Hands-on report : Generation is about 8x faster, but end-to-end agent tasks speed up only 2–4x because tool latency dominates ( @sayashk ). Computer use : Gains are largest here, since UI actions respond in milliseconds. Cost : The tester exhausted a weekly limit in about 2 hours. Product layer : DevDay introduced dots (persistent agents with their own cloud computers), a Decisions API and computer use ( @latentspacepod ). Sites : ChatGPT Sites can now host MCP servers and turn them into installable plugins ( @mxstbr ). Usage limits : Users report one-off credits worth about $2,500. Others complain that usage limits were cut ( @kimmonismus , @kimmonismus ). Other Releases: Embeddings, Image/Video and Open Models Perplexity contextual embeddings : pplx-embed-v2-context-9b-preview is open on Hugging Face ( @perplexity_ai ). Method : The model encodes the whole document once and pools chunk vectors afterward. Training distills relevance from a context-compression model instead of using single gold-chunk labels ( @denisyarats ). Results : It sets a new state of the art on ConTEB. On turbopuffer’s private context-bench it beats voyage-context-4 by 14.4 points in answer recall@10, using 1 KB int8 vectors against 8 KB ( @turbopuffer ). Cohere Embed 5 : The family has Pro and Fast variants in a shared embedding space, so you can index with one and retrieve with the other ( @cohere ). Fast tier : Cohere says it beats other fast-tier models by at least 6 points at a third less cost than Pro. Evaluation uses its new RCP-nDCG@10 metric ( @cohere ). Ideogram 4.5 : The editing model targets artifact-free multi-turn edits, with open weights promised ( @ideogram_ai ). Edit fidelity : Over ten consecutive edits, 94–99% of untouched content stays identical ( @fal ). Ranking : It is #18 in Image Edit Arena at 1351 ( @arena ). Video benchmark : Artificial Analysis launched AA-Video-T2V v2.0, judged at 1080p with more than 68K human votes ( @ArtificialAnlys ). Leaders : Wan 3.0 is #1 at $12/min. Seedance 2.5 is #2 at $34.12/min, and MiniMax H3 is statistically tied at $4.80/min. Utopai X : This post-train of MiniMax H3 debuts at #2 ( @ArtificialAnlys ). Open and small models : Ling-3.1-flash : A 500B model reported close to GPT-5.6 Sol and Opus 5 ( @kimmonismus ). It ranks #2 among open-weight models in Mobile App Arena ( @DesignArena ). Praxis-1 : Runway released an open-weight world-action model and says robotics policy performance scales predictably with third-person video ( @agermanidis ). Solar Mini 4 : Upstage reports 35B total / 3B active parameters. It scores 24 on the Intelligence Index at $0.10/$0.40 ( @ArtificialAnlys ). Caching penalty : It still costs about 5x Luna per task, because only 48% of its repeated context hits cache versus 99% for Luna ( @ArtificialAnlys ). Agent Research, Inference and Systems Context Language Model

    Don't miss tomorrow's

    The Daily Pulse in your inbox each morning — sourced and linked.

    How often
    Keep going — across the app