All AI News
    Latent SpaceFriday, July 24, 2026 20 min read
    AI

    Black Forest Labs FLUX 3: Multimodal Flow Models That Beat Seedance 2.0, Gemini Omni, and Grok Imagine, Plus a FLUX-Mimic Video-Action Robotics Model

    Black Forest Labs enters video generation with multimodal FLUX 3, claiming SOTA and extending into robotics control.

    Koko brief

    Black Forest Labs enters video generation with multimodal FLUX 3, claiming SOTA and extending into robotics control.

    Black Forest Labs' FLUX 3 Video arrives two years after the company teased the capability at launch. The multimodal system handles text, image, and video inputs with native audio output, multilingual dialogue, and agentic clip chaining—benchmarking above Seedance 2.0, Gemini Omni, and Grok Imagine. More consequentially, the FLUX3-mimic variant pairs the same backbone with dexterous robot learning, suggesting generative video models are becoming viable world models for factory automation.

    Watch: whether the promised open-weights Dev release shifts enterprise video tooling away from closed APIs and toward self-hosted pipelines.

    Thursdays are the heaviest days for AI releases, and even though OpenAI scored a victory over Anthropic in launching the new ChatGPT Voice (consumer) and OpenAI Presence (enterprise) and getting more impressions than Claude Voice today (a completely accidental coincidence in timing, we are sure), neither seem as monumental as BFL’s launch of FLUX 3 Video today: We last covered BFL in our very well received Anjney Midha podcast : $5000 w…","cta":null,"showBylines":true,"showDescription":true,"showImage":true,"size":"sm","isEditorNode":true,"title":"The Professor of Outputmaxxing — Anjney Midha, AMP","publishedBylines":[],"post_date":"2026-06-18T17:30:00.811Z","cover_image":"https://substack-video.s3.amazonaws.com/video_upload/post/202359797/8dbbb3fa-e808-473c-af72-b9aee4fe0026/transcoded-1781652240.png","cover_image_alt":null,"canonical_url":"https://www.latent.space/p/anj","section_name":null,"video_upload_id":null,"id":202359797,"type":"podcast","reaction_count":22,"comment_count":4,"publication_id":1084089,"publication_name":"Latent.Space","publication_logo_url":"https://substackcdn.com/image/fetch/$s_!DbYa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73b0838a-bd14-46a1-801c-b6a2046e5c1e_1130x1130.png","belowTheFold":false,"youtube_url":null,"show_links":null,"feed_url":null}"> Most GenMedia people will remember the BFL homepage when they initially launched Flux 1 in 2024, hinting at video models next , with their logo in a forest. Well, 2 years later, it’s finally real: The blogpost outlines Self Flow, covering ALL their modalities together with strong preference claims: “ Its core capabilities include the following ( all outputs come with native audio generation ): Text-to-video generation. Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references. Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context. Generative video-audio continuation from input video and audio. Keyframe-to-video generation for controlled transitions between defined moments. Multilingual dialogue. A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output. Agentic chaining of individual clips into longer, multi-shot sequences. High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics. Strong typography generation and animated designs.” Some of the above are SOTA features from other frontier lab models, like we discussed in our Grok Imagine pod , so the community has very much been put on notice that there has now been independent, perhaps SOTA, reproduction of these capabilities, with an open weights Dev version on the way. As if this release wasn’t enough, the team also announced FLUX3-mimic , which proves that the FLUX 3 model is learning a sufficient world model capable of driving robots… @mimicrobotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous","username":"bfl_ai","name":"Black Forest Labs","profile_image_url":"https://pbs.substack.com/profile_images/1954888731053142016/NDyG-4-j_normal.jpg","date":"2026-07-23T15:08:16.000Z","photos":[],"quoted_tweet":{},"reply_count":2,"retweet_count":7,"like_count":138,"impression_count":14471,"expanded_url":null,"video_url":null,"video_preview_media_key":null,"belowTheFold":true}" data-component-name="Twitter2ToDOM"> … and predicting their impact in real factory settings… AI News for 7/22/2026-7/23/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies! AI Twitter Recap Open Code, Open Models, and the Policy Fault Line Around Distillation The Stack v3 is the day’s most consequential open-data release : @anton_lozhkov announced The Stack v3 , now the largest open code dataset publicly released: 114 TB raw , 224M repositories , 44B files , 770 languages , and roughly 5T deduplicated/filtered tokens . Relative to v2, the filtered corpus jumps from ~ 550B to ~5T tokens , with especially large gains in C++ (x15) , TypeScript (x7.5) , Rust (x7) , and Python (x4.8) . The notable operational changes are that v3 ships contents inline rather than Software Heritage IDs, includes a fresh GitHub recrawl through Aug 2025, excludes restrictively licensed code, and offers both a ready-to-train split and a full bucket for custom dedup/filtering. Hugging Face researchers framed it explicitly as infrastructure for the next generation of open code models and cyber-defense tooling: see @LoubnaBenAllal1 , @lvwerra , and commentary from @eliebakouch noting prior Stack versions were used in many disclosed code-model training mixtures. Distillation remains the live ideological fault line : several high-signal posts pushed back on attempts to sharply separate “internet-scale pretraining” from output-level distillation. @GergelyOrosz compared model inspection via prompting to reverse-engineering a competitor’s product, while @SchmidhuberAI emphasized distillation’s long lineage. @Suhail argued the practical response is not prohibition but stronger investment in open-weight domestic models , and @garrytan put it more simply: open weights are strategically important. The subtext across these posts is that open datasets like The Stack v3 materially raise the floor for every lab that wants to build competitive code models without relying on closed ecosystems. Multimodal Frontier: FLUX 3, Robotics Transfer, and New Audio/TTS Systems Black Forest Labs’ FLUX 3 expands the multimodal frontier beyond image/video : @bfl_ai launched FLUX 3 , a unified multimodal model spanning image, video, audio, and action prediction , with early access for FLUX 3 Video and an explicit claim that the same architecture can be extended toward robotics. Team members connected it back to the earlier Self-Flow research, including @hila_chefer and @robrombach . What matters technically is the unified training story: not a loose family of specialized generators, but one architecture intended to bridge media generation and control. mimic’s FLUX-mimic is a concrete robotics instantiation of that thesis : @mimicrobotics described FLUX-mimic as a Video-Action Model built on top of FLUX 3 , trained on robot and wearable data for general-purpose dexterity and deployable on a single on-prem GPU . Their central claim is that better video world modeling transfers directly into robot control quality and sample efficiency; they’re already testing with Audi . This dovetails with @GeneralistAI , whose GEN-1 now supports varied end effectors and can adapt when the “hand” changes mid-rollout, reinforcing the idea that embodiment-general policies may come from conditioning on morphology rather than specializing per manipulator. Audio saw two notable launches at opposite ends of the stack : @Alibaba_Qwen introduced Qwen-Audio-3.0-TTS in Flash and Plus variants, with 16 languages , inline control tags like [whisper] /

    Key takeaways
    • 01Black Forest Labs' FLUX 3 Video arrives two years after the company teased the capability at launch.
    • 02The multimodal system handles text, image, and video inputs with native audio output, multilingual dialogue, and agentic clip chaining—benchmarking above Seedance 2.0, Gemini Omni, and Grok Imagine.
    • 03More consequentially, the FLUX3-mimic variant pairs the same backbone with dexterous robot learning, suggesting generative video models are becoming viable world models for factory automation.
    Keep going — across the app