FLUX 3
AssessPlatforms
A multimodal foundation model for generating and understanding images, video, audio, and language.
Why it's here
Placed in Assess: 3 article(s) of evidence from 2 source(s), led by open-source activity, with 2 in the last 30 days. Confidence 50%.
Evidence (3)
- 7Hacker News·7/24/2026researchFLUX 3 x mimic adds action prediction to video generation
Black Forest Labs says an early version of FLUX 3, its multimodal foundation model, is now running on robots through a collaboration with mimic robotics. The company reports that adding action prediction briefly reduced video quality, but the model later recovered while retaining the new capability, positioning physical AI as an extension of the same backbone.
- 8Hacker News·7/24/2026model_releaseFlux 3 Early Access Released
FLUX 3 is a new multimodal foundation model that jointly learns from images, videos, audio, and language in a single architecture. The company says it can generate and transform video with native audio, support text-to-video and image-to-video workflows, and has shown early preference results over competing video models in preliminary evaluations.
- 6Hugging Face Blog·3/5/2026open_sourceHugging Face introduces Modular Diffusers for composable diffusion pipelines
Hugging Face announced Modular Diffusers, a new way to build diffusion workflows from reusable blocks instead of writing full pipelines from scratch. The system works alongside DiffusionPipeline, supports custom blocks, and can be integrated with the Mellon visual workflow interface.