Back to news
AI Tools & Products
Jul 24, 2026

FLUX 3 Launches in Early Access as a New Multimodal Foundation Model

Jul 24, 2026
AI Summary

FLUX 3, a new multimodal foundation model, is now available in Early Access. It integrates learning from images, videos, and audio to create a unified understanding of the world, enhancing capabilities in content creation and physical AI.

FLUX 3 is designed to learn from multiple modalities—images, videos, and audio—simultaneously, improving the model's understanding of the world.

The model can generate diverse videos with audio up to 20 seconds long and is capable of synthesizing and editing images in various styles and resolutions.

Preliminary evaluations show that FLUX 3 outperformed several existing models in various comparisons, with a preference rate of up to 93% over Luma Ray 3.2.

Key strengths include capturing human facial expressions, associating sounds with physical events, and supporting multilingual output.

The model also integrates action prediction and has been tested in collaboration with partners like mimic robotics for applications in dexterous manipulation.

Further capabilities will be rolled out in phases, with ongoing improvements expected during the early access period. The development team is also focused on creating the next generation of models that unify perception, action, and language prediction.

fluxai toolsmachine learningsoftwaretechnology