# FLUX 3: Black Forest Labs' First Multimodal Frontier Model

> FLUX 3 launched July 23, 2026: Black Forest Labs' first multimodal model, generating image, video (up to 20s with synced native audio) and physical action in one architecture. What it does, how it rolls out, and how it compares to FLUX.2.

FLUX 3 is Black Forest Labs' first truly integrated multimodal model, launched in early access on July 23, 2026. Where FLUX.2 focused on images, FLUX 3 generates across images, video and audio, and even predicts physical action, all in a single architecture. Its headline feature is text-to-video up to 20 seconds long with native, in-sync audio. Here is what it does, how it is rolling out, and how it compares.

## What is FLUX 3?

FLUX 3 is Black Forest Labs' first truly integrated foundation model, unveiled in early access on July 23, 2026. Instead of doing one thing, it works across images, video, audio and physical action prediction inside a single architecture, built on the company's Self-Flow approach for aligning multimodal generation and understanding. The headline capability is text-to-video that produces clips up to 20 seconds long with native audio in sync with the visuals, dialogue, sound effects and ambient noise all matched to what is on screen. Here is how the family breaks down.

| Component | What it does | Availability |
| --- | --- | --- |
| FLUX 3 Video | Text-to-video up to 20s with native synced audio | Early access (API + private weights to partners) |
| FLUX 3 Action / FLUX-mimic | Physical action prediction for robotics | Early access (robotics partners) |
| FLUX 3 Image | Image generation | Rolling out in the coming weeks |
| FLUX 3 Dev | Open-weight version | Planned later in 2026 |

## What can FLUX 3 do?

FLUX 3's leap is that generation and understanding share one model across modalities. In practice that means a single prompt can drive a video with matching audio, an image, or an action sequence, without stitching separate tools together. The video side is the standout: up to 20-second clips with native audio, where dialogue, sound effects and ambient sound stay aligned with the picture, something most text-to-video models still handle as a separate step. The Action variant, FLUX-mimic, extends the same architecture to predicting physical movement for robotics.

## How is FLUX 3 rolling out?

Black Forest Labs staged the launch rather than shipping everything at once. FLUX 3 Video went out first, in early access through an API and private weights to initial partners. FLUX 3 Action / FLUX-mimic reached selected research and commercial robotics partners. FLUX 3 Image, the part most creators will use, is rolling out in the coming weeks. An open-weight FLUX 3 Dev is planned for later in 2026 for developers who want to self-host. So if you cannot access it yet, that is expected, broad availability is still ramping.

## FLUX 3 vs FLUX.2

FLUX.2, released in November 2025, is an image model: strong multi-reference editing, 4-megapixel output and sharp text. FLUX 3 is a different class of product, a multimodal model that adds video with synced audio and physical-action prediction on top of imagery. If you need production image generation today, FLUX.2 is the widely available choice. If you are tracking where AI video is going, FLUX 3's 20-second, audio-synced clips are the headline. For most creators the practical path is to use FLUX.2 images now and adopt FLUX 3 video as its image and open-weight tiers become available.

## How to use FLUX 3

During early access FLUX 3 Video and Action are limited to Black Forest Labs' partners via API and private weights, with FLUX 3 Image arriving for broader use in the coming weeks and an open-weight Dev release later in 2026. If your goal is finished video today, without waiting on early-access slots, FluxNote already turns a prompt into a narrated, captioned video using current top models, and generates FLUX-style images in its Image Studio, so you can produce and publish now and fold in FLUX 3 as it opens up.

## Frequently asked questions

### When did FLUX 3 launch?

Black Forest Labs unveiled FLUX 3 in early access on July 23, 2026. It is their first multimodal foundation model, with FLUX 3 Video and Action available to partners first and FLUX 3 Image rolling out in the coming weeks.

### What can FLUX 3 do?

FLUX 3 generates across images, video and audio and predicts physical action, all in one architecture. Its standout feature is text-to-video up to 20 seconds long with native audio, dialogue, sound effects and ambient sound, kept in sync with the visuals.

### Is FLUX 3 available to the public yet?

Not fully. FLUX 3 launched in staged early access on July 23, 2026: Video and Action went to partners via API and private weights, FLUX 3 Image is rolling out in the coming weeks, and an open-weight FLUX 3 Dev is planned for later in 2026.

### What is the difference between FLUX 3 and FLUX.2?

FLUX.2 (November 2025) is an image model with strong multi-reference editing and 4-megapixel output. FLUX 3 (July 2026) is multimodal, adding video with synced audio and physical-action prediction. FLUX.2 is what most people can use today; FLUX 3 is still ramping availability.

### Can FLUX 3 generate video with sound?

Yes. FLUX 3's headline feature is text-to-video up to 20 seconds with native audio generated in sync with the visuals, including dialogue, sound effects and ambient noise, rather than adding sound as a separate step.

---

Source: https://fluxnote.io/guides/flux-3
