FLUX 3 logo

FLUX 3

Visit

Black Forest Labs multimodal model for video, audio, upcoming images, and action-prediction, with up to 20-second clips, native audio, and draft mode.

Share:

FLUX 3

FLUX 3 is Black Forest Labs' multimodal generation model for video, audio, upcoming image generation, and action-prediction. The company describes it as one model for multiple modalities rather than a video-only successor to FLUX.1. Official pages at bfl.ai emphasize stylistically diverse video with optional native audio, clips up to 20 seconds in a single generation, multilingual speech, and an agent-friendly draft-then-render workflow. Image generation is listed as "soon" on the product page.

Key Features

  • One model, several modalities: Video and audio are available now; image generation is marked soon; action-prediction targets robotics by turning visual observations and text instructions into predicted outcomes and control actions.
  • Video controls: Text-to-video, image-to-video, start/end and ordered keyframes, multiple scenes and camera angles in one take, video continuation from the last frame, and agentic chaining of clips.
  • Native audio: Optional multilingual speech, effects, and ambience generated with the frames, plus lipsync-focused dialogue.
  • Draft mode: Generate a cheap, fast preview, then send the same prompt back for a full-quality render.
  • Self-Flow architecture: BFL presents Self-Flow as its approach for aligning multimodal generation and understanding in one underlying model, with charts comparing generation error and manipulation-task success against flow matching.
  • API, playground, and open weights path: The site offers a playground, a production API, and a separate open-weights licensing track. FLUX 3's own open-weight availability should be checked on the current model page before you assume downloadable checkpoints.

Use Cases

  • Agentic video pipelines: Nous Research's Dillon Rolnick describes Hermes generating a shot, checking it, then continuing so 20-second clips become longer pieces.
  • Product and brand video: Magnific and Envato quotes on the official page highlight motion design, product, and on-brand creative work, not only cinematic film.
  • Storyboarding and advertising: Multiple scenes, typography, and multilingual dialogue support explainer, documentary, and campaign workflows.
  • Robotics research: Action-prediction is positioned for perception, simulation, and execution, not consumer video.

Pricing

Black Forest Labs uses pay-as-you-go API pricing with no seat fees. The public calculator on bfl.ai/pricing (accessed 2026-08-16) listed FLUX 3 video at $0.17 per second in HD draft for a text/image-to-video job, or $0.85 for a 5-second clip at that selected rate. Resolution, duration, draft versus normal, and video-to-video options change the total, so treat $0.17/s as the calculator's displayed draft HD rate, not a universal list price.

Open-weights licensing on the same page is still framed around FLUX.2 families (Builder, Platform, Professional, Enterprise, Synthetic Data). Enterprise API customers can ask sales for volume, SLA, and combo pricing.

Getting Started

  1. Open the FLUX 3 page and try the playground.
  2. Create an API key from the BFL dashboard if you need production volume.
  3. Start in draft mode to explore prompts cheaply, then render the selected draft at full quality.
  4. For longer stories, continue from the last frame or chain clips instead of asking for one long take.
  5. If you need on-prem weights, read the current open-weights terms. Do not assume FLUX 3 checkpoints are in the FLUX.2 klein/dev packages.

Frequently Asked Questions

Is FLUX 3 only a video model?

No. Video and audio are the current public surface. Image generation is listed as coming soon, and action-prediction is a robotics-oriented modality.

Can I download the weights?

BFL sells open-weights licenses for some FLUX models. The FLUX 3 marketing page talks about API, playground, and "download our models," but the public pricing page still details FLUX.2 klein and dev packages. Verify the exact FLUX 3 weight offering before planning a local deploy.

How long can a clip be?

Official copy says up to 20 seconds in a single generation. Longer pieces are expected to come from continuation and agentic chaining.

Alternatives

  • Flux.1 Pro: BFL's earlier commercial image model, still the reference if you only need stills.
  • Veo 3: Google's video generation product if you already live in Gemini.
  • MiniMax: another multimodal vendor with video and audio APIs.

Tips & Best Practices

  1. Draft first: the official workflow is explore cheap, then render. Do not start on full quality for prompt search.
  2. Write for shots, not films: 20-second clips plus continuation match how the model is demonstrated.
  3. Check audio separately: multilingual speech and lipsync are a feature, but they also add failure modes. Preview dialogue before you chain a long sequence.

Conclusion

FLUX 3 is BFL's bet that video, audio, images, and robot action can share one multimodal backbone. If you need agent-controllable clips with native sound, start in the official playground, keep generations short, and only then wire the API into a longer pipeline.

Comments

No comments yet. Be the first to comment!