FLUX 3
FLUX 3 is Black Forest Labs' multimodal generation model for video, audio, upcoming image generation, and action-prediction. The company describes it as one model for multiple modalities rather than a video-only successor to FLUX.1. Official pages at bfl.ai emphasize stylistically diverse video with optional native audio, clips up to 20 seconds in a single generation, multilingual speech, and an agent-friendly draft-then-render workflow. Image generation is listed as "soon" on the product page.
Key Features
- One model, several modalities: Video and audio are available now; image generation is marked soon; action-prediction targets robotics by turning visual observations and text instructions into predicted outcomes and control actions.
- Video controls: Text-to-video, image-to-video, start/end and ordered keyframes, multiple scenes and camera angles in one take, video continuation from the last frame, and agentic chaining of clips.
- Native audio: Optional multilingual speech, effects, and ambience generated with the frames, plus lipsync-focused dialogue.
- Draft mode: Generate a cheap, fast preview, then send the same prompt back for a full-quality render.
- Self-Flow architecture: BFL presents Self-Flow as its approach for aligning multimodal generation and understanding in one underlying model, with charts comparing generation error and manipulation-task success against flow matching.
- API, playground, and open weights path: The site offers a playground, a production API, and a separate open-weights licensing track. FLUX 3's own open-weight availability should be checked on the current model page before you assume downloadable checkpoints.
Use Cases
- Agentic video pipelines: Nous Research's Dillon Rolnick describes Hermes generating a shot, checking it, then continuing so 20-second clips become longer pieces.
- Product and brand video: Magnific and Envato quotes on the official page highlight motion design, product, and on-brand creative work, not only cinematic film.
- Storyboarding and advertising: Multiple scenes, typography, and multilingual dialogue support explainer, documentary, and campaign workflows.
- Robotics research: Action-prediction is positioned for perception, simulation, and execution, not consumer video.
Pricing
Black Forest Labs uses pay-as-you-go API pricing with no seat fees. The public calculator on bfl.ai/pricing (accessed 2026-08-16) listed FLUX 3 video at $0.17 per second in HD draft for a text/image-to-video job, or $0.85 for a 5-second clip at that selected rate. Resolution, duration, draft versus normal, and video-to-video options change the total, so treat $0.17/s as the calculator's displayed draft HD rate, not a universal list price.
Open-weights licensing on the same page is still framed around FLUX.2 families (Builder, Platform, Professional, Enterprise, Synthetic Data). Enterprise API customers can ask sales for volume, SLA, and combo pricing.
Getting Started
- Open the FLUX 3 page and try the playground.
- Create an API key from the BFL dashboard if you need production volume.
- Start in draft mode to explore prompts cheaply, then render the selected draft at full quality.
- For longer stories, continue from the last frame or chain clips instead of asking for one long take.
- If you need on-prem weights, read the current open-weights terms. Do not assume FLUX 3 checkpoints are in the FLUX.2 klein/dev packages.
Frequently Asked Questions
Is FLUX 3 only a video model?
No. Video and audio are the current public surface. Image generation is listed as coming soon, and action-prediction is a robotics-oriented modality.
Can I download the weights?
BFL sells open-weights licenses for some FLUX models. The FLUX 3 marketing page talks about API, playground, and "download our models," but the public pricing page still details FLUX.2 klein and dev packages. Verify the exact FLUX 3 weight offering before planning a local deploy.
How long can a clip be?
Official copy says up to 20 seconds in a single generation. Longer pieces are expected to come from continuation and agentic chaining.
Alternatives
- Flux.1 Pro: BFL's earlier commercial image model, still the reference if you only need stills.
- Veo 3: Google's video generation product if you already live in Gemini.
- MiniMax: another multimodal vendor with video and audio APIs.
Tips & Best Practices
- Draft first: the official workflow is explore cheap, then render. Do not start on full quality for prompt search.
- Write for shots, not films: 20-second clips plus continuation match how the model is demonstrated.
- Check audio separately: multilingual speech and lipsync are a feature, but they also add failure modes. Preview dialogue before you chain a long sequence.
Conclusion
FLUX 3 is BFL's bet that video, audio, images, and robot action can share one multimodal backbone. If you need agent-controllable clips with native sound, start in the official playground, keep generations short, and only then wire the API into a longer pipeline.
Comments
No comments yet. Be the first to comment!
Related Tools
Flux.1 Dev
blackforestlabs.ai
Black Forest Labs' open-source non-commercial image generation model, maintaining Pro-level image quality, fully open source.
Flux.1 Pro
blackforestlabs.ai
Black Forest Labs' most powerful commercial image generation model, surpassing Midjourney and DALL-E 3 in image quality.
Google: Gemini 3.7 Flash
gemini.google.com
Google's most capable Flash model for agentic workflows and multimodal reasoning, with 1M-token context, 65k output tokens, and Computer Use preview.
Related Insights
What Locks Codex Is the Picker, Not the Models
You already pay for OpenCode Go, Grok, and Z.ai, but the Codex picker still shows mostly GPT. The community project codex-router does not teach another install ritual. It puts subscriptions you already bought back into the selector, keeps keys on the machine, and leaves native GPT alone.
Seven AI Coding CLIs, Six Months: No Matter How Strong the Model, Work Needs Supervision
Claude Code, Codex, opencode, pi, omp and DeepSeek Harness all have personalities. After six months of deep use I run a division of labor: pi for the fastest cheapest reviews, omp for complex PRs, DeepSeek Harness on V4 Flash for high-frequency low-cost review, and Claude Code, Qoder and Cursor for writing. No matter how strong the model, work needs supervision — ideally from an independent third party.
After I Connected Obsidian to OpenClaw, It Started Helping Me Make Decisions
Once Obsidian stopped being just a place to store notes and started working with OpenClaw, it began helping me organize context, connect information, and improve real decisions.