Black Forest Labs Launches FLUX 3 Image Generation Model
The new model introduces bounding-box composition, multi-reference editing for up to ten images, and native rendering up to 4K resolution.

Key takeaways
- Black Forest Labs officially released FLUX 3 Image, featuring bounding-box layout control and native rendering up to 4K resolution.
- The model operates via a single API endpoint supporting both image generation and multi-element localized editing without mask modes.
- FLUX 3 Image accepts up to ten reference images and features native web search grounding to improve generation accuracy.
- A commercial weights license is currently available for enterprise infrastructure, with an open-weights release expected in the coming weeks.
Black Forest Labs has released FLUX 3 Image, expanding its multimodal FLUX 3 foundation model family into visual generation and precision editing. Announced on October 1, 2026, the model introduces native bounding-box layout composition, localized pixel editing without manual masks, and direct rendering up to 4K resolution across 15 aspect ratios, according to documentation on Black Forest Labs.
The release represents the visual synthesis counterpart to Black Forest Labs' multimodal architecture, which originally debuted its video capabilities on August 4, 2026, as noted by Morphic. FLUX 3 Image is designed to give creators, developers, and autonomous software agents granular control over spatial placement and graphic layouts directly through text and coordinate prompts.

Bounding-Box Layouts and Spatial Control
A primary architectural addition in FLUX 3 Image is native coordinate awareness. Instead of relying solely on natural language interpretations to determine composition, the model accepts bounding-box coordinates at the end of a prompt. According to the Black Forest Labs API Overview, the system maps layout elements across a standardized 0 to 1000 coordinate grid formatted as [top, left, bottom, right].
Users can write a global caption alongside an element table in JSON format, assigning identifiers such as dome_1 or crowd_1 to specific regions. For autonomous systems, the model is designed to work directly with large language models that generate layout bounding boxes automatically from a single line of instructional text, as detailed by Black Forest Labs.
This structured approach prevents common issues in generative media where subjects drift across frames or crowd scenes lose coherence. Graphic designers can place headlines, focal subjects, and secondary objects precisely where needed for posters, editorial spreads, and promotional assets.
Multi-Reference Synthesis and Pixel-Exact Edits
FLUX 3 Image handles both new generations and surgical image modifications through a unified API endpoint without requiring a dedicated edit mode parameter, as stated in the Black Forest Labs Release Notes. The prompt itself determines whether the model generates a scene from scratch or updates an existing image.

When editing, the model supports localized changes while preserving untouched pixels, Black Forest Labs claims. Users can recolor objects, swap props, resize figures, or remove background items in a single pass. For multi-image workflows, FLUX 3 Image accepts between 1 and 10 reference images ranging from 256 × 256 pixels up to 16 megapixels each, according to the Black Forest Labs Documentation. Prompts can reference source assets by numerical position—for example, taking a specific garment from one reference image and placing it on a character defined in another.
The model also incorporates built-in web and image search grounding by default. When enabled, this grounding feature searches the web before generation to accurately depict real-world entities, places, and objects, though developers can disable it via API flags to prioritize processing speed.
Technical Specifications and Output Capabilities
FLUX 3 Image supports multiple resolution tiers, ranging from lightweight 768-pixel square drafts up to 4K resolutions measuring approximately 5456 × 3072 pixels (16.8 megapixels), as documented by Morphic. Supported aspect ratios span 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, 9:21, and an automatic mode that adapts to the dimensions of an input reference image.
According to partner platform Krea, the underlying foundation is built on Self-Flow architecture, where generation and world understanding are trained jointly. This foundation enables high-accuracy multilingual text rendering inside images, making the system capable of embedding legible signage, labels, and typography into complex compositions.

Pricing, Availability, and Licensing
FLUX 3 Image is accessible immediately through the official Black Forest Labs API and third-party developer platforms including fal.ai. Under the standard API pricing listed in the Black Forest Labs Release Notes, per-image rates are scaled by resolution:
- 768sq: $0.041 per image
- 1k: $0.048 per image
- 2k: $0.100 per image
- 4k: $0.607 per image
As reported by The Decoder, an introductory 50 percent launch discount is available across API endpoints through October 8, 2026. On fal.ai, promotional pricing starts at $0.0205 per image, with 1K outputs temporarily priced at $0.024 before returning to standard rates.
For enterprise infrastructure deployments, Black Forest Labs offers commercial weight licensing to allow self-hosted inference and custom fine-tuning. An open-weights release is expected to arrive in the coming weeks, as reported by The Decoder. Integration with consumer platforms is also underway, with FLUX 3 coming soon to Krea, where users currently generate with FLUX.2 while FLUX 3 rolls out in early access.
Frequently asked questions
What is FLUX 3 Image?
FLUX 3 Image is the image generation and editing component of Black Forest Labs' multimodal FLUX 3 family. It allows users to compose and edit scenes using text prompts, bounding boxes, and reference images.
How do bounding boxes work in FLUX 3 Image?
Users specify element coordinates formatted as [top, left, bottom, right] on a normalized 0 to 1000 grid. The coordinates are appended as a JSON table directly in the text prompt.
How many reference images does FLUX 3 Image support?
The model accepts up to 10 reference images per request, each sized between 256 × 256 pixels and 16 megapixels.
Is FLUX 3 Image open-source?
Not yet. Black Forest Labs currently provides a commercial weights license for self-hosted enterprise deployments, with an open-weights version expected in the coming weeks.
Sources
- FLUX 3 Image: Maximum control over every pixel | Black Forest LabsBlack Forest Labs · Official
- FLUX 3 Image - Black Forest LabsBlack Forest Labs · Official
- FLUX 3 by Black Forest Labs — AI Image Generator | KreaKrea · Jul 23, 2026 · Official
- Black Forest Labs launches Flux 3 Image with multi-step editing that leaves the rest of your picture aloneThe Decoder · Oct 2, 2026
- Release Notes - Black Forest LabsBlack Forest Labs
- Flux 3 Image: bounding-box layouts, precise edits, 4K outputMorphic
- Flux 3 Image (Text to Image) API on fal - Fal.aifal.ai
How this story was made: the newsroom picked it up from Reddit, the-decoder.com and Google News, gathered the full text of the sources above, and drafted it with AI assistance. Every factual claim was then checked against those sources before publishing (44 claims checked). Illustrations marked as AI-generated are not photographs. Spotted an error? Tell us.
Published October 3, 2026 at 00:40 UTC


