FLUX GPU requirements: VRAM, licences and RTX PRO cards for on-premise image generation
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- FLUX.1 [dev] and [schnell] have 12 billion parameters and a full checkpoint of about 23 GB; NVIDIA puts the FLUX.1 Kontext [dev] model at 24 GB, reduced to 12 GB in FP8 and 7 GB in FP4
- FLUX.2 [dev] has 32 billion parameters and needs 90 GB to load completely, by NVIDIA’s figure; FP8 cuts that by 40 per cent, about 54 GB by our arithmetic, which fits the 72 GB RTX PRO 5000 or the 96 GB RTX PRO 6000
- Memory decides which model and resolution fit; the time per image follows compute, so FP8 and FP4 speed it up: one FLUX.1 Kontext step takes 607 ms in BF16 and 254 ms in FP4 on an RTX PRO 6000, per NVIDIA, at a resolution it does not state
- FLUX.1 [schnell], FLUX.2 [klein] 4B and Qwen-Image are Apache 2.0; FLUX.1 [dev], FLUX.2 [dev] and [klein] 9B are non-commercial, and the Stable Diffusion 3.5 licence ends above the revenue threshold it sets
- NVIDIA lists one RTX PRO 6000 Server Edition at 0.20 FLUX images per second in FP4 at batch 1, about 720 an hour, and an L40S at 0.08 in FP8; neither figure states resolution or steps
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
FLUX GPU requirements in short
For on-premise image generation, the FLUX GPU requirements start from the model size. FLUX.1 [dev] and FLUX.1 [schnell] are 12-billion-parameter models whose full checkpoint is about 23 GB in 16-bit, so a 24 GB card runs them in FP8 or FP4. For the 16-bit model with the 16-bit T5 text encoder, ComfyUI recommends more than 32 GB of VRAM, which means a 48 GB card or larger. FLUX.2 [dev], with 32 billion parameters, needs 90 GB to load completely by NVIDIA’s figure, and about 54 GB in FP8 by our arithmetic, which points to the 72 GB RTX PRO 5000 or the 96 GB RTX PRO 6000. Memory decides which model and which resolution fit on a card. The time per image depends on compute, because a diffusion model runs its whole transformer once for every denoising step.
An LLM generates one token after another, and for a single user its speed follows memory bandwidth, as our RTX PRO 6000 benchmark overview explains. An image model processes the whole latent image in each step, so lower precision shortens the step as well as reducing the memory. NVIDIA measured one denoising step of FLUX.1 Kontext [dev] on the RTX PRO 6000 Blackwell at 607 ms in BF16, 317 ms in FP8 and 254 ms in FP4 (NVIDIA Technical Blog, 2 July 2025). The blog does not state the resolution.
The step count multiplies that time. BFL’s model card says FLUX.1 [schnell] produces images “in only 1 to 4 steps”, and the Diffusers example for FLUX.2 [klein] 4B uses 4. For FLUX.1 [dev], the Diffusers code sets 28 steps by default, while the same page says the guidance-distilled variant “takes about 50 sampling steps for good-quality generation”.
FLUX.1 also loads two text encoders, CLIP-L and T5-XXL, and a VAE that turns the latent image into pixels. ComfyUI’s FLUX tutorial recommends the 16-bit T5 encoder “when your VRAM is greater than 32GB” and offers an FP8 encoder for smaller cards. NVIDIA names Mistral Small 3 as the text encoder of FLUX.2 [dev] (January 2026). Higher resolutions need more working memory, and FLUX.2 generates images of “up to 4 megapixel resolution”, NVIDIA says.
Offloading to system memory runs large models on small cards at the cost of speed: the Diffusers documentation calls CPU offloading “extremely slow”, and NVIDIA says ComfyUI’s weight streaming works “albeit with some performance loss”. For a team that generates all day, choose a card that holds the model and its encoders at the chosen precision.
Open image models: parameters, licences and memory
| MODEL | PARAMETERS | LICENCE | MEMORY AS STATED |
|---|---|---|---|
| FLUX.1 [schnell] | 12B | Apache 2.0 | about 23 GB for the original FLUX.1 files (ComfyUI); FP8 checkpoint available |
| FLUX.1 [dev], Kontext [dev] | 12B | FLUX.1 [dev] Non-Commercial License | Kontext: 24 GB, 12 GB in FP8, 7 GB in FP4 (NVIDIA) |
| FLUX.2 [dev] | 32B | FLUX Non-Commercial License | 90 GB to load completely; FP8 reduces this by 40 per cent (NVIDIA) |
| FLUX.2 [klein] 4B | 4B | Apache 2.0 | about 13 GB (BFL) |
| FLUX.2 [klein] 9B | 9B | FLUX Non-Commercial License | not stated |
| Stable Diffusion 3.5 Large | 8B | Stability AI Community License | over 18 GB; 11 GB in FP8 (NVIDIA) |
| Qwen-Image | 20B | Apache 2.0 | not stated; 40 GB of BF16 weights for the transformer by our arithmetic |
Hugging Face model cards of black-forest-labs, stabilityai and Qwen, read on 10 October 2026; ComfyUI FLUX.1 tutorial; NVIDIA blogs of 12 June 2025 (SD 3.5), 13 August 2025 (Kontext NIM) and 25 November 2025 (FLUX.2). The Qwen-Image figure is 20 billion parameters at 2 bytes, without its text encoder.
The memory figures measure different things: NVIDIA’s FLUX.2 figure covers loading the whole pipeline, the Kontext figures the model size in each precision. Add a margin for the resolution you plan and for a second model kept loaded.
FLUX licences: which models a company may run in production
For a company, check the licence before sizing the GPU. FLUX.1 [schnell], FLUX.2 [klein] 4B and Qwen-Image are under Apache 2.0. BFL’s card for [schnell] says that under it “the model can be used for personal, scientific, and commercial purposes”, and the [klein] 4B card calls its weights “available for commercial use under the Apache 2.0 license”.
FLUX.1 [dev] is under the FLUX.1 [dev] Non-Commercial License v1.1.1. It grants use “solely for your Non-Commercial Purposes”, excludes use “for revenue-generating activity” from that purpose, and forbids use “for any commercial or production purposes”. For the generated images the licence says: “You may use Output for any purpose (including for commercial purposes), except as expressly prohibited herein.” The model card for FLUX.2 [dev] says BFL released the weights “under a non-commercial license to support third-party research and development”, and the [klein] 9B models fall under the same FLUX Non-Commercial License. For running these models in a company’s own workflow, BFL offers self-hosted commercial licences under its FLUX Commercial Weights License, in tiers with a monthly image allowance; its licensing page lists FLUX.2 [dev] and the [klein] models and does not name FLUX.1 [dev] in a tier.
Stable Diffusion 3.5 is under the Stability AI Community License, last updated on 5 July 2024 on Stability AI’s site and on Hugging Face. It covers people and organisations whose annual revenue stays below a threshold set in the licence. Above that threshold, “any licenses granted to You under this Agreement shall terminate”, and a licence must be requested from Stability AI. Commercial use also requires registration with Stability AI, and you own the outputs “to the extent permitted by applicable law”.
The FLUX.1 [dev] licence bars using outputs to train a model “competitive with” FLUX.1 [dev], and the Stability licence bars using them to create or improve “any foundational generative AI model” other than its own models. Whether a given internal use counts as commercial under these licences is a legal assessment for the company’s legal department.
FP8, FP4 and TensorRT on RTX PRO Blackwell
NVIDIA’s blog on FLUX.1 Kontext reports “approximately 2x and 3x memory savings” for the transformer from BF16 to FP8 and to FP4, with the quantisation done by TensorRT Model Optimizer and the engines built with TensorRT. NVIDIA’s blog on the Kontext NIM gives the FP8 size for “NVIDIA Ada Generation GPUs” and the FP4 size for the “NVIDIA Blackwell architecture”. Among the cards we supply, Blackwell means the RTX PRO 2000 to 6000 range and the DGX Spark, while the L40S, the L4 and the RTX Ada generation cards run FP8. NVIDIA’s SD 3.5 blog says “the Ada Lovelace generation of NVIDIA RTX PRO GPUs support FP8 quantization”.
For Stable Diffusion 3.5 Large, NVIDIA and Stability AI moved from over 18 GB to 11 GB with FP8, and NVIDIA says “FP8 TensorRT boosts SD3.5 Large performance by 2.3x vs. BF16 PyTorch”. For FLUX.2, NVIDIA says FP8 checkpoints reduce “the VRAM requirements by 40% at comparable quality”. BFL publishes ONNX exports of FLUX.1 [dev] in “BF16, FP8, and FP4 precision”, under the same non-commercial licence.
NVIDIA publishes side-by-side images for these formats, not quality metrics, so generate the same prompts and seeds in BF16 and in the quantised format and let the people who use the images judge before you size for FP4.
Images per minute: published figures for RTX PRO and L40S
NVIDIA’s AI inference performance page, dated 5 October 2026 in its metadata, lists one RTX PRO 6000 Server Edition at 0.20 FLUX images per second in FP4 at batch 1, with a latency of 5.0 seconds, and one L40S at 0.08 images per second in FP8, 12.5 seconds per image. Neither row names the FLUX variant, the resolution or the step count, and our benchmark overview lists NVIDIA’s Stable Diffusion XL and video rows next to them.
By our arithmetic, 0.20 images per second is 12 a minute and 720 an hour per card. From NVIDIA’s per-step Kontext figures, an edit at 28 steps spends about 7.1 seconds in the transformer in FP4, 8.9 in FP8 and 17 in BF16 on an RTX PRO 6000, before text encoding and VAE decoding. For the DGX Spark, NVIDIA measured FLUX.1 [schnell] in FP4 at 1024×1024 with 4 steps at 23 images per minute (24 October 2025).
Our RTX PRO 6000 vs RTX 5090 comparison covers single-image speed on both cards and the GeForce driver licence clause on data-centre deployment.
Which GPU for which team: configurations
Size an image server in four steps; the table gives our estimates from the figures above.
- Choose the models the licence allows for your use, and the precision you accept after a quality check.
- Take the memory from the table above, add the text encoders and a margin for the largest resolution, and pick the card that holds every model you keep loaded.
- Estimate the seconds per image at your step count and resolution, and divide the images per peak hour by the images one card delivers per hour.
- Check the waiting time: at batch 1, ten people who submit at once wait for ten images in turn, so the peak number of people generating together often sets the card count before the daily volume does.
| TEAM USE | MODELS | CONFIGURATION | BASIS OF ESTIMATE |
|---|---|---|---|
| One designer, evaluation | FLUX.2 [klein] 4B, FLUX.1 in FP8, SD 3.5 Large in FP8 | workstation with an RTX PRO 4500, 32 GB, or RTX PRO 4000, 24 GB | 11 to 13 GB models, FP4 option on Blackwell |
| Design team, licensed [dev] | FLUX.1 [dev] or FLUX.2 [dev] in FP8 | RTX PRO 5000, 72 GB, or RTX PRO 6000 Workstation or Max-Q, 96 GB | FLUX.2 in FP8 about 54 GB; unquantised at 90 GB only on the 96 GB card, with little margin |
| Several users, small models | [klein] 4B, SD 3.5 Large in FP8 | one RTX PRO 6000 Server Edition split by MIG into four 24 GB instances | one isolated instance per user or workflow |
| Production, thousands a day | FLUX in FP4 or FP8 | server with 2 to 4 RTX PRO 6000 Server Edition, one instance per card | 720 images an hour per card at NVIDIA’s batch-1 figure |
| Existing L40S servers | FLUX.1 in FP8, SD 3.5 Large | L40S, 48 GB | 12.5 s per FLUX image in FP8 (NVIDIA) |
| Model evaluation on a desk | FLUX.2 [dev] unquantised | one DGX Spark, 128 GB | 90 GB model loads; 23 [schnell] images a minute in FP4 |
Estimates by Eurokommerz from Table 1 and from NVIDIA’s published figures; card memory from NVIDIA’s product pages. Measure your own workflow at your resolution and step count before ordering several cards.
As a worked example, a marketing team of 20 that generates 2,000 images on a busy day, half of them in one peak hour, needs 1,000 images in that hour. At NVIDIA’s 720 images an hour per card that is two RTX PRO 6000 cards, and our estimate rises in proportion if your step count or resolution is higher than NVIDIA’s unstated test. Our RTX PRO 6000, 5000 and 4500 comparison has the memory, bandwidth and MIG profiles of each tier.
We supply every card in the table and build servers and workstations around them. Tell us the models, the resolution and the images per day in the form below, and we reply with a configuration within one business day.
ComfyUI and serving on a shared GPU server
ComfyUI, which BFL and NVIDIA both name for running FLUX.2 locally, is under GPL-3.0, lists Flux.1, Flux.2, SD3.5 and Qwen Image among its supported image models, and describes “asynchronous queueing”. On a team server, one ComfyUI instance per GPU or per MIG instance keeps one long job from holding up everyone else.
MIG on the RTX PRO 6000 gives up to four isolated 24 GB instances, enough for FLUX.2 [klein] 4B or Stable Diffusion 3.5 Large in FP8 each, and needs no licence on bare metal. Four Max-Q cards in one tower give a design team 384 GB without a server room; our four-card Max-Q workstation guide covers the power supply and the roughly 2 kW of heat.
We build AI servers with two or four RTX PRO 6000 Server Edition cards and install the operating system, drivers, CUDA and a container runtime on request. Describe your image workflows in the form below, with the number of people generating at peak.
What we supply
We supply the cards this article discusses, the RTX PRO 4000, 4500, 5000 (48 and 72 GB) and 6000 in all three editions, the L40S and the DGX Spark, with manufacturer warranty on one EU contract and invoice. The GPU page lists each card’s memory and power. We build AI servers to order for image generation, sized by the models, the resolution and the peak number of users, assembled and burn-in tested, and we check the rack, power and airflow before we quote. NVIDIA AI Enterprise and vGPU licences come on the same invoice where the setup needs them.
FAQ
What GPU do I need for FLUX?
How much VRAM does FLUX.1 dev need?
How much VRAM does FLUX.2 need?
How much VRAM does Stable Diffusion 3.5 Large need?
Can FLUX be used commercially?
How many images per minute does an RTX PRO 6000 generate with FLUX?
Send us the image models you plan to run, the resolution, the number of images a day and how many people generate at the same time. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day