FLUX is a diffusion-transformer image model. Two routes exist locally: ComfyUI’s node graph, or a Python script using diffusers. This guide uses diffusers — it is one file, it downloads the weights for you, and there is nothing to wire up by hand, which makes it far easier to get working the first time.
We use FLUX.1-schnell, the fast Apache-2.0 variant that produces an image in four steps.
1. Know your VRAM budget
Loading every FLUX component at once needs about 50 GB of RAM/VRAM. Nobody does that on a desktop — you offload instead:
| Your GPU | What to use |
|---|---|
| 24 GB+ | enable_model_cpu_offload() (the script below) |
| 4–32 GB | enable_sequential_cpu_offload() plus VAE slicing/tiling — slower, much less VRAM |
| No GPU | Technically possible on CPU, but minutes per image. Use a hosted API instead. |
2. Get access to the weights
FLUX.1-schnell is a gated repo: open its Hugging Face model page, accept the license, then authenticate.
pip install -U huggingface_hub
hf auth login
Stores a token so from_pretrained can fetch the weights. Skipping this gives a 401 on download, not a helpful message.
3. Install
pip install -U diffusers transformers accelerate torch
Installs the pipeline, the text encoders, the offloading helpers, and PyTorch. On an NVIDIA machine, install torch from the PyTorch site first if you need a specific CUDA build.
4. Run it
import torch
from diffusers import FluxPipeline
pipe = FluxPipeline.from_pretrained(
"black-forest-labs/FLUX.1-schnell", dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
image = pipe(
prompt="A cat holding a sign that says hello world",
guidance_scale=0.0,
height=768,
width=1360,
num_inference_steps=4,
max_sequence_length=256,
).images[0]
image.save("image.png")
Downloads about 34 GB of weights on the first run (the transformer is ~24 GB, the T5 text encoder ~9.5 GB), caches them in ~/.cache/huggingface, then writes image.png next to the script. Budget the disk space before you start.
Three values are not optional on schnell: guidance_scale must be 0.0, max_sequence_length must not exceed 256, and four steps is the point of this model — raising it mostly costs time.
Tight on VRAM? Replace the offload line with:
pipe.enable_sequential_cpu_offload()
pipe.vae.enable_slicing()
pipe.vae.enable_tiling()
Moves layers to the CPU one at a time and decodes the image in tiles, which fits GPUs down to about 4 GB at the cost of speed.
5. Your first prompt
FLUX responds to plain descriptive sentences, not keyword soup. Say the subject, the setting, the lighting, and the framing — and it is unusually good at rendering readable text in the image, so quote any words you want to appear.
Next: write prompts that actually work or edit part of an image with inpainting.