§12.05

How to Generate Images Locally with FLUX (Python)

Run FLUX.1-schnell locally on your GPU with a short Python diffusers script — accept the Hugging Face license and offload for low VRAM.

published 06 Sept 2026 checked against docs 06 Sept 2026 4 min in AI images Markdown

On this page5 sections
  1. 1. Know your VRAM budget
  2. 2. Get access to the weights
  3. 3. Install
  4. 4. Run it
  5. 5. Your first prompt

FLUX is a diffusion-transformer image model. Two routes exist locally: ComfyUI’s node graph, or a Python script using diffusers. This guide uses diffusers — it is one file, it downloads the weights for you, and there is nothing to wire up by hand, which makes it far easier to get working the first time.

We use FLUX.1-schnell, the fast Apache-2.0 variant that produces an image in four steps.

1. Know your VRAM budget

Loading every FLUX component at once needs about 50 GB of RAM/VRAM. Nobody does that on a desktop — you offload instead:

Your GPUWhat to use
24 GB+enable_model_cpu_offload() (the script below)
4–32 GBenable_sequential_cpu_offload() plus VAE slicing/tiling — slower, much less VRAM
No GPUTechnically possible on CPU, but minutes per image. Use a hosted API instead.

2. Get access to the weights

FLUX.1-schnell is a gated repo: open its Hugging Face model page, accept the license, then authenticate.

pip install -U huggingface_hub
hf auth login

Stores a token so from_pretrained can fetch the weights. Skipping this gives a 401 on download, not a helpful message.

3. Install

pip install -U diffusers transformers accelerate torch

Installs the pipeline, the text encoders, the offloading helpers, and PyTorch. On an NVIDIA machine, install torch from the PyTorch site first if you need a specific CUDA build.

4. Run it

import torch
from diffusers import FluxPipeline

pipe = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-schnell", dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()

image = pipe(
    prompt="A cat holding a sign that says hello world",
    guidance_scale=0.0,
    height=768,
    width=1360,
    num_inference_steps=4,
    max_sequence_length=256,
).images[0]
image.save("image.png")

Downloads about 34 GB of weights on the first run (the transformer is ~24 GB, the T5 text encoder ~9.5 GB), caches them in ~/.cache/huggingface, then writes image.png next to the script. Budget the disk space before you start.

Three values are not optional on schnell: guidance_scale must be 0.0, max_sequence_length must not exceed 256, and four steps is the point of this model — raising it mostly costs time.

Tight on VRAM? Replace the offload line with:

pipe.enable_sequential_cpu_offload()
pipe.vae.enable_slicing()
pipe.vae.enable_tiling()

Moves layers to the CPU one at a time and decodes the image in tiles, which fits GPUs down to about 4 GB at the cost of speed.

5. Your first prompt

FLUX responds to plain descriptive sentences, not keyword soup. Say the subject, the setting, the lighting, and the framing — and it is unusually good at rendering readable text in the image, so quote any words you want to appear.


Next: write prompts that actually work or edit part of an image with inpainting.

← All AI images plates · Search all guides

↑↓ move↵ openalt+↵ copy first command

Keyboard

⌘/ctrl+K or /
Search all guides
alt+↵
In search: copy the guide's first command
j / k
Move through a list of guides
c
On a guide: copy its first command
t
Toggle light / dark
?
This list