# How to Generate Images Locally with FLUX (Python)

> Run FLUX.1-schnell locally on your GPU with a short Python diffusers script — accept the Hugging Face license and offload for low VRAM.

- Canonical: https://guides-ai.pages.dev/guides/generate-images-with-flux-locally/
- Plate 12.05 · Topic: AI images (https://guides-ai.pages.dev/topics/image/)
- Published: 06 Sept 2026 · 4 min read
- Source site: guides-ai — https://guides-ai.pages.dev/

FLUX is a diffusion-transformer image model. Two routes exist locally: ComfyUI's node graph, or a Python script using **diffusers**. This guide uses diffusers — it is one file, it downloads the weights for you, and there is nothing to wire up by hand, which makes it far easier to get working the first time.

We use **FLUX.1-schnell**, the fast Apache-2.0 variant that produces an image in four steps.

## 1. Know your VRAM budget

Loading every FLUX component at once needs about **50 GB** of RAM/VRAM. Nobody does that on a desktop — you offload instead:

| Your GPU | What to use |
| --- | --- |
| 24 GB+ | `enable_model_cpu_offload()` (the script below) |
| 4–32 GB | `enable_sequential_cpu_offload()` plus VAE slicing/tiling — slower, much less VRAM |
| No GPU | Technically possible on CPU, but minutes per image. Use a hosted API instead. |

## 2. Get access to the weights

FLUX.1-schnell is a **gated** repo: open its Hugging Face model page, accept the license, then authenticate.

```bash
pip install -U huggingface_hub
hf auth login
```

Stores a token so `from_pretrained` can fetch the weights. Skipping this gives a 401 on download, not a helpful message.

## 3. Install

```bash
pip install -U diffusers transformers accelerate torch
```

Installs the pipeline, the text encoders, the offloading helpers, and PyTorch. On an NVIDIA machine, install `torch` from the PyTorch site first if you need a specific CUDA build.

## 4. Run it

```python
import torch
from diffusers import FluxPipeline

pipe = FluxPipeline.from_pretrained(
    "black-forest-labs/FLUX.1-schnell", dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()

image = pipe(
    prompt="A cat holding a sign that says hello world",
    guidance_scale=0.0,
    height=768,
    width=1360,
    num_inference_steps=4,
    max_sequence_length=256,
).images[0]
image.save("image.png")
```

Downloads about 34 GB of weights on the first run (the transformer is ~24 GB, the T5 text encoder ~9.5 GB), caches them in `~/.cache/huggingface`, then writes `image.png` next to the script. Budget the disk space before you start.

Three values are not optional on schnell: `guidance_scale` must be `0.0`, `max_sequence_length` must not exceed `256`, and four steps is the point of this model — raising it mostly costs time.

Tight on VRAM? Replace the offload line with:

```python
pipe.enable_sequential_cpu_offload()
pipe.vae.enable_slicing()
pipe.vae.enable_tiling()
```

Moves layers to the CPU one at a time and decodes the image in tiles, which fits GPUs down to about 4 GB at the cost of speed.

## 5. Your first prompt

FLUX responds to plain descriptive sentences, not keyword soup. Say the subject, the setting, the lighting, and the framing — and it is unusually good at rendering readable text in the image, so quote any words you want to appear.

---

Next: [write prompts that actually work](/guides/midjourney-prompts-that-work/) or [edit part of an image with inpainting](/guides/ai-image-inpainting/).
