Skip to content

Latest commit

 

History

2,658 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

sd-scripts

English / 日本語

Table of Contents

Click to expand

Introduction

This repository contains training, generation and utility scripts for Stable Diffusion and other image generation models.

Sponsors

We are grateful to the following companies for their generous sponsorship:

AiHUB Inc.

Support the Project

If you find this project helpful, please consider supporting its development via GitHub Sponsors. Your support is greatly appreciated!

Change History

  • Version 0.12.0 (2026-09-24):

    • Added support for Windows on ARM64 (e.g. NVIDIA RTX Spark PCs). PR #2430, PR #2431, PR #2433
      • opencv-python is now optional (a Pillow/NumPy fallback is used when it is missing), and requirements.txt selects the packages that have Windows ARM64 wheels automatically. For details, please refer to Installing without OpenCV / Windows on ARM64.
      • transformers, schedulefree and safetensors in requirements.txt have been updated to versions that provide Windows ARM64 wheels.
    • Updated the dependencies in requirements.txt: transformers 4.57.6 -> 5.5.4, diffusers 0.32.1 -> 0.40.0, accelerate 1.6.0 -> 1.15.0, huggingface-hub 0.34.3 -> 1.32.0. PR #2436
      • This is mainly a security maintenance update (the 4.x line of transformers and diffusers < 0.38 no longer receive fixes). The previous versions of the libraries still work with this release, so you do not have to update them immediately, but it is recommended to run pip install --upgrade -r requirements.txt at your earliest convenience.
      • diffusers 0.40 requires PyTorch 2.6 or later (PyTorch 2.6.0 or later has been the requirement of sd-scripts already). CI now tests with PyTorch 2.6.0 and 2.8.0.
      • In transformers 5.x, CLIPTokenizer no longer applies the ftfy text normalization of the original CLIP tokenizer (straightening curly quotes, converting full-width characters, etc.). sd-scripts now applies it itself, so tokenization is unchanged from previous versions.
      • Text encoder outputs, VAE outputs and the noise schedulers were verified to be identical to the previous versions with the local regression tests in tests/local.
      • diffusers 0.40 prints a spurious warning "There are modules in AutoencoderKL that should be kept in float32: [] ..." on every .to(dtype) call (a bug in diffusers: the check fires even when the list is empty). sd-scripts suppresses this warning when the list is empty.
    • Added support for transformers 5.6 and later, and updated transformers in requirements.txt to 5.17.0. PR #2437
      • transformers 5.6 changed the internal structure of CLIPTextModel (the text_model submodule was removed). sd-scripts now wraps the model so that the checkpoint keys, the LoRA weight names of the text encoders (lora_te_text_model_...) and text_encoder.text_model.* access stay the same as before. Nothing changes for transformers < 5.6, where the wrapper is not applied.
      • transformers 5.6 also switched the attention implementation of T5 (T5-XXL of FLUX.1 / SD3, byT5 of HunyuanImage) to SDPA, which should be faster and use less memory. The bf16/fp16 outputs of T5 differ very slightly from previous versions (cosine similarity ≈ 0.998 for T5-XXL, the accuracy against fp32 is the same). This may change generated images or trained weights in minor details, and cached text encoder outputs from previous versions are still usable. The outputs of the other text encoders are identical.
      • As above, the previous versions of transformers still work, but updating with pip install --upgrade -r requirements.txt is recommended.
    • Added OFTv2 and BOFT network modules (networks.oft_v2, networks.boft) for SD1.x / SD2.x / SDXL training. PR #2357
      • Orthogonal fine-tuning adapters following the PEFT implementation. Weights in PEFT format can also be loaded. Thanks to umisetokikaze.
      • Note that --network_dim means the block size for these modules. For details, please refer to the documentation.
    • Added per-subset timestep sampling offset (custom_attributes.timestep_sampling.offset) for FLUX.1 and Anima LoRA training. PR #2401 Thanks to okdsf.
      • Shifts the timestep sampling distribution of each dataset subset toward lower- or higher-noise timesteps. For details, please refer to the documentation.
    • Added --show_timesteps_offset to preview the timestep distribution with the offset applied when using --show_timesteps. PR #2410
      • The documentation also describes how the offset behaves with shift / flux_shift timestep sampling.
    • Removed gen_img_diffusers.py, the old image generation script for SD1.x / SD2.x. It had not worked for a while (it depended on a function removed by the refactoring) and gen_img.py supports everything it did except the experimental CLIP / VGG16 guidance. Please use gen_img.py instead (see gen_img_README.md). The file is still available in the previous releases. PR #2439
    • Fixed the dpmsolver and dpmsingle samplers of --sample_sampler (sample image generation during training) and --sampler of gen_img.py / sdxl_gen_img.py, which failed with an error on recent versions of diffusers. PR #2438
      • The lms / k_lms samplers require the scipy package, which is not included in requirements.txt. A clear error message is now shown at startup (instead of at the first sample generation) when scipy is missing. Please run pip install scipy to use them.
  • Version 0.11.1 (2026-06-16):

    • Added support for torch.compile in Anima LoRA/LLLite training. PR #2379
      • It seems to speed up training by about 20%. It requires Triton and MSVC compiler. For details, please refer to the documentation.
    • Added 2D-only Qwen-Image VAE. PR #2382
      • Based on the suggestion by woct0rdho in issue #2369. Thanks to woct0rdho.
      • Enabled by specifying --qwen_image_vae_2d. The weights are the same as the standard (3D) version.
      • Expected to speed up latent pre-caching (training itself remains unchanged). For details, please refer to the documentation.
    • Added support for LLLite inpainting model training. PR #2378
    • Added logging of timestep sampling settings and visualization of timesteps distribution. PR #2384
      • Visualization makes it easier to understand how training is conducted at different timesteps.
      • For details, please refer to the documentation.
  • Version 0.11.0 (2026-06-12):

    • A major internal refactoring of the codebase has been performed to improve code quality and maintainability. PR #2372
      • We have made efforts to minimize direct impact on users. For details and bug reports, please refer to this discussion.
  • Version 0.10.6 (2026-06-12):

    • Stable version before refactoring merge.
  • Version 0.10.5 (2026-05-08):

    • Support for transformers version 5 and later has been added. Thanks to marcus165090-spec for PR #2315 (followed by PR #2316).
      • The transformers version in requirements.txt remains 4.x, but it also works with 5.x. If you use 5.x for any reason, please also update diffusers to the latest version.
    • Support for ControlNet-LLLite training for Anima has been added. Thanks to PR #2317.
  • Version 0.10.4 (2026-05-07):

    • Improved compatibility with Intel GPUs. Thanks to WhitePr for PR #2307.
    • Support for training inpainting models for SD 1.5/SDXL has been added. Thanks to allanoepping for PR #2309 (followed by PR #2318).

Supported Models

  • Stable Diffusion 1.x/2.x
  • SDXL
  • SD3/SD3.5
  • FLUX.1
  • LUMINA
  • HunyuanImage-2.1
  • Anima

Features

  • LoRA training
  • Fine-tuning (native training, DreamBooth): except for HunyuanImage-2.1
  • Textual Inversion training: SD/SDXL
  • Inpainting model training: SD1.5 and SDXL
  • Image generation
  • Other utilities such as model conversion, image tagging, LoRA merging, etc.

Documentation

Training Documentation (English and Japanese)

Other Documentation (English and Japanese)

For Developers Using AI Coding Agents

This repository provides recommended instructions to help AI agents like Claude and Gemini understand our project context and coding standards.

To use them, you need to opt-in by creating your own configuration file in the project root.

Quick Setup:

  1. Create a CLAUDE.md and/or GEMINI.md file in the project root.

  2. Add the following line to your CLAUDE.md to import the repository's recommended prompt:

    @./.ai/claude.prompt.md

    or for Gemini:

    @./.ai/gemini.prompt.md
  3. You can now add your own personal instructions below the import line (e.g., Always respond in Japanese.).

This approach ensures that you have full control over the instructions given to your agent while benefiting from the shared project context. Your CLAUDE.md and GEMINI.md are already listed in .gitignore, so they won't be committed to the repository.

Windows Installation

Windows Required Dependencies

Python 3.10.x and Git:

Python 3.11.x, and 3.12.x will work but not tested.

Give unrestricted script access to powershell so venv can work:

  • Open an administrator powershell window
  • Type Set-ExecutionPolicy Unrestricted and answer A
  • Close admin powershell window

Installation Steps

Open a regular Powershell terminal and type the following inside:

git clone https://github.com/kohya-ss/sd-scripts.git
cd sd-scripts

python -m venv venv
.\venv\Scripts\activate

pip install torch==2.6.0 torchvision==0.21.0 --index-url https://download.pytorch.org/whl/cu124
pip install --upgrade -r requirements.txt

accelerate config

If python -m venv shows only python, change python to py.

Note: bitsandbytes, prodigyopt and lion-pytorch are included in the requirements.txt. If you'd like to use another version, please install it manually.

This installation is for CUDA 12.4. If you use a different version of CUDA, please install the appropriate version of PyTorch. For example, if you use CUDA 12.1, please install pip install torch==2.6.0 torchvision==0.21.0 --index-url https://download.pytorch.org/whl/cu121.

Answers to accelerate config:

- This machine
- No distributed training
- NO
- NO
- NO
- all
- fp16

If you'd like to use bf16, please answer bf16 to the last question.

Note: Some user reports ValueError: fp16 mixed precision requires a GPU is occurred in training. In this case, answer 0 for the 6th question: What GPU(s) (by id) should be used for training on this machine as a comma-separated list? [all]:

(Single GPU with id 0 will be used.)

About requirements.txt and PyTorch

The file does not contain requirements for PyTorch. Because the version of PyTorch depends on the environment, it is not included in the file. Please install PyTorch first according to the environment. See installation instructions below.

The scripts are tested with PyTorch 2.6.0. PyTorch 2.6.0 or later is required.

For RTX 50 series GPUs, PyTorch 2.8.0 with CUDA 12.8/12.9 should be used. requirements.txt will work with this version.

Installing without OpenCV / Windows on ARM64

opencv-python is listed in requirements.txt, but the core training / dataset pipeline only uses a small subset of OpenCV (mainly cv2.resize, cv2.cvtColor, and a debug-only cv2.imshow). When opencv-python is not available, a lightweight Pillow/NumPy fallback under library/_cv2_stub is automatically registered as cv2, so existing scripts continue to work. If you would rather avoid the large OpenCV install, simply uninstall it after installing the requirements:

pip uninstall opencv-python

On Windows on ARM64 (e.g. NVIDIA RTX Spark PCs), opencv-python has no prebuilt wheel, so requirements.txt skips it automatically through an environment marker and the usual pip install --upgrade -r requirements.txt works as is. For the same reason tensorboardX is installed instead of tensorboard on that platform (TensorBoard 2.x depends on grpcio, which has no Windows ARM64 wheel). Logging with --log_with tensorboard works unchanged through tensorboardX; view the logs with TensorBoard on another machine.

Note that:

  • The default install with OpenCV remains the recommended path. The fallback reproduces OpenCV's INTER_AREA and INTER_LINEAR resizing (the modes the dataset pipeline uses by default) in NumPy, so training results match up to rounding, but it is slower than OpenCV (roughly 0.1 s per 24-megapixel image). INTER_CUBIC / INTER_LANCZOS4 go through Pillow and differ slightly.
  • The following tools still require real opencv-python and will exit with a clear message when it is missing: tools/canny.py, tools/detect_face_rotate.py, and the ControlNet canny preprocessor used by gen_img.py / sdxl_gen_img.py.
  • Debug-only features such as cv2.imshow during dataset inspection fall back to Pillow's default image viewer (PIL.Image.show), and cv2.waitKey blocks on input() in the terminal so you can page through images one at a time.

xformers installation (optional)

To install xformers, run the following command in your activated virtual environment:

pip install xformers --index-url https://download.pytorch.org/whl/cu124

Please change the CUDA version in the URL according to your environment if necessary. xformers may not be available for some GPU architectures.

Linux/WSL2 Installation

Linux or WSL2 installation steps are almost the same as Windows. Just change venv\Scripts\activate to source venv/bin/activate.

Note: Please make sure that NVIDIA driver and CUDA toolkit are installed in advance.

DeepSpeed installation (experimental, Linux or WSL2 only)

To install DeepSpeed, run the following command in your activated virtual environment:

pip install deepspeed==0.16.7 

Upgrade

When a new release comes out you can upgrade your repo with the following command:

cd sd-scripts
git pull
.\venv\Scripts\activate
pip install --use-pep517 --upgrade -r requirements.txt

Once the commands have completed successfully you should be ready to use the new version.

Upgrade PyTorch

If you want to upgrade PyTorch, you can upgrade it with pip install command in Windows Installation section.

Credits

The implementation for LoRA is based on cloneofsimo's repo. Thank you for great work!

The LoRA expansion to Conv2d 3x3 was initially released by cloneofsimo and its effectiveness was demonstrated at LoCon by KohakuBlueleaf. Thank you so much KohakuBlueleaf!

License

The majority of scripts is licensed under ASL 2.0 (including codes from Diffusers, cloneofsimo's and LoCon), however portions of the project are available under separate license terms:

Memory Efficient Attention Pytorch: MIT

bitsandbytes: MIT

BLIP: BSD-3-Clause

About

No description, website, or topics provided.

Resources

Stars

7.2k stars

Watchers

55 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages