README.md

# AnimateDiff

This repository is the official implementation of [AnimateDiff](https://arxiv.org/abs/2307.04725).

**[AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning](https://arxiv.org/abs/2307.04725)**
</br>
Yuwei Guo,
Ceyuan Yang*,
Anyi Rao,
Yaohui Wang,
Yu Qiao,
Dahua Lin,
Bo Dai
<p style="font-size: 0.8em; margin-top: -1em">*Corresponding Author</p>

<!-- [Arxiv Report](https://arxiv.org/abs/2307.04725) | [Project Page](https://animatediff.github.io/) -->
[![arXiv](https://img.shields.io/badge/arXiv-2307.04725-b31b1b.svg)](https://arxiv.org/abs/2307.04725)
[![Project Page](https://img.shields.io/badge/Project-Website-green)](https://animatediff.github.io/)
[![Open in OpenXLab](https://cdn-static.openxlab.org.cn/app-center/openxlab_app.svg)](https://openxlab.org.cn/apps/detail/Masbfca/AnimateDiff)
[![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/guoyww/AnimateDiff)


## Features
- **[2023/09/25]** Release **MotionLoRA** and its model zoo, **enabling camera movement controls**! Please download the MotionLoRA models (**74 MB per model**, available at [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules) ) and save them to the `models/MotionLoRA` folder. Example:
  ```
  python -m scripts.animate --config configs/prompts/v2/5-RealisticVision-MotionLoRA.yaml
  ```
    <table class="center">
      <tr style="line-height: 0">
      <td colspan="2" style="border: none; text-align: center">Zoom In</td>
      <td colspan="2" style="border: none; text-align: center">Zoom Out</td>
      <td colspan="2" style="border: none; text-align: center">Zoom Pan Left</td>
      <td colspan="2" style="border: none; text-align: center">Zoom Pan Right</td>
      </tr>
      <tr>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/01.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/02.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/02.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/01.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/03.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/04.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/04.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/03.gif"></td>
      </tr>
      <tr style="line-height: 0">
      <td colspan="2" style="border: none; text-align: center">Tilt Up</td>
      <td colspan="2" style="border: none; text-align: center">Tilt Down</td>
      <td colspan="2" style="border: none; text-align: center">Rolling Anti-Clockwise</td>
      <td colspan="2" style="border: none; text-align: center">Rolling Clockwise</td>
      </tr>
      <tr>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/05.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/05.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/06.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/06.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/07.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/07.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/08.gif"></td>
      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/08.gif"></td>
      </tr>
  </table>

- **[2023/09/10]** New Motion Module release! `mm_sd_v15_v2.ckpt` was trained on larger resolution & batch size, and gains noticeable quality improvements. Check it out at [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules) and use it with `configs/inference/inference-v2.yaml`. Example:
  ```
  python -m scripts.animate --config configs/prompts/v2/5-RealisticVision.yaml
  ```
  Here is a qualitative comparison between `mm_sd_v15.ckpt` (left) and `mm_sd_v15_v2.ckpt` (right):
  <table class="center">
      <tr>
      <td><img src="__assets__/animations/compare/old_0.gif"></td>
      <td><img src="__assets__/animations/compare/new_0.gif"></td>
      <td><img src="__assets__/animations/compare/old_1.gif"></td>
      <td><img src="__assets__/animations/compare/new_1.gif"></td>
      <td><img src="__assets__/animations/compare/old_2.gif"></td>
      <td><img src="__assets__/animations/compare/new_2.gif"></td>
      <td><img src="__assets__/animations/compare/old_3.gif"></td>
      <td><img src="__assets__/animations/compare/new_3.gif"></td>
      </tr>
  </table>
- GPU Memory Optimization, ~12GB VRAM to inference
- User Interface: [Gradio](#gradio-demo), A1111 WebUI Extension [sd-webui-animatediff](https://github.com/continue-revolution/sd-webui-animatediff) (by [@continue-revolution](https://github.com/continue-revolution))
- Google Colab: [Colab](https://colab.research.google.com/github/camenduru/AnimateDiff-colab/blob/main/AnimateDiff_colab.ipynb) (by [@camenduru](https://github.com/camenduru))


## Model Zoo
<details open>
<summary>Motion Modules</summary>

  | Name                 | Parameter | Storage Space |
  |----------------------|-----------|---------------|
  | mm_sd_v14.ckpt       | 417 M     | 1.6 GB        |
  | mm_sd_v15.ckpt       | 417 M     | 1.6 GB        |
  | mm_sd_v15_v2.ckpt    | 453 M     | 1.7 GB        |

</details>

<details open>
<summary>MotionLoRAs</summary>

  | Name                                 | Parameter | Storage Space |
  |--------------------------------------|-----------|---------------|
  | v2_lora_ZoomIn.ckpt                  | 19 M      | 74 MB         |
  | v2_lora_ZoomOut.ckpt                 | 19 M      | 74 MB         |
  | v2_lora_PanLeft.ckpt                 | 19 M      | 74 MB         |
  | v2_lora_PanRight.ckpt                | 19 M      | 74 MB         |
  | v2_lora_PanUp.ckpt                   | 19 M      | 74 MB         |
  | v2_lora_PanDown.ckpt                 | 19 M      | 74 MB         |
  | v2_lora_RollingClockwise.ckpt        | 19 M      | 74 MB         |
  | v2_lora_RollingAnticlockwise.ckpt    | 19 M      | 74 MB         |

</details>

## Common Issues
<details>
<summary>Installation</summary>

Please ensure the installation of [xformer](https://github.com/facebookresearch/xformers) that is applied to reduce the inference memory.
</details>


<details>
<summary>Various resolution or number of frames</summary>
Currently, we recommend users to generate animation with 16 frames and 512 resolution that are aligned with our training settings. Notably, various resolution/frames may affect the quality more or less. 
</details>


<details>
<summary>How to use it without any coding</summary>

1) Get lora models: train lora model with [A1111](https://github.com/continue-revolution/sd-webui-animatediff) based on a collection of your own favorite images (e.g., tutorials [English](https://www.youtube.com/watch?v=mfaqqL5yOO4), [Japanese](https://www.youtube.com/watch?v=N1tXVR9lplM), [Chinese](https://www.bilibili.com/video/BV1fs4y1x7p2/)) 
or download Lora models from [Civitai](https://civitai.com/).

2) Animate lora models: using gradio interface or A1111 
(e.g., tutorials [English](https://github.com/continue-revolution/sd-webui-animatediff), [Japanese](https://www.youtube.com/watch?v=zss3xbtvOWw), [Chinese](https://941ai.com/sd-animatediff-webui-1203.html)) 

3) Be creative togther with other techniques, such as, super resolution, frame interpolation, music generation, etc.
</details>


<details>
<summary>Animating a given image</summary>

We totally agree that animating a given image is an appealing feature, which we would try to support officially in future. For now, you may enjoy other efforts from the [talesofai](https://github.com/talesofai/AnimateDiff).  
</details>

<details>
<summary>Contributions from community</summary>
Contributions are always welcome!! The <code>dev</code> branch is for community contributions. As for the main branch, we would like to align it with the original technical report :)
</details>


## Setups for Inference

### Prepare Environment

***We updated our inference code with xformers and a sequential decoding trick. Now AnimateDiff takes only ~12GB VRAM to inference, and run on a single RTX3090 !!***

```
git clone https://github.com/guoyww/AnimateDiff.git
cd AnimateDiff

conda env create -f environment.yaml
conda activate animatediff
```

### Download Base T2I & Motion Module Checkpoints
We provide two versions of our Motion Module, which are trained on stable-diffusion-v1-4 and finetuned on v1-5 seperately.
It's recommanded to try both of them for best results.
```
git lfs install
git clone https://huggingface.co/runwayml/stable-diffusion-v1-5 models/StableDiffusion/

bash download_bashscripts/0-MotionModule.sh
```
You may also directly download the motion module checkpoints from [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules), then put them in `models/Motion_Module/` folder.

### Prepare Personalize T2I
Here we provide inference configs for 6 demo T2I on CivitAI.
You may run the following bash scripts to download these checkpoints.
```
bash download_bashscripts/1-ToonYou.sh
bash download_bashscripts/2-Lyriel.sh
bash download_bashscripts/3-RcnzCartoon.sh
bash download_bashscripts/4-MajicMix.sh
bash download_bashscripts/5-RealisticVision.sh
bash download_bashscripts/6-Tusun.sh
bash download_bashscripts/7-FilmVelvia.sh
bash download_bashscripts/8-GhibliBackground.sh
```

### Inference
After downloading the above peronalized T2I checkpoints, run the following commands to generate animations. The results will automatically be saved to `samples/` folder.
```
python -m scripts.animate --config configs/prompts/1-ToonYou.yaml
python -m scripts.animate --config configs/prompts/2-Lyriel.yaml
python -m scripts.animate --config configs/prompts/3-RcnzCartoon.yaml
python -m scripts.animate --config configs/prompts/4-MajicMix.yaml
python -m scripts.animate --config configs/prompts/5-RealisticVision.yaml
python -m scripts.animate --config configs/prompts/6-Tusun.yaml
python -m scripts.animate --config configs/prompts/7-FilmVelvia.yaml
python -m scripts.animate --config configs/prompts/8-GhibliBackground.yaml
```

To generate animations with a new DreamBooth/LoRA model, you may create a new config `.yaml` file in the following format:
```
NewModel:
  inference_config: "[path to motion module config file]"

  motion_module:
    - "models/Motion_Module/mm_sd_v14.ckpt"
    - "models/Motion_Module/mm_sd_v15.ckpt"
    
    motion_module_lora_configs:
    - path:  "[path to MotionLoRA model]"
      alpha: 1.0
    - ...

  dreambooth_path: "[path to your DreamBooth model .safetensors file]"
  lora_model_path: "[path to your LoRA model .safetensors file, leave it empty string if not needed]"

  steps:          25
  guidance_scale: 7.5

  prompt:
    - "[positive prompt]"

  n_prompt:
    - "[negative prompt]"
```
Then run the following commands:
```
python -m scripts.animate --config [path to the config file]
```


## Steps for Training

### Dataset
Before training, download the videos files and the `.csv` annotations of [WebVid10M](https://maxbain.com/webvid-dataset/) to the local mechine.
Note that our examplar training script requires all the videos to be saved in a single folder. You may change this by modifying `animatediff/data/dataset.py`.

### Configuration
After dataset preparations, update the below data paths in the config `.yaml` files in `configs/training/` folder:
```
train_data:
  csv_path:     [Replace with .csv Annotation File Path]
  video_folder: [Replace with Video Folder Path]
  sample_size:  256
```
Other training parameters (lr, epochs, validation settings, etc.) are also included in the config files.

### Training
To train motion modules
```
torchrun --nnodes=1 --nproc_per_node=1 train.py --config configs/training/training.yaml
```

To finetune the unet's image layers
```
torchrun --nnodes=1 --nproc_per_node=1 train.py --config configs/training/image_finetune.yaml
```


## Gradio Demo
We have created a Gradio demo to make AnimateDiff easier to use. To launch the demo, please run the following commands:
```
conda activate animatediff
python app.py
```
By default, the demo will run at `localhost:7860`.
<br><img src="__assets__/figs/gradio.jpg" style="width: 50em; margin-top: 1em">

## Gallery
Here we demonstrate several best results we found in our experiments.

<table class="center">
    <tr>
    <td><img src="__assets__/animations/model_01/01.gif"></td>
    <td><img src="__assets__/animations/model_01/02.gif"></td>
    <td><img src="__assets__/animations/model_01/03.gif"></td>
    <td><img src="__assets__/animations/model_01/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/30240/toonyou">ToonYou</a></p>

<table>
    <tr>
    <td><img src="__assets__/animations/model_02/01.gif"></td>
    <td><img src="__assets__/animations/model_02/02.gif"></td>
    <td><img src="__assets__/animations/model_02/03.gif"></td>
    <td><img src="__assets__/animations/model_02/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/4468/counterfeit-v30">Counterfeit V3.0</a></p>

<table>
    <tr>
    <td><img src="__assets__/animations/model_03/01.gif"></td>
    <td><img src="__assets__/animations/model_03/02.gif"></td>
    <td><img src="__assets__/animations/model_03/03.gif"></td>
    <td><img src="__assets__/animations/model_03/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/4201/realistic-vision-v20">Realistic Vision V2.0</a></p>

<table>
    <tr>
    <td><img src="__assets__/animations/model_04/01.gif"></td>
    <td><img src="__assets__/animations/model_04/02.gif"></td>
    <td><img src="__assets__/animations/model_04/03.gif"></td>
    <td><img src="__assets__/animations/model_04/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model： <a href="https://civitai.com/models/43331/majicmix-realistic">majicMIX Realistic</a></p>

<table>
    <tr>
    <td><img src="__assets__/animations/model_05/01.gif"></td>
    <td><img src="__assets__/animations/model_05/02.gif"></td>
    <td><img src="__assets__/animations/model_05/03.gif"></td>
    <td><img src="__assets__/animations/model_05/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/66347/rcnz-cartoon-3d">RCNZ Cartoon</a></p>

<table>
    <tr>
    <td><img src="__assets__/animations/model_06/01.gif"></td>
    <td><img src="__assets__/animations/model_06/02.gif"></td>
    <td><img src="__assets__/animations/model_06/03.gif"></td>
    <td><img src="__assets__/animations/model_06/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/33208/filmgirl-film-grain-lora-and-loha">FilmVelvia</a></p>

#### Community Cases
Here are some samples contributed by the community artists. Create a Pull Request if you would like to show your results here😚.

<table>
    <tr>
    <td><img src="__assets__/animations/model_07/init.jpg"></td>
    <td><img src="__assets__/animations/model_07/01.gif"></td>
    <td><img src="__assets__/animations/model_07/02.gif"></td>
    <td><img src="__assets__/animations/model_07/03.gif"></td>
    <td><img src="__assets__/animations/model_07/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">
Character Model：<a href="https://civitai.com/models/13237/genshen-impact-yoimiya">Yoimiya</a> 
(with an initial reference image, see <a href="https://github.com/talesofai/AnimateDiff">WIP fork</a> for the extended implementation.)


<table>
    <tr>
    <td><img src="__assets__/animations/model_08/01.gif"></td>
    <td><img src="__assets__/animations/model_08/02.gif"></td>
    <td><img src="__assets__/animations/model_08/03.gif"></td>
    <td><img src="__assets__/animations/model_08/04.gif"></td>
    </tr>
</table>
<p style="margin-left: 2em; margin-top: -1em">
Character Model：<a href="https://civitai.com/models/9850/paimon-genshin-impact">Paimon</a>;
Pose Model：<a href="https://civitai.com/models/107295/or-holdingsign">Hold Sign</a></p>

## BibTeX
```
@article{guo2023animatediff,
  title={AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning},
  author={Guo, Yuwei and Yang, Ceyuan and Rao, Anyi and Wang, Yaohui and Qiao, Yu and Lin, Dahua and Dai, Bo},
  journal={arXiv preprint arXiv:2307.04725},
  year={2023}
}
```

## Contact Us
**Yuwei Guo**: [guoyuwei@pjlab.org.cn](mailto:guoyuwei@pjlab.org.cn)  
**Ceyuan Yang**: [yangceyuan@pjlab.org.cn](mailto:yangceyuan@pjlab.org.cn)  
**Bo Dai**: [daibo@pjlab.org.cn](mailto:daibo@pjlab.org.cn)

## Acknowledgements
Codebase built upon [Tune-a-Video](https://github.com/showlab/Tune-A-Video).
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
+								# AnimateDiff
-												update readme

											
										
										
											2023-07-11 10:54:17 +08:00
+								This repository is the official implementation of [AnimateDiff](https://arxiv.org/abs/2307.04725).
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
-												update readme

											
										
										
											2023-07-11 10:54:17 +08:00
+								**[AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning](https://arxiv.org/abs/2307.04725)**
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
+								</br>
 								Yuwei Guo,
-												update readme

											
										
										
											2023-07-03 01:13:53 +08:00
+								Ceyuan Yang*,
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
+								Anyi Rao,
 								Yaohui Wang,
 								Yu Qiao,
 								Dahua Lin,
-												update readme

											
										
										
											2023-07-03 01:13:53 +08:00
+								Bo Dai
-												update

											
										
										
											2023-07-09 23:37:00 +08:00
+								<p style="font-size: 0.8em; margin-top: -1em">*Corresponding Author</p>
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
-												update

											
										
										
											2023-07-21 10:36:35 +08:00
+								<!-- [Arxiv Report](https://arxiv.org/abs/2307.04725) | [Project Page](https://animatediff.github.io/) -->
 								[![arXiv](https://img.shields.io/badge/arXiv-2307.04725-b31b1b.svg)](https://arxiv.org/abs/2307.04725)
 								[![Project Page](https://img.shields.io/badge/Project-Website-green)](https://animatediff.github.io/)
 								[![Open in OpenXLab](https://cdn-static.openxlab.org.cn/app-center/openxlab_app.svg)](https://openxlab.org.cn/apps/detail/Masbfca/AnimateDiff)
 								[![Hugging Face Spaces](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Spaces-yellow)](https://huggingface.co/spaces/guoyww/AnimateDiff)
-												update readme

											
										
										
											2023-07-03 00:32:41 +08:00
-												update readme

											
										
										
											2023-09-25 11:39:05 +08:00
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
+								## Features
-												update readme

											
										
										
											2023-09-25 11:39:05 +08:00
+								- **[2023/09/25]** Release **MotionLoRA** and its model zoo, **enabling camera movement controls**! Please download the MotionLoRA models (**74 MB per model**, available at [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules) ) and save them to the `models/MotionLoRA` folder. Example:
 								  ```
 								  python -m scripts.animate --config configs/prompts/v2/5-RealisticVision-MotionLoRA.yaml
 								  ```
 								    <table class="center">
 								      <tr style="line-height: 0">
 								      <td colspan="2" style="border: none; text-align: center">Zoom In</td>
 								      <td colspan="2" style="border: none; text-align: center">Zoom Out</td>
 								      <td colspan="2" style="border: none; text-align: center">Zoom Pan Left</td>
 								      <td colspan="2" style="border: none; text-align: center">Zoom Pan Right</td>
 								      </tr>
 								      <tr>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/01.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/02.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/02.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/01.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/03.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/04.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/04.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/03.gif"></td>
 								      </tr>
 								      <tr style="line-height: 0">
-												Update shot camera movement term
											
										
										
											2023-09-25 19:51:52 -07:00
+								      <td colspan="2" style="border: none; text-align: center">Tilt Up</td>
 								      <td colspan="2" style="border: none; text-align: center">Tilt Down</td>
-												update readme

											
										
										
											2023-09-25 11:39:05 +08:00
+								      <td colspan="2" style="border: none; text-align: center">Rolling Anti-Clockwise</td>
 								      <td colspan="2" style="border: none; text-align: center">Rolling Clockwise</td>
 								      </tr>
 								      <tr>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/05.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/05.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/06.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/06.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/07.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/07.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_01/08.gif"></td>
 								      <td style="border: none"><img src="__assets__/animations/motion_lora/model_02/08.gif"></td>
 								      </tr>
 								  </table>
 								- **[2023/09/10]** New Motion Module release! `mm_sd_v15_v2.ckpt` was trained on larger resolution & batch size, and gains noticeable quality improvements. Check it out at [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules) and use it with `configs/inference/inference-v2.yaml`. Example:
-												update readme

											
										
										
											2023-09-10 21:25:55 +08:00
+								  ```
 								  python -m scripts.animate --config configs/prompts/v2/5-RealisticVision.yaml
 								  ```
 								  Here is a qualitative comparison between `mm_sd_v15.ckpt` (left) and `mm_sd_v15_v2.ckpt` (right):
 								  <table class="center">
 								      <tr>
 								      <td><img src="__assets__/animations/compare/old_0.gif"></td>
 								      <td><img src="__assets__/animations/compare/new_0.gif"></td>
 								      <td><img src="__assets__/animations/compare/old_1.gif"></td>
 								      <td><img src="__assets__/animations/compare/new_1.gif"></td>
 								      <td><img src="__assets__/animations/compare/old_2.gif"></td>
 								      <td><img src="__assets__/animations/compare/new_2.gif"></td>
 								      <td><img src="__assets__/animations/compare/old_3.gif"></td>
 								      <td><img src="__assets__/animations/compare/new_3.gif"></td>
 								      </tr>
 								  </table>
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
+								- GPU Memory Optimization, ~12GB VRAM to inference
-												update

											
										
										
											2023-07-21 10:36:35 +08:00
+								- User Interface: [Gradio](#gradio-demo), A1111 WebUI Extension [sd-webui-animatediff](https://github.com/continue-revolution/sd-webui-animatediff) (by [@continue-revolution](https://github.com/continue-revolution))
 								- Google Colab: [Colab](https://colab.research.google.com/github/camenduru/AnimateDiff-colab/blob/main/AnimateDiff_colab.ipynb) (by [@camenduru](https://github.com/camenduru))
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
-												update readme

											
										
										
											2023-09-25 11:39:05 +08:00
 								## Model Zoo
 								<details open>
 								<summary>Motion Modules</summary>
 								  | Name                 | Parameter | Storage Space |
 								  |----------------------|-----------|---------------|
 								  | mm_sd_v14.ckpt       | 417 M     | 1.6 GB        |
 								  | mm_sd_v15.ckpt       | 417 M     | 1.6 GB        |
 								  | mm_sd_v15_v2.ckpt    | 453 M     | 1.7 GB        |
 								</details>
 								<details open>
 								<summary>MotionLoRAs</summary>
 								  | Name                                 | Parameter | Storage Space |
 								  |--------------------------------------|-----------|---------------|
 								  | v2_lora_ZoomIn.ckpt                  | 19 M      | 74 MB         |
 								  | v2_lora_ZoomOut.ckpt                 | 19 M      | 74 MB         |
 								  | v2_lora_PanLeft.ckpt                 | 19 M      | 74 MB         |
 								  | v2_lora_PanRight.ckpt                | 19 M      | 74 MB         |
 								  | v2_lora_PanUp.ckpt                   | 19 M      | 74 MB         |
 								  | v2_lora_PanDown.ckpt                 | 19 M      | 74 MB         |
 								  | v2_lora_RollingClockwise.ckpt        | 19 M      | 74 MB         |
 								  | v2_lora_RollingAnticlockwise.ckpt    | 19 M      | 74 MB         |
 								</details>
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								## Common Issues
 								<details>
 								<summary>Installation</summary>
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								Please ensure the installation of [xformer](https://github.com/facebookresearch/xformers) that is applied to reduce the inference memory.
 								</details>
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								<details>
 								<summary>Various resolution or number of frames</summary>
 								Currently, we recommend users to generate animation with 16 frames and 512 resolution that are aligned with our training settings. Notably, various resolution/frames may affect the quality more or less.
 								</details>
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
 								<details>
 								<summary>How to use it without any coding</summary>
 ) Get lora models: train lora model with [A1111](https://github.com/continue-revolution/sd-webui-animatediff) based on a collection of your own favorite images (e.g., tutorials [English](https://www.youtube.com/watch?v=mfaqqL5yOO4), [Japanese](https://www.youtube.com/watch?v=N1tXVR9lplM), [Chinese](https://www.bilibili.com/video/BV1fs4y1x7p2/))
 								or download Lora models from [Civitai](https://civitai.com/).
 ) Animate lora models: using gradio interface or A1111
 								(e.g., tutorials [English](https://github.com/continue-revolution/sd-webui-animatediff), [Japanese](https://www.youtube.com/watch?v=zss3xbtvOWw), [Chinese](https://941ai.com/sd-animatediff-webui-1203.html))
 ) Be creative togther with other techniques, such as, super resolution, frame interpolation, music generation, etc.
 								</details>
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								<details>
 								<summary>Animating a given image</summary>
-												update "usage without coding", fix hyperlink


											
										
										
											2023-07-20 12:58:29 -07:00
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								We totally agree that animating a given image is an appealing feature, which we would try to support officially in future. For now, you may enjoy other efforts from the [talesofai](https://github.com/talesofai/AnimateDiff).
 								</details>
 								<details>
 								<summary>Contributions from community</summary>
-												update

											
										
										
											2023-07-19 23:32:30 +08:00
+								Contributions are always welcome!! The <code>dev</code> branch is for community contributions. As for the main branch, we would like to align it with the original technical report :)
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								</details>
-												training script

											
										
										
											2023-08-20 17:02:57 +08:00
+								## Setups for Inference
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
 								### Prepare Environment
-												update readme

											
										
										
											2023-07-12 17:07:44 +08:00
-												update readme

											
										
										
											2023-07-12 17:40:23 +08:00
+								***We updated our inference code with xformers and a sequential decoding trick. Now AnimateDiff takes only ~12GB VRAM to inference, and run on a single RTX3090 !!***
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
 								```
-												update readme

											
										
										
											2023-07-12 17:07:44 +08:00
+								git clone https://github.com/guoyww/AnimateDiff.git
-												update readme

											
										
										
											2023-07-12 16:57:50 +08:00
+								cd AnimateDiff
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
-												update readme

											
										
										
											2023-07-12 16:57:50 +08:00
+								conda env create -f environment.yaml
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
+								conda activate animatediff
 								```
 								### Download Base T2I & Motion Module Checkpoints
 								We provide two versions of our Motion Module, which are trained on stable-diffusion-v1-4 and finetuned on v1-5 seperately.
 								It's recommanded to try both of them for best results.
 								```
 								git lfs install
 								git clone https://huggingface.co/runwayml/stable-diffusion-v1-5 models/StableDiffusion/
 								bash download_bashscripts/0-MotionModule.sh
 								```
-												update readme

											
										
										
											2023-09-25 11:39:05 +08:00
+								You may also directly download the motion module checkpoints from [Google Drive](https://drive.google.com/drive/folders/1EqLC65eR1-W-sGD0Im7fkED6c8GkiNFI?usp=sharing) / [HuggingFace](https://huggingface.co/guoyww/animatediff) / [CivitAI](https://civitai.com/models/108836/animatediff-motion-modules), then put them in `models/Motion_Module/` folder.
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
 								### Prepare Personalize T2I
 								Here we provide inference configs for 6 demo T2I on CivitAI.
 								You may run the following bash scripts to download these checkpoints.
 								```
 								bash download_bashscripts/1-ToonYou.sh
 								bash download_bashscripts/2-Lyriel.sh
 								bash download_bashscripts/3-RcnzCartoon.sh
 								bash download_bashscripts/4-MajicMix.sh
 								bash download_bashscripts/5-RealisticVision.sh
 								bash download_bashscripts/6-Tusun.sh
 								bash download_bashscripts/7-FilmVelvia.sh
 								bash download_bashscripts/8-GhibliBackground.sh
 								```
 								### Inference
-												update readme

											
										
										
											2023-07-12 16:57:50 +08:00
+								After downloading the above peronalized T2I checkpoints, run the following commands to generate animations. The results will automatically be saved to `samples/` folder.
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
+								```
 								python -m scripts.animate --config configs/prompts/1-ToonYou.yaml
 								python -m scripts.animate --config configs/prompts/2-Lyriel.yaml
 								python -m scripts.animate --config configs/prompts/3-RcnzCartoon.yaml
 								python -m scripts.animate --config configs/prompts/4-MajicMix.yaml
 								python -m scripts.animate --config configs/prompts/5-RealisticVision.yaml
 								python -m scripts.animate --config configs/prompts/6-Tusun.yaml
 								python -m scripts.animate --config configs/prompts/7-FilmVelvia.yaml
 								python -m scripts.animate --config configs/prompts/8-GhibliBackground.yaml
 								```
-												update readme

											
										
										
											2023-07-03 01:09:13 +08:00
-												update readme

											
										
										
											2023-07-14 21:36:01 +08:00
+								To generate animations with a new DreamBooth/LoRA model, you may create a new config `.yaml` file in the following format:
 								```
 								NewModel:
-												fix error

											
										
										
											2023-09-25 19:41:24 +08:00
+								  inference_config: "[path to motion module config file]"
-												update readme

											
										
										
											2023-07-14 21:36:01 +08:00
 								  motion_module:
 								    - "models/Motion_Module/mm_sd_v14.ckpt"
 								    - "models/Motion_Module/mm_sd_v15.ckpt"
-												fix error

											
										
										
											2023-09-25 19:41:24 +08:00
+								    motion_module_lora_configs:
 								    - path:  "[path to MotionLoRA model]"
 								      alpha: 1.0
 								    - ...
 								  dreambooth_path: "[path to your DreamBooth model .safetensors file]"
 								  lora_model_path: "[path to your LoRA model .safetensors file, leave it empty string if not needed]"
-												update readme

											
										
										
											2023-07-14 21:36:01 +08:00
+								  steps:          25
 								  guidance_scale: 7.5
 								  prompt:
 								    - "[positive prompt]"
 								  n_prompt:
 								    - "[negative prompt]"
 								```
 								Then run the following commands:
 								```
 								python -m scripts.animate --config [path to the config file]
 								```
-												training script

											
										
										
											2023-08-20 17:02:57 +08:00
 								## Steps for Training
 								### Dataset
 								Before training, download the videos files and the `.csv` annotations of [WebVid10M](https://maxbain.com/webvid-dataset/) to the local mechine.
 								Note that our examplar training script requires all the videos to be saved in a single folder. You may change this by modifying `animatediff/data/dataset.py`.
 								### Configuration
 								After dataset preparations, update the below data paths in the config `.yaml` files in `configs/training/` folder:
 								```
 								train_data:
 								  csv_path:     [Replace with .csv Annotation File Path]
 								  video_folder: [Replace with Video Folder Path]
 								  sample_size:  256
 								```
 								Other training parameters (lr, epochs, validation settings, etc.) are also included in the config files.
 								### Training
 								To train motion modules
 								```
 								torchrun --nnodes=1 --nproc_per_node=1 train.py --config configs/training/training.yaml
 								```
 								To finetune the unet's image layers
 								```
 								torchrun --nnodes=1 --nproc_per_node=1 train.py --config configs/training/image_finetune.yaml
 								```
-												gradio

											
										
										
											2023-07-18 16:41:22 +08:00
+								## Gradio Demo
-												gradio

											
										
										
											2023-07-18 16:43:10 +08:00
+								We have created a Gradio demo to make AnimateDiff easier to use. To launch the demo, please run the following commands:
-												gradio

											
										
										
											2023-07-18 16:41:22 +08:00
+								```
 								conda activate animatediff
 								python app.py
 								```
-												gradio

											
										
										
											2023-07-18 16:43:10 +08:00
+								By default, the demo will run at `localhost:7860`.
-												gradio

											
										
										
											2023-07-18 16:41:22 +08:00
+								<br><img src="__assets__/figs/gradio.jpg" style="width: 50em; margin-top: 1em">
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
+								## Gallery
-												update readme

											
										
										
											2023-07-13 19:32:50 +08:00
+								Here we demonstrate several best results we found in our experiments.
-												update readme

											
										
										
											2023-07-03 01:09:13 +08:00
 								<table class="center">
 								    <tr>
 								    <td><img src="__assets__/animations/model_01/01.gif"></td>
 								    <td><img src="__assets__/animations/model_01/02.gif"></td>
 								    <td><img src="__assets__/animations/model_01/03.gif"></td>
 								    <td><img src="__assets__/animations/model_01/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/30240/toonyou">ToonYou</a></p>
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_02/01.gif"></td>
 								    <td><img src="__assets__/animations/model_02/02.gif"></td>
 								    <td><img src="__assets__/animations/model_02/03.gif"></td>
 								    <td><img src="__assets__/animations/model_02/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/4468/counterfeit-v30">Counterfeit V3.0</a></p>
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_03/01.gif"></td>
 								    <td><img src="__assets__/animations/model_03/02.gif"></td>
 								    <td><img src="__assets__/animations/model_03/03.gif"></td>
 								    <td><img src="__assets__/animations/model_03/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/4201/realistic-vision-v20">Realistic Vision V2.0</a></p>
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_04/01.gif"></td>
 								    <td><img src="__assets__/animations/model_04/02.gif"></td>
 								    <td><img src="__assets__/animations/model_04/03.gif"></td>
 								    <td><img src="__assets__/animations/model_04/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">Model： <a href="https://civitai.com/models/43331/majicmix-realistic">majicMIX Realistic</a></p>
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_05/01.gif"></td>
 								    <td><img src="__assets__/animations/model_05/02.gif"></td>
 								    <td><img src="__assets__/animations/model_05/03.gif"></td>
 								    <td><img src="__assets__/animations/model_05/04.gif"></td>
 								    </tr>
 								</table>
-												update

											
										
										
											2023-07-09 23:49:59 +08:00
+								<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/66347/rcnz-cartoon-3d">RCNZ Cartoon</a></p>
-												update readme

											
										
										
											2023-07-03 01:09:13 +08:00
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_06/01.gif"></td>
 								    <td><img src="__assets__/animations/model_06/02.gif"></td>
 								    <td><img src="__assets__/animations/model_06/03.gif"></td>
 								    <td><img src="__assets__/animations/model_06/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">Model：<a href="https://civitai.com/models/33208/filmgirl-film-grain-lora-and-loha">FilmVelvia</a></p>
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
-												update readme

											
										
										
											2023-07-13 19:32:50 +08:00
+								#### Community Cases
 								Here are some samples contributed by the community artists. Create a Pull Request if you would like to show your results here😚.
-												update readme: community cases

											
										
										
											2023-07-13 17:40:54 +08:00
 								<table>
 								    <tr>
-												update readme

											
										
										
											2023-07-13 19:32:50 +08:00
+								    <td><img src="__assets__/animations/model_07/init.jpg"></td>
-												update readme: community cases

											
										
										
											2023-07-13 17:40:54 +08:00
+								    <td><img src="__assets__/animations/model_07/01.gif"></td>
 								    <td><img src="__assets__/animations/model_07/02.gif"></td>
 								    <td><img src="__assets__/animations/model_07/03.gif"></td>
 								    <td><img src="__assets__/animations/model_07/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">
 								Character Model：<a href="https://civitai.com/models/13237/genshen-impact-yoimiya">Yoimiya</a>
-												update readme

											
										
										
											2023-07-13 19:35:09 +08:00
+								(with an initial reference image, see <a href="https://github.com/talesofai/AnimateDiff">WIP fork</a> for the extended implementation.)
-												update readme

											
										
										
											2023-07-13 19:32:50 +08:00
-												update readme: community cases

											
										
										
											2023-07-13 17:40:54 +08:00
 								<table>
 								    <tr>
 								    <td><img src="__assets__/animations/model_08/01.gif"></td>
 								    <td><img src="__assets__/animations/model_08/02.gif"></td>
 								    <td><img src="__assets__/animations/model_08/03.gif"></td>
 								    <td><img src="__assets__/animations/model_08/04.gif"></td>
 								    </tr>
 								</table>
 								<p style="margin-left: 2em; margin-top: -1em">
-												update readme

											
										
										
											2023-07-13 19:32:50 +08:00
+								Character Model：<a href="https://civitai.com/models/9850/paimon-genshin-impact">Paimon</a>;
 								Pose Model：<a href="https://civitai.com/models/107295/or-holdingsign">Hold Sign</a></p>
-												update readme: community cases

											
										
										
											2023-07-13 17:40:54 +08:00
-												update

											
										
										
											2023-07-10 13:46:01 +08:00
+								## BibTeX
-												update readme

											
										
										
											2023-07-11 10:54:17 +08:00
+								```
-												update README

											
										
										
											2023-07-16 10:40:59 +08:00
+								@article{guo2023animatediff,
 								  title={AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning},
 								  author={Guo, Yuwei and Yang, Ceyuan and Rao, Anyi and Wang, Yaohui and Qiao, Yu and Lin, Dahua and Dai, Bo},
 								  journal={arXiv preprint arXiv:2307.04725},
 								  year={2023}
-												update readme

											
										
										
											2023-07-11 10:54:17 +08:00
+								}
 								```
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
-												update

											
										
										
											2023-07-10 13:52:01 +08:00
+								## Contact Us
 								**Yuwei Guo**: [guoyuwei@pjlab.org.cn](mailto:guoyuwei@pjlab.org.cn)
 								**Ceyuan Yang**: [yangceyuan@pjlab.org.cn](mailto:yangceyuan@pjlab.org.cn)
 								**Bo Dai**: [daibo@pjlab.org.cn](mailto:daibo@pjlab.org.cn)
-												add bashscripts

											
										
										
											2023-07-09 22:40:48 +08:00
+								## Acknowledgements
-												Update shot camera movement term
											
										
										
											2023-09-25 19:51:52 -07:00
+								Codebase built upon [Tune-a-Video](https://github.com/showlab/Tune-A-Video).