Anima
Anima is a 2B-parameter, anime-focused text-to-image model built on NVIDIA’s Cosmos Predict2 diffusion transformer. Instead of a large text encoder, it pairs a small Qwen3 0.6B encoder with a built-in LLM adapter that translates the encoder’s output for the transformer. It decodes with the 16-channel Wan 2.1 / Qwen Image VAE.
InvokeAI supports Anima for text-to-image, and on the Canvas for image-to-image, inpainting, outpainting and regional prompts. Two community finetunes are supported as well: Anima-2.9B and Anima-3.8B (see below).
License
Section titled “License”Anima is released under the CircleStone Labs Non-Commercial License, and, because it is built on Cosmos-Predict2, the NVIDIA Open Model License also applies. The weights may not be used commercially; the model card says images you generate may be. The download is not gated.
This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.
Hardware
Section titled “Hardware”Anima is small by current standards: the transformer is a ~4.5 GB download, the Qwen3 encoder ~1.2 GB and the VAE ~200 MB. On low-VRAM GPUs, Low-VRAM mode and FP8 Storage both apply to Anima.
Installing
Section titled “Installing”Install the Anima bundle from the Model Manager’s Starter Models. It contains:
- Anima Base 1.0 — the main model (single-file checkpoint, ~4.5 GB)
- Anima Qwen3 0.6B Text Encoder (~1.2 GB)
- Anima QwenImage VAE (~200 MB)
- Anima LLLite Inpainting and Anima LLLite Sketch — ControlNet-LLLite adapters (see below)
Installing Anima Base 1.0 on its own also installs the encoder and VAE as dependencies. Anima is distributed as a single file only, so the three components are always separate models:
| Component | Model |
|---|---|
| Transformer | the Anima checkpoint (Cosmos Predict2 DiT + LLM adapter) |
| Text encoder | a Qwen3 0.6B encoder |
| VAE | the 16-channel Wan 2.1 / Qwen Image VAE |
Select the Qwen3 Encoder and VAE in the Components section next to the model; both are required. Only the 0.6B Qwen3 encoder is offered — the larger Qwen3 encoders used by Z-Image and FLUX.2 are not compatible. The same VAE file may be installed under the Anima, Qwen Image or Wan base, and all three are listed. FLUX VAEs are not compatible.
Anima also needs a T5-XXL tokenizer, which ships with InvokeAI; no T5 model has to be installed.
Community finetunes
Section titled “Community finetunes”Both finetunes keep Anima’s VAE, schedulers and Canvas support, and install from the Starter Models.
| Model | What changes | Components |
|---|---|---|
| Anima-2.9B (Preview v1) | 40 transformer blocks instead of 28, trained on 1.7M more images (cutoff July 2026). An int8 build (~3.1 GB) halves the memory of the bf16 one (~5.8 GB); it is not faster. | Same as Anima |
| Anima-3.8B (v1.1) | 52 blocks, plus a semantic connector bundled in the checkpoint that adds a second text encoder, Qwen3.5 4B, for better prompt adherence, multi-character binding and mixed natural-language/tag prompts. | Qwen3 0.6B and Qwen3.5 4B encoder, VAE |
With an Anima-3.8B model selected, the Components section shows a third picker, Qwen3.5 Encoder, and generation needs it. Its starter model (~4.8 GB) installs together with Anima-3.8B. The connector runs once per denoising step, so Anima-3.8B is slower per step than Anima-2.9B (about 1.3 against 1.7 iterations per second at 832×1216 on an RTX 4090), and its Qwen3.5 encoder adds memory while the prompt is encoded. Reset all to model defaults sets 40 steps at CFG 6, the settings of the author’s reference workflow.
FP8 Storage changes Anima-3.8B’s images more than Anima’s: at the same seed the composition can change, where Anima keeps it and only details differ.
Only the v1.1 checkpoint of Anima-3.8B is supported. The earlier v1.0 transformer needs a separate adapter file and is refused at install with a message that says so.
Both finetunes are derivatives of Anima and fall under the same non-commercial license.
Generation settings
Section titled “Generation settings”Selecting an Anima model applies these defaults:
- Steps: 35
- CFG Scale: 4.5 (the minimum is 1, which turns CFG off)
- Scheduler: Euler
- Size: 1024×1024; width and height snap to multiples of 8.
The Scheduler menu offers Anima’s own set: Euler, Heun (2nd order), DPM++ 2M, DPM++ 2M SDE, ER-SDE and LCM.
The negative prompt is used only when CFG Scale is above 1; at CFG 1 it is ignored.
Regional prompting
Section titled “Regional prompting”On the Canvas, Regional Guidance layers with a positive prompt work with Anima: each region’s prompt is steered to its masked area, alongside the global prompt. The restriction is applied on alternating transformer blocks, leaving the others unrestricted to keep the image coherent. Regional negative prompts and regional reference images are not supported.
In the workflow editor, the Prompt - Anima node has an optional mask input, and Denoise - Anima accepts a collection of conditionings for its positive and negative inputs.
Anima LoRAs (Kohya / LyCORIS format) are supported. They apply to the transformer and, where the LoRA includes text-encoder layers, to the Qwen3 encoder. In the workflow editor use Apply LoRA - Anima or Apply LoRA Collection - Anima.
ControlNet-LLLite adapters
Section titled “ControlNet-LLLite adapters”Anima supports kohya-ss’s ControlNet-LLLite adapters, small models (8–66 MB) that condition the transformer on an image. Each adapter model may be used only once per generation, and each has Weight, Begin Step Percent and End Step Percent settings.
- On the canvas, add a control layer while an Anima model is selected. Its adapter type is Anima ControlNet-LLLite, its pixels are the control image, and several control layers combine as long as each uses a different model. The Model list holds the control adapters (Sketch, Depth, Scribble, Lineart, Pose); the Inpainting adapter needs an inpaint mask rather than a control image, so it is not offered there. An adapter installed before Invoke recorded which kind it is is not listed either: re-identify it in the Model Manager. A control layer from an older project may still be set to ControlNet, which Anima cannot run; its settings say so and offer Switch to Anima ControlNet-LLLite, which keeps the selected adapter.
- In the workflow editor, add an Anima ControlNet-LLLite node, give it the conditioning image and a control model, and connect it to the Control LLLite input of Denoise - Anima, collecting several nodes to combine them.
| Starter model | Purpose |
|---|---|
| Anima LLLite Inpainting | Conditions on the masked image during inpainting/outpainting. Requires the node’s mask input (white = area to inpaint). |
| Anima LLLite Sketch | Mixed scribble / HED / lineart / grayscale conditioning |
| Anima LLLite Depth (Preview3) | Depth |
| Anima LLLite Scribble (Preview3) | Scribble |
| Anima LLLite Lineart (Preview3) | Lineart |
| Anima LLLite Pose (Preview3) | Pose |
Workflow nodes
Section titled “Workflow nodes”| Node | Purpose |
|---|---|
| Main Model - Anima | Loads the transformer, Qwen3 encoder and VAE, and for Anima-3.8B the Qwen3.5 encoder |
| Prompt - Anima | Encodes a prompt, with an optional regional mask; connect the Qwen3.5 encoder for Anima-3.8B |
| Denoise - Anima | Runs sampling; accepts img2img latents, masks and LLLite adapters |
| Image to Latents - Anima | VAE encode |
| Latents to Image - Anima | VAE decode |
| Anima ControlNet-LLLite | Configures one ControlNet-LLLite adapter |