Skip to content

Ideogram 4

Ideogram 4 is an open-weight text-to-image model with a distinctive structured JSON prompt: instead of a single sentence, the model is trained to read an overall scene description plus a list of regions, each with a bounding box and its own text. Invoke assembles this JSON for you.

Ideogram 4 is released under the Ideogram 4 Non-Commercial License. The weights may not be used commercially, and outputs may not be used to build competing models. Ideogram claims no rights in images you generate, but the license does not explicitly grant commercial use of them. The ideogram-ai repositories are gated; the Comfy-Org copy is not, but the same license applies.

This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.

The weights are gated on HuggingFace under a non-commercial license. Open the model page, accept the terms, and make sure your HuggingFace token is set up in the config file before installing. Two bundled builds are available:

Paste either repo ID into the Model Manager’s HuggingFace / URL field to install. Each folder carries everything Ideogram 4 needs.

Comfy-Org’s Comfy-Org/Ideogram-4 repackage is ungated (the same non-commercial license still applies) and ships the parts separately. Ideogram 4 runs two transformers — a conditional branch and an unconditional one, both active at every step — so a single-file install is four models, and each starter installs all four together:

  • diffusion_models/ideogram4_<build>.safetensors — the conditional branch, selected as the model
  • diffusion_models/ideogram4_unconditional_<build>.safetensors — the unconditional branch
  • text_encoders/qwen3vl_8b_fp8_scaled.safetensors — the Qwen3-VL 8B encoder (not the 4B one Krea-2 uses). A llama.cpp GGUF of the 8B encoder works as well; see Qwen3-VL GGUF encoders.
  • vae/flux2-vae.safetensors — the 32-channel VAE, shared with FLUX.2, which is why it installs under the FLUX.2 base

Pick the unconditional branch, the encoder and the VAE in the Components section next to the model — and pick the unconditional branch of the same build as the model. The nvfp4 repackage in that repo is not supported yet and is refused at install.

Which branch a file holds is read from the file itself, and from the filename for a build that does not record it, so keep these names as published.

Two safetensors builds of the transformers are supported, and they behave differently in memory:

download, per branchresident, per branch
Ideogram 4 (single file, int8) — int8_convrot8.9 GB8.9 GB, on every device
Ideogram 4 (single file, fp8) — fp8_scaled8.6 GB8.7 GB with fp8_compute or FP8 Storage; 17.3 GB with neither

Both branches are resident at once, so double those figures for the pair.

Community GGUF conversions of the same two transformers load too, and they are the smallest builds. The weights stay packed in memory and are unpacked layer by layer during generation, on every device, so the download size is about the resident size. Three pairs are available as starter models, each of which installs its unconditional branch, the Qwen3-VL 8B encoder and the VAE with it:

download, per branch1024px, 12 steps, 4090
Ideogram 4 (GGUF, Q4_0) — molbal/ideogram-4-gguf5.6 GB27.9 s
Ideogram 4 (GGUF, Q5_K) — rectangleworm/ideogram-4-gguf6.4 GB32.2 s
Ideogram 4 (GGUF, Q5_1) — molbal/ideogram-4-gguf7.3 GBnot measured

For comparison, the fp8 single-file pair took 37–38 s on the same card. Both GGUF pairs stay fully on the GPU next to the encoder (5.3 and 6.0 GB per branch), where the fp8 pair has to stream part of its second branch. Q5_K is the closer of the two to the fp8 build’s images.

Pick the model and the unconditional branch in Components exactly as for the safetensors builds. Two things are different:

  • The branch comes from the filename alone. No published GGUF records which branch it holds, so a file whose name contains unconditional or uncond installs as the unconditional branch and anything else as the conditional one. Keep the published names; to fix a renamed file, remove the model, give the file its published name back and install it again.
  • Other GGUF repositories of Ideogram 4 (for example leejet’s or stduhpf’s) load as well. IQ4_NL files work but are unpacked on the CPU at every step, which is much slower. Q8_0 is not offered as a starter: at about 10.1 GB per branch it is larger than the fp8 build.

Ideogram 4 keeps both transformer branches loaded for the whole generation. With the single-file builds on a CUDA or ROCm GPU (with partial loading on, the default), InvokeAI checks whether the pair fits next to the working memory a generation needs. If it does not, it keeps only part of the second branch on the GPU and streams the rest in each step. On a 24 GB card with the fp8 pair this keeps about 2.4 GB free instead of 1.2 GB. The bundled nf4 and fp8 folders are not affected. Three things follow:

  • A streamed layer computes on a slightly different path, so images at the same seed can differ from those generated with both branches fully on the GPU.
  • If generations still run out of memory, raise device_working_mem_gb in invokeai.yaml (for example to 8).
  • If you set max_cache_vram_gb, InvokeAI does not limit the second branch this way, because the cache already plans from that limit.

When an Ideogram 4 model is selected, Invoke builds the structured JSON prompt automatically:

  • The positive prompt becomes the overall scene description.
  • Each enabled Regional Guidance layer on the Canvas contributes one element: its drawn box becomes the region’s bounding box and its prompt becomes that region’s description. Draw a box where you want something and describe it there.
  • To drive the model directly, paste a raw JSON object into the prompt box — anything starting with { is passed through unchanged.

The exact JSON that was encoded is stored in the image metadata as Structured Caption, and can be recalled straight back into the prompt box from the metadata viewer.

  • Sampler Preset — the primary quality/speed control. Quality (48 steps), Default (20 steps), and Turbo (12 steps) each bundle a step count, a guidance schedule, and the schedule shift.
  • Advanced overrides (all optional, leave on Auto to use the preset’s values): Steps, Guidance Scale, Schedule Shift (mu), and a Color Palette that biases the generated colors.
This site was designed and developed by Aether Fox Studio.