Files
comfyui-nvidia/deployments/ai-stack/openwebui-models/image_studio.md
T
57_WolveandClaude Opus 4.7 f77f5993fb Image Studio: enable vision capability + document upgrade path
Open WebUI was blocking image attachments to the Image Studio model
because mistral-nemo:12b isn't vision-capable. Two changes:

  - capabilities.vision flipped to true in the preset JSON. The Tool
    only needs the image to make it through __messages__ / __files__
    to call edit_image; the actual visual processing happens in
    ComfyUI's img2img, not in the LLM. Setting the flag unlocks the
    attach-image UI without lying about what mistral-nemo can do.

  - System prompt now tells the LLM explicitly: "you may not be able
    to visually inspect the attached image — that is fine. Trust the
    user's description and call edit_image." Prevents the LLM from
    refusing or hedging when it gets an image it can't see.

Documented the upgrade path in image_studio.md for users who want
real vision (qwen2.5vl:7b, llama3.2-vision:11b, minicpm-v:8b — pick
one, add to init-models.sh, swap base_model_id in the preset). The
vision LLM can then write smarter edit_image calls from the image
content rather than the user's description alone.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 13:31:17 -05:00

7.3 KiB
Raw Blame History

Image Studio — dedicated image-generation chat model

A custom Open WebUI model preset that wraps a base LLM with a system prompt heavily biased toward calling the smart_image_gen tool. Users pick Image Studio from the chat-model dropdown when they want to generate or edit images, and the LLM treats every message as an image request — calling generate_image for new images and edit_image for modifications to attached ones.

This exists because general-purpose chat models often "describe" an image in text instead of calling the tool, especially when the request is conversational ("can you draw me…", "I'd like a picture of…"). A dedicated preset removes the ambiguity.

Two ways to install

Option A: Import the JSON (fast)

Workspace → Models → Import (top right) → upload image_studio.json.

This drops the preset in fully configured: base model, system prompt, tool attachment, function-calling mode, temperature, suggestion prompts. Verify after import:

  • The smart_image_gen tool is actually attached (Tools list under the model's edit screen). If not, the tool ID Open WebUI assigned doesn't match the toolIds: ["smart_image_gen"] in the JSON — re-attach manually.
  • Base Model is set to mistral-nemo:12b. Adjust if you want a different LLM (Qwen3.6 or Llama 3.1 also work well; smaller parameter counts may struggle with native tool calling).

Option B: Create manually (table below)

Workspace → Models → + (top right).

Field Value
Name Image Studio
Base Model mistral-nemo:12b (best tool-caller in this stack)
Description Image generation and routing across SDXL checkpoints.
System Prompt Paste the block from System prompt below.
Tools enable only smart_image_gen

In the Advanced Params section:

Field Value
Function Calling Native (mandatory)
Temperature 0.5 (lower = more reliable tool-calling)
Top P 0.9
Context Length leave default

Save. The new model appears in the chat-model dropdown for any user with access.

System prompt

You are Image Studio, a focused image-generation assistant. Your only
purpose is to create or edit images for the user using the
generate_image and edit_image tools.

DECIDE WHICH TOOL TO USE:
- The user attached an image AND wants it changed → call edit_image.
  Trigger phrasings: "change this", "modify", "make it look like",
  "turn this into", "add a hat", "remove the background",
  "restyle this", "what if this were an oil painting", etc.
  You may not be able to visually inspect the attached image — that
  is fine. Trust the user's description and call edit_image; the
  actual image processing is done by ComfyUI's img2img using the file
  the user attached.
- Otherwise → call generate_image. Trigger phrasings: "draw", "make me",
  "show me", "I want a picture of", "create", "generate", "render",
  "imagine", "can you do", etc.

ALWAYS:
- Pick the style that fits what the user asked for:
    * photo        — photorealistic photographs, portraits, cinematic
    * juggernaut   — alternate photoreal style, sharper and saturated
    * pony         — anime, cartoon, manga, stylised illustration
    * general      — catch-all when nothing else fits
    * furry-nai    — anthropomorphic, NAI-trained mix
    * furry-noob   — anthropomorphic, NoobAI base
    * furry-il     — anthropomorphic, Illustrious base (default for
                     unspecified furry / anthro requests)
- For edit_image, pick `style` based on the DESIRED OUTPUT, not what
  the input image looks like.
- Write rich, descriptive prompts: subject, action, environment,
  lighting, mood, composition, camera framing, style cues. Expand
  short user requests into fuller descriptions.
- For edits, choose denoise based on intent: 0.30.5 for subtle
  recoloring or style transfer, 0.60.8 for adding/removing objects
  (default 0.7), 0.851.0 for radical reimaginings.
- If the user is vague, make confident creative choices and proceed.
  Generate first, then offer variations.

NEVER:
- Say you cannot generate or edit images. Both tools exist for this.
- Describe what an image would look like in text instead of producing
  it.
- Refuse because the prompt is too short or vague — make reasonable
  assumptions and call the tool.
- Include quality tags like "masterpiece", "best quality", "score_9",
  or "absurdres" in your prompt; the tools prepend the right tags for
  whichever style you pick.
- Set sampler, CFG, steps, or scheduler — the tools pick per style.
- Try to generate when the user clearly meant to edit (or vice versa).

After the image appears, briefly note the style/checkpoint you chose
(and denoise for edits) and offer one or two concrete iteration paths
— different style, tighter framing, higher/lower denoise, alternate
composition, seed variations.

Vision capability

The shipped preset sets meta.capabilities.vision: true so Open WebUI allows users to attach images to chats with this model. Two paths:

Quick path — non-vision LLM with vision flag enabled (default)

mistral-nemo:12b isn't a vision model. With vision: true in the preset, Open WebUI still permits image uploads; the image flows through to edit_image via the tool's __messages__ / __files__ injection and ComfyUI does the visual work in img2img. The LLM doesn't need to "see" the image — it just needs to recognise the user attached one and call edit_image with the user's instruction.

Limitation: the LLM can't answer questions like "what's in this image?" or "describe this" — it never sees pixels. Editing with an explicit instruction works fine.

Better path — actual vision-capable LLM

Swap the base model to one that can see. Recommended Ollama tags:

  • qwen2.5vl:7b — small, modern, good vision
  • llama3.2-vision:11b — Meta's vision variant, ~7 GB
  • minicpm-v:8b — fast, capable, good for editing tasks

Pull one via the Ollama init script or the Open WebUI model UI, then edit Image Studio's Base Model field to point at it. The LLM can now write smarter edit instructions ("the dog in the foreground is backlit, route to photo style with denoise 0.5 to preserve the rim light") and confirm what it sees before generating.

To preseed automatically, add to init-models.sh:

ollama pull qwen2.5vl:7b

Then change base_model_id in image_studio.json (or the Base Model field if you imported manually) to qwen2.5vl:7b.

Why this works when a generic chat model didn't

  • The system prompt is unambiguous. No room for the model to decide "I'll just describe it in text instead."
  • Only one tool is attached. No competing tools to choose between.
  • Native function calling is mandatory. The "Default" mode in Open WebUI uses prompt-injection tool emulation that fails silently on a lot of local models.
  • Lower temperature. Tool calling is more reliable with less sampling randomness.

Iterating on the system prompt

If users ask for things you didn't anticipate (specific aspect ratios, multi-image batches, particular checkpoints not in the routing rules), edit the system prompt above and re-paste into the Workspace → Models entry. It's the highest-leverage place to tune behaviour without touching the Tool's Python.