Image Studio: enable vision capability + document upgrade path

Open WebUI was blocking image attachments to the Image Studio model
because mistral-nemo:12b isn't vision-capable. Two changes:

  - capabilities.vision flipped to true in the preset JSON. The Tool
    only needs the image to make it through __messages__ / __files__
    to call edit_image; the actual visual processing happens in
    ComfyUI's img2img, not in the LLM. Setting the flag unlocks the
    attach-image UI without lying about what mistral-nemo can do.

  - System prompt now tells the LLM explicitly: "you may not be able
    to visually inspect the attached image — that is fine. Trust the
    user's description and call edit_image." Prevents the LLM from
    refusing or hedging when it gets an image it can't see.

Documented the upgrade path in image_studio.md for users who want
real vision (qwen2.5vl:7b, llama3.2-vision:11b, minicpm-v:8b — pick
one, add to init-models.sh, swap base_model_id in the preset). The
vision LLM can then write smarter edit_image calls from the image
content rather than the user's description alone.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-04-19 13:31:17 -05:00
co-authored by Claude Opus 4.7
parent b604e3f509
commit f77f5993fb
3 changed files with 54 additions and 2 deletions
+7
View File
@@ -201,6 +201,13 @@ Setup — two paths:
Users then pick **Image Studio** from the chat-model dropdown when
they want to generate or edit images.
The preset ships with `vision: true` so users can attach images for
editing even though `mistral-nemo:12b` isn't a vision model — see the
[**Vision capability** section in image_studio.md](openwebui-models/image_studio.md#vision-capability)
for the trade-offs and the upgrade path to a real vision LLM
(`qwen2.5vl:7b`, `llama3.2-vision:11b`, etc.) if the LLM needs to
actually see the image to write smarter edit instructions.
To extend (new checkpoint, new style):
- Add the filename to `comfyui-init-models.sh` so it gets pulled.
@@ -4,7 +4,7 @@
"base_model_id": "mistral-nemo:12b",
"name": "Image Studio",
"params": {
"system": "You are Image Studio, a focused image-generation assistant. Your only purpose is to create or edit images for the user using the generate_image and edit_image tools.\n\nDECIDE WHICH TOOL TO USE:\n- The user attached an image AND wants it changed → call edit_image. Trigger phrasings: \"change this\", \"modify\", \"make it look like\", \"turn this into\", \"add a hat\", \"remove the background\", \"restyle this\", \"what if this were an oil painting\", etc.\n- Otherwise → call generate_image. Trigger phrasings: \"draw\", \"make me\", \"show me\", \"I want a picture of\", \"create\", \"generate\", \"render\", \"imagine\", \"can you do\", etc.\n\nALWAYS:\n- Pick the style that fits what the user asked for:\n * photo — photorealistic photographs, portraits, cinematic\n * juggernaut — alternate photoreal style, sharper and saturated\n * pony — anime, cartoon, manga, stylised illustration\n * general — catch-all when nothing else fits\n * furry-nai — anthropomorphic, NAI-trained mix\n * furry-noob — anthropomorphic, NoobAI base\n * furry-il — anthropomorphic, Illustrious base (default for unspecified furry / anthro requests)\n- For edit_image, pick `style` based on the DESIRED OUTPUT, not what the input image looks like.\n- Write rich, descriptive prompts: subject, action, environment, lighting, mood, composition, camera framing, style cues. Expand short user requests into fuller descriptions.\n- For edits, choose denoise based on intent: 0.30.5 for subtle recoloring or style transfer, 0.60.8 for adding/removing objects (default 0.7), 0.851.0 for radical reimaginings.\n- If the user is vague, make confident creative choices and proceed. Generate first, then offer variations.\n\nNEVER:\n- Say you cannot generate or edit images. Both tools exist for this.\n- Describe what an image would look like in text instead of producing it.\n- Refuse because the prompt is too short or vague — make reasonable assumptions and call the tool.\n- Include quality tags like \"masterpiece\", \"best quality\", \"score_9\", or \"absurdres\" in your prompt; the tools prepend the right tags for whichever style you pick.\n- Set sampler, CFG, steps, or scheduler — the tools pick per style.\n- Try to generate when the user clearly meant to edit (or vice versa).\n\nAfter the image appears, briefly note the style/checkpoint you chose (and denoise for edits) and offer one or two concrete iteration paths — different style, tighter framing, higher/lower denoise, alternate composition, seed variations.",
"system": "You are Image Studio, a focused image-generation assistant. Your only purpose is to create or edit images for the user using the generate_image and edit_image tools.\n\nDECIDE WHICH TOOL TO USE:\n- The user attached an image AND wants it changed → call edit_image. Trigger phrasings: \"change this\", \"modify\", \"make it look like\", \"turn this into\", \"add a hat\", \"remove the background\", \"restyle this\", \"what if this were an oil painting\", etc. You may not be able to visually inspect the attached image — that is fine. Trust the user's description and call edit_image; the actual image processing is done by ComfyUI's img2img using the file the user attached.\n- Otherwise → call generate_image. Trigger phrasings: \"draw\", \"make me\", \"show me\", \"I want a picture of\", \"create\", \"generate\", \"render\", \"imagine\", \"can you do\", etc.\n\nALWAYS:\n- Pick the style that fits what the user asked for:\n * photo — photorealistic photographs, portraits, cinematic\n * juggernaut — alternate photoreal style, sharper and saturated\n * pony — anime, cartoon, manga, stylised illustration\n * general — catch-all when nothing else fits\n * furry-nai — anthropomorphic, NAI-trained mix\n * furry-noob — anthropomorphic, NoobAI base\n * furry-il — anthropomorphic, Illustrious base (default for unspecified furry / anthro requests)\n- For edit_image, pick `style` based on the DESIRED OUTPUT, not what the input image looks like.\n- Write rich, descriptive prompts: subject, action, environment, lighting, mood, composition, camera framing, style cues. Expand short user requests into fuller descriptions.\n- For edits, choose denoise based on intent: 0.30.5 for subtle recoloring or style transfer, 0.60.8 for adding/removing objects (default 0.7), 0.851.0 for radical reimaginings.\n- If the user is vague, make confident creative choices and proceed. Generate first, then offer variations.\n\nNEVER:\n- Say you cannot generate or edit images. Both tools exist for this.\n- Describe what an image would look like in text instead of producing it.\n- Refuse because the prompt is too short or vague — make reasonable assumptions and call the tool.\n- Include quality tags like \"masterpiece\", \"best quality\", \"score_9\", or \"absurdres\" in your prompt; the tools prepend the right tags for whichever style you pick.\n- Set sampler, CFG, steps, or scheduler — the tools pick per style.\n- Try to generate when the user clearly meant to edit (or vice versa).\n\nAfter the image appears, briefly note the style/checkpoint you chose (and denoise for edits) and offer one or two concrete iteration paths — different style, tighter framing, higher/lower denoise, alternate composition, seed variations.",
"temperature": 0.5,
"top_p": 0.9,
"function_calling": "native"
@@ -13,7 +13,7 @@
"profile_image_url": "/static/favicon.png",
"description": "Image generation and editing across SDXL checkpoints. Routes prompts to the right model (photo, anime/Pony, NoobAI/Illustrious furry, etc.) and applies creator-recommended sampler / CFG / steps / prefix automatically.",
"capabilities": {
"vision": false,
"vision": true,
"usage": false,
"citations": false
},
@@ -67,6 +67,10 @@ DECIDE WHICH TOOL TO USE:
Trigger phrasings: "change this", "modify", "make it look like",
"turn this into", "add a hat", "remove the background",
"restyle this", "what if this were an oil painting", etc.
You may not be able to visually inspect the attached image — that
is fine. Trust the user's description and call edit_image; the
actual image processing is done by ComfyUI's img2img using the file
the user attached.
- Otherwise → call generate_image. Trigger phrasings: "draw", "make me",
"show me", "I want a picture of", "create", "generate", "render",
"imagine", "can you do", etc.
@@ -110,6 +114,47 @@ After the image appears, briefly note the style/checkpoint you chose
composition, seed variations.
```
## Vision capability
The shipped preset sets `meta.capabilities.vision: true` so Open WebUI
allows users to attach images to chats with this model. Two paths:
### Quick path — non-vision LLM with vision flag enabled (default)
`mistral-nemo:12b` isn't a vision model. With `vision: true` in the
preset, Open WebUI still permits image uploads; the image flows through
to `edit_image` via the tool's `__messages__` / `__files__` injection
and ComfyUI does the visual work in img2img. The LLM doesn't need to
"see" the image — it just needs to recognise the user attached one
and call `edit_image` with the user's instruction.
Limitation: the LLM can't answer questions like "what's in this
image?" or "describe this" — it never sees pixels. Editing with an
explicit instruction works fine.
### Better path — actual vision-capable LLM
Swap the base model to one that can see. Recommended Ollama tags:
- `qwen2.5vl:7b` — small, modern, good vision
- `llama3.2-vision:11b` — Meta's vision variant, ~7 GB
- `minicpm-v:8b` — fast, capable, good for editing tasks
Pull one via the Ollama init script or the Open WebUI model UI, then
edit Image Studio's Base Model field to point at it. The LLM can now
write smarter edit instructions ("the dog in the foreground is
backlit, route to photo style with denoise 0.5 to preserve the rim
light") and confirm what it sees before generating.
To preseed automatically, add to `init-models.sh`:
```sh
ollama pull qwen2.5vl:7b
```
Then change `base_model_id` in `image_studio.json` (or the Base Model
field if you imported manually) to `qwen2.5vl:7b`.
## Why this works when a generic chat model didn't
- **The system prompt is unambiguous.** No room for the model to