smart_image_gen v0.7.1: rename edit_image arg + parse file id from URL
Two bugs in one screenshot: 1. LLM called edit_image(prompt=..., ...) but the signature was edit_image(edit_instruction=..., ...) — mismatch, missing-arg crash. Renamed the first param to `prompt` so both tools have a matching, predictable name. System prompt updated with an explicit 'do not invent edit_instruction' line for stubborn models. 2. After fix #1, edit_image still couldn't find the prior generated image because Open WebUI assistant-message file attachments only carry {type, url} (no id, no path). _read_file_dict now also greps the file id out of /api/v1/files/<uuid>/content URLs and feeds it to Files.get_file_by_id. Verified pattern matches absolute URLs (https://llm-1.srvno.de/api/v1/files/.../content). System prompt also now says 'including images you previously generated in this chat' to nudge the LLM to pick up assistant outputs as edit candidates. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -4,7 +4,7 @@
|
||||
"base_model_id": "huihui_ai/qwen3.5-abliterated:9b",
|
||||
"name": "Image Studio",
|
||||
"params": {
|
||||
"system": "/no_think\n\nYou are an image-tool dispatcher. You do not respond in prose. Every user message MUST result in exactly one tool call.\n\nROUTING:\n- If the user attached an image → call edit_image\n- Otherwise → call generate_image\n\nFire the tool on the FIRST message, with no preamble. Do not write a 'plan', 'approach', 'steps', 'breakdown', or any explanation before calling. Do not ask clarifying questions. Do not say what you are about to do. If the request is vague, pick reasonable defaults and call the tool — the user iterates after.\n\nSTYLES (pick one):\n photo photorealistic photo / portrait / cinematic\n juggernaut alternate photoreal — sharper, more saturated\n pony anime, cartoon, manga, stylised illustration\n general catch-all when nothing else fits\n furry-nai anthropomorphic, NAI-trained mix\n furry-noob anthropomorphic, NoobAI base\n furry-il anthropomorphic, Illustrious base (default for any furry/anthro request)\n\nedit_image has TWO MODES — pick based on whether the change is local or global:\n- LOCAL change (\"change the ball to a basketball\", \"add a hat to the dog\", \"remove the bird\", \"recolor the car red\") → set `mask_text` to a brief noun phrase naming the region (\"the ball\", \"the dog\", \"the bird\", \"the car\"). Only that region is repainted; rest stays pixel-perfect.\n- GLOBAL change (\"make this a sunset\", \"turn this into anime\", \"restyle as oil painting\") → leave mask_text unset. The whole image is reimagined.\nALWAYS prefer LOCAL when the user names a specific object, person, or region. GLOBAL is only for whole-image style/lighting transformations.\n\nDenoise:\n- LOCAL (mask_text set): default 1.0. Drop to 0.6–0.8 only for subtle local edits that should retain some original structure.\n- GLOBAL (no mask_text): default 0.7. Use 0.3–0.5 for subtle restyle, 0.85–1.0 for radical reimagining.\n\nPick style for the DESIRED OUTPUT, not the input image.\n\nWrite rich, descriptive prompts (subject, action, environment, lighting, mood, framing). Do NOT add quality tags like 'masterpiece', 'best quality', 'score_9', 'absurdres' — the tool prepends the correct tags per style. Do NOT set sampler, CFG, steps, scheduler — the tool picks them.\n\nAFTER the tool returns, write at most one short sentence noting your style/mode choice and offering one iteration idea. The image is already shown to the user; do not describe it.",
|
||||
"system": "/no_think\n\nYou are an image-tool dispatcher. You do not respond in prose. Every user message MUST result in exactly one tool call.\n\nROUTING:\n- If the user attached an image (including images you previously generated in this chat) → call edit_image(prompt=..., ...)\n- Otherwise → call generate_image(prompt=..., ...)\nBoth tools take `prompt` as the first argument — same name on both. Do NOT invent `edit_instruction`.\n\nFire the tool on the FIRST message, with no preamble. Do not write a 'plan', 'approach', 'steps', 'breakdown', or any explanation before calling. Do not ask clarifying questions. Do not say what you are about to do. If the request is vague, pick reasonable defaults and call the tool — the user iterates after.\n\nSTYLES (pick one):\n photo photorealistic photo / portrait / cinematic\n juggernaut alternate photoreal — sharper, more saturated\n pony anime, cartoon, manga, stylised illustration\n general catch-all when nothing else fits\n furry-nai anthropomorphic, NAI-trained mix\n furry-noob anthropomorphic, NoobAI base\n furry-il anthropomorphic, Illustrious base (default for any furry/anthro request)\n\nedit_image has TWO MODES — pick based on whether the change is local or global:\n- LOCAL change (\"change the ball to a basketball\", \"add a hat to the dog\", \"remove the bird\", \"recolor the car red\") → set `mask_text` to a brief noun phrase naming the region (\"the ball\", \"the dog\", \"the bird\", \"the car\"). Only that region is repainted; rest stays pixel-perfect.\n- GLOBAL change (\"make this a sunset\", \"turn this into anime\", \"restyle as oil painting\") → leave mask_text unset. The whole image is reimagined.\nALWAYS prefer LOCAL when the user names a specific object, person, or region. GLOBAL is only for whole-image style/lighting transformations.\n\nDenoise:\n- LOCAL (mask_text set): default 1.0. Drop to 0.6–0.8 only for subtle local edits that should retain some original structure.\n- GLOBAL (no mask_text): default 0.7. Use 0.3–0.5 for subtle restyle, 0.85–1.0 for radical reimagining.\n\nPick style for the DESIRED OUTPUT, not the input image.\n\nWrite rich, descriptive prompts (subject, action, environment, lighting, mood, framing). Do NOT add quality tags like 'masterpiece', 'best quality', 'score_9', 'absurdres' — the tool prepends the correct tags per style. Do NOT set sampler, CFG, steps, scheduler — the tool picks them.\n\nAFTER the tool returns, write at most one short sentence noting your style/mode choice and offering one iteration idea. The image is already shown to the user; do not describe it.",
|
||||
"temperature": 0.5,
|
||||
"top_p": 0.9,
|
||||
"function_calling": "native",
|
||||
|
||||
@@ -65,8 +65,11 @@ You are an image-tool dispatcher. You do not respond in prose. Every
|
||||
user message MUST result in exactly one tool call.
|
||||
|
||||
ROUTING:
|
||||
- If the user attached an image → call edit_image
|
||||
- Otherwise → call generate_image
|
||||
- If the user attached an image (including images you previously
|
||||
generated in this chat) → call edit_image(prompt=..., ...)
|
||||
- Otherwise → call generate_image(prompt=..., ...)
|
||||
Both tools take `prompt` as the first argument — same name on both.
|
||||
Do NOT invent `edit_instruction`.
|
||||
|
||||
Fire the tool on the FIRST message, with no preamble. Do not write a
|
||||
'plan', 'approach', 'steps', 'breakdown', or any explanation before
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
"""
|
||||
title: Smart Image Generator & Editor (ComfyUI)
|
||||
author: ai-stack
|
||||
version: 0.7.0
|
||||
version: 0.7.1
|
||||
description: Generate or edit images via ComfyUI with automatic SDXL
|
||||
checkpoint routing. Two methods — generate_image (txt2img) and
|
||||
edit_image (img2img on the user's most recently attached image). The
|
||||
@@ -345,12 +345,19 @@ def _file_dict_is_image(f: dict) -> bool:
|
||||
return "image" in ftype or fname.endswith((".png", ".jpg", ".jpeg", ".webp"))
|
||||
|
||||
|
||||
_FILE_URL_ID_RE = re.compile(r"/(?:api/v1/)?files/([0-9a-fA-F-]{8,})(?:/content)?")
|
||||
|
||||
|
||||
def _read_file_dict(f: dict) -> Optional[bytes]:
|
||||
"""
|
||||
Try to read raw bytes for one file dict. Path keys first (covers local
|
||||
uploads), then Open WebUI's Files model lookup by id (covers assistant-
|
||||
emitted images that only have an id + relative URL). Returns None if
|
||||
no method worked.
|
||||
Try to read raw bytes for one file dict. Tries in order:
|
||||
1. Local filesystem path keys (covers user uploads with `path`).
|
||||
2. Open WebUI's Files.get_file_by_id with f["id"] (covers files
|
||||
the user uploaded via the file API).
|
||||
3. Same lookup with the id parsed out of f["url"] (covers
|
||||
assistant-emitted files where the message attachment is just
|
||||
{"type":"image","url":"/api/v1/files/<uuid>/content"} —
|
||||
no id field, no path field, but the URL has the id).
|
||||
"""
|
||||
for path_key in ("path", "filepath", "file_path"):
|
||||
path = f.get(path_key)
|
||||
@@ -361,12 +368,21 @@ def _read_file_dict(f: dict) -> Optional[bytes]:
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
fid = f.get("id")
|
||||
if _OPENWEBUI_RUNTIME and fid:
|
||||
try:
|
||||
file_model = Files.get_file_by_id(fid)
|
||||
if file_model is not None:
|
||||
# FileModel may expose path directly or under .meta
|
||||
candidate_ids = []
|
||||
if f.get("id"):
|
||||
candidate_ids.append(f["id"])
|
||||
url = f.get("url")
|
||||
if url:
|
||||
m = _FILE_URL_ID_RE.search(url)
|
||||
if m:
|
||||
candidate_ids.append(m.group(1))
|
||||
|
||||
if _OPENWEBUI_RUNTIME:
|
||||
for fid in candidate_ids:
|
||||
try:
|
||||
file_model = Files.get_file_by_id(fid)
|
||||
if file_model is None:
|
||||
continue
|
||||
path = getattr(file_model, "path", None)
|
||||
if not path:
|
||||
meta = getattr(file_model, "meta", None) or {}
|
||||
@@ -380,8 +396,8 @@ def _read_file_dict(f: dict) -> Optional[bytes]:
|
||||
return fh.read()
|
||||
except OSError:
|
||||
pass
|
||||
except Exception:
|
||||
pass
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
return None
|
||||
|
||||
@@ -713,7 +729,7 @@ class Tools:
|
||||
|
||||
async def edit_image(
|
||||
self,
|
||||
edit_instruction: str,
|
||||
prompt: str,
|
||||
style: Optional[StyleName] = None,
|
||||
mask_text: Optional[str] = None,
|
||||
denoise: Optional[float] = None,
|
||||
@@ -757,7 +773,7 @@ class Tools:
|
||||
|
||||
Pick `style` for the DESIRED OUTPUT, not the input image.
|
||||
|
||||
:param edit_instruction: What the changed area should look like.
|
||||
:param prompt: What the changed area should look like.
|
||||
Tool auto-prepends quality tags — don't include those.
|
||||
:param style: One of the StyleName values. Omit to auto-detect.
|
||||
:param mask_text: Noun phrase describing the region to edit. Set
|
||||
@@ -768,7 +784,7 @@ class Tools:
|
||||
:param seed: 0 to randomize, otherwise specific.
|
||||
:return: Markdown image of the result, or an error if no image is attached.
|
||||
"""
|
||||
chosen = style or _route_style(edit_instruction)
|
||||
chosen = style or _route_style(prompt)
|
||||
settings = STYLES.get(chosen)
|
||||
if not settings:
|
||||
return f"Unknown style '{chosen}'. Available: {', '.join(STYLES.keys())}"
|
||||
@@ -809,7 +825,7 @@ class Tools:
|
||||
+ (f", mask='{mask_text}'" if mask_text else "")
|
||||
)
|
||||
|
||||
positive = f"{settings['prefix']}{edit_instruction}"
|
||||
positive = f"{settings['prefix']}{prompt}"
|
||||
negative = settings["negative"]
|
||||
if negative_prompt:
|
||||
negative = f"{negative}, {negative_prompt}"
|
||||
|
||||
Reference in New Issue
Block a user