From 698136c37e7981ae4cf7ee20ee1729b51c8c227c Mon Sep 17 00:00:00 2001 From: William Gill Date: Sat, 6 Jun 2026 20:15:31 -0500 Subject: [PATCH] smart_image_gen v0.7.11: docstrings steer follow-up edits to edit_image MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit When the user generated an image and then asked to modify it ("make her hair red", "now at sunset"), the OWUI default-mode tool decider was picking generate_image and producing a fresh, unrelated image — or, if the base model output the call as text instead, narrating generate_image(...) in the chat instead of dispatching anything. Root cause was in the tool docstrings, which the decider weights heavily. Both were asymmetric in the same direction: - edit_image led with "an image the user has ATTACHED to the chat" and "the user uploads an image" — both phrasings exclude assistant- emitted images, so the decider read follow-up turns as "no source, edit_image invalid." - generate_image's exclusion clause ("they have NOT attached an existing image") matched: assistant-emitted images aren't "attached," so generate_image stayed valid for follow-ups too. Result: on iteration turns, the decider saw generate_image as the only valid choice and dispatched it (or the base model emitted a pseudo-call when the decider declined). Rewrite both leads symmetrically: - edit_image now covers "any image already in this chat" with explicit mention of assistant-emitted sources. - generate_image now defers to edit_image whenever ANY image is visible above, even when the user's phrasing sounds like a fresh request — that last clause is what catches "make her hair red." The extraction code itself already handled assistant-emitted images correctly (path #2 + #4 in _extract_attached_image, including the chat-DB fallback from f26dfbe) — only the docstrings were lying to the decider. smart_image_gen.py 0.7.10 -> 0.7.11 Co-Authored-By: Claude Opus 4.7 (1M context) --- .../openwebui-tools/smart_image_gen.py | 31 +++++++++++++------ 1 file changed, 21 insertions(+), 10 deletions(-) diff --git a/deployments/ai-stack/openwebui-tools/smart_image_gen.py b/deployments/ai-stack/openwebui-tools/smart_image_gen.py index 2b94ea2..526b0f5 100644 --- a/deployments/ai-stack/openwebui-tools/smart_image_gen.py +++ b/deployments/ai-stack/openwebui-tools/smart_image_gen.py @@ -1,7 +1,7 @@ """ title: Smart Image Generator & Editor (ComfyUI) author: ai-stack -version: 0.7.10 +version: 0.7.11 description: Generate or edit images via ComfyUI with automatic SDXL checkpoint routing. Two methods — generate_image (txt2img) and edit_image (img2img on the user's most recently attached image). The @@ -754,11 +754,17 @@ class Tools: ) -> str: """ Create a NEW image from scratch and show it to the user. Use this - whenever the user asks you to draw, generate, create, make, paint, - render, or imagine any visual content — photographs, portraits, - characters, scenes, illustrations, anime, drawings — and they have - NOT attached an existing image. If they did attach an image and - want it modified, use edit_image instead. + ONLY when there is no existing image in the chat at all and the + user is asking for a fresh image (draw, generate, create, make, + paint, render, imagine) — photographs, portraits, characters, + scenes, illustrations, anime, drawings. If ANY image is already + visible in the chat (either user-attached OR one you generated + earlier in this conversation) and the user is asking to change, + modify, restyle, recolor, add to, remove from, or iterate on it, + use edit_image instead — even if the user's word choice sounds + like a fresh request ("make her hair red", "now show her at + sunset"). Follow-up requests on an existing image are ALMOST + ALWAYS edits, not new generations. Pick `style` to match what the user wants: - "photo" — photorealistic photographs, portraits, cinematic shots. @@ -854,9 +860,13 @@ class Tools: __event_emitter__: Optional[Callable[[dict], Awaitable[None]]] = None, ) -> str: """ - Edit, modify, transform, or restyle an image the user has ATTACHED - to the chat. Use whenever the user uploads an image and asks to - change it. If no image is attached, use generate_image instead. + Edit, modify, transform, or restyle an image already in this + chat — either one the user attached, OR one you (the assistant) + generated earlier in this same conversation via generate_image + or edit_image. Use whenever the user asks to change, modify, + restyle, recolor, add to, remove from, or iterate on an image + that is already visible above in the chat. If no image exists + in the chat at all, use generate_image instead. TWO MODES — choose based on whether the change is local or global: @@ -905,7 +915,8 @@ class Tools: 1.0 with mask_text, 0.7 without. :param negative_prompt: Extra terms to exclude. Usually unneeded. :param seed: 0 to randomize, otherwise specific. - :return: Markdown image of the result, or an error if no image is attached. + :return: Markdown image of the result, or an error if no image + exists anywhere in this chat (user-uploaded or assistant-generated). """ # Resolve style — inheritance DOMINATES for edits. Vision LLMs # misclassify subject types (observed in the wild: juggernaut