Text AI nodes & helpers
Image Describer
Analyzes a reference image with a vision-capable LLM. Its default Reconstruct mode reverse-engineers the visual into prompts that an Image Generator can use to reproduce the subject, composition, palette, lighting, visible text, and overall style.
Recommended style-copy workflow
- Connect the reference image to
Image Describer.image. - Connect
Image Describer.outtoImage Generator.prompt. - Connect
Image Describer.negativetoImage Generator.negativewhen the selected generator supports a negative prompt. - Connect the same reference image directly to
Image Generator.image(or its reference input). This second image connection is intentional and gives the generator the pixels as well as the reconstructed prompt. - Run Image Describer first, inspect its Prompt/Negative/Analysis/OCR tabs, then run the Image Generator.
Without the direct reference connection, Reconstruct still produces a useful text-to-image prompt, but identity, exact layout, typography, and small details are less likely to match.
Modes
| Mode | What it does | Outputs |
|---|---|---|
| Reconstruct | Performs structured visual reconstruction and produces generator-ready positive and negative prompts. This is the default for newly created nodes. | out, negative, analysis, ocr, bundle |
| Describe | Writes a vivid human-readable description of composition, colors, lighting, mood, and important objects. | out |
| Generate prompt | Produces one general text-to-image prompt for a similar look, without the structured reconstruction pass. | out |
| Tags | Extracts concise visual keywords for search, classification, or later prompt composition. | out |
| Custom | Runs the instruction entered in the node. | out |
Reconstruct controls
| Control | Meaning |
|---|---|
| Quality · Fast | One vision-analysis call using the normalized full image. Use it for quick drafts. |
| Quality · Deep | Analyzes the full image plus detail tiles for large references, then performs a second prompt-optimization call. Use it for final reconstruction. |
| Target · Auto | Detects one Image Generator connected directly from out to prompt and adapts the prompt to that generator's provider/model dialect. If none or several are connected, it safely falls back to Universal. |
| Target · Universal | Produces portable natural-language prompts without assuming a particular generator. |
| Target · Manual | Lets you name the intended provider and model when no target generator is connected. It changes prompt wording only; it does not configure or run that generator. |
| Provider / Model in Properties | Selects the preferred vision LLM that examines the reference. It is separate from the Image Generator that creates the new image. Only vision-capable providers are offered. If an older saved node points to a provider that no longer has credentials, the helper uses the current managed vision model for that run; unsupported or unknown model IDs still stop before any provider call. |
Reconstruct outputs
| Port / tab | Purpose |
|---|---|
out / Prompt | Positive generation prompt describing what should be reproduced. Connect it to Image Generator.prompt. |
negative / Negative | Artifacts and unwanted traits to avoid. Connect it to Image Generator.negative when supported. |
analysis / Analysis | Structured JSON with subjects, spatial composition, camera, lighting, materials, style, post-processing, palette, source dimensions, and resolved target information. Use it for inspection or downstream automation. |
ocr / OCR | Text recognized inside the image, preserving observed language, case, and punctuation as closely as the vision model can read it. |
bundle / Bundle | One versioned JSON object containing the positive prompt, negative prompt, analysis, OCR, source fingerprint/asset metadata, and validated editable layout for the same source image. Use this port for downstream automation so a batch of N references stays N items instead of multiplying parallel N-item streams. Older bundles remain accepted. |
Reconstruct copies a visual direction; it is not a pixel clone. The final similarity still depends on the target generator and its reference-image support. OCR also does not guarantee perfect text rendering in the generated image.
Image Style does not inspect a reference: it adds one of several built-in style and palette presets to a subject prompt. Brand Style manually supplies colors and fonts to an HTML Creative. Use Image Describer → Reconstruct when the style must be learned from an existing image.
The legacy modes are also useful for:
- Accessibility: generate alt text for images.
- Tag extraction: pull keywords for organization or search.
Visual Creative Editor
Use this node when the final creative must preserve the source composition while reproducing exact editable copy. It makes no provider call: all drag, resize, rotate, typography, shape, and inline-text edits are deterministic local scene changes.
Hybrid reconstruction workflow
- Connect the original image to
Image Describer.image. - Connect
Image Describer.bundletoVisual Creative Editor.reconstruction. - Click Create & Open Visual Editor after Image Describer has produced a bundle. The node does not run Image Describer automatically.
- Connect
Visual Creative Editor.outandnegativeto the matching Image Generator prompt inputs. - Connect the original image directly to that Image Generator's image/reference input. Run preflight blocks the workflow before generation when this exact reference path is absent.
- Connect
Visual Creative Editor.layouttoCreative.layout, and connectImage Generator.outtoCreative.image.
The Image Generator receives no literal overlay copy. Its positive prompt removes baked-in words, logos, watermarks, and pseudo-text while reserving geometry-only safe zones; the negative prompt suppresses text and composition drift. Creative then compiles the whitelisted scene server-side and overlays HTML-escaped text and simple shapes at exact normalized coordinates.
The editor stores scenes by source fingerprint inside the node, so drafts survive project save, versioning, clone, and reload. Batch bundles keep one scene and three outputs per fingerprint, subject to the normal batch limit. When Image Describer produces a different source fingerprint, existing edits are preserved and Source changed appears; choose Rebuild from latest explicitly before the new source can run. Raw HTML, arbitrary CSS, scripts, executable media, and external URLs are not accepted in a linked scene.
Prompt Rewriter
Rewrites an Image Describer reconstruction before generation while keeping an auditable change log. Connect Image Describer.bundle to Prompt Rewriter.reconstruction, then connect out and negative to the matching Image Generator inputs. A legacy raw prompt is also accepted as reconstruction input.
- Instructions can be saved on the node or supplied through the
instructionsport. A connected value always takes precedence, and one value broadcasts across every reconstruction item. - Text policy · Remove is the safe default: source copy, prices, logos, watermarks, labels, and pseudo-text are removed while clean zones are retained for deterministic HTML copy.
- Text policy · Replace also reserves safe zones for later replacement copy.
- Text policy · Preserve keeps source text only when the trusted rewrite instructions explicitly request it.
- Reconstruction, analysis, visible text, and OCR are treated as untrusted data. Instructions embedded in the image or OCR never override the connected or saved rewrite instructions.
- The LLM must return exactly one structured response containing a positive prompt, negative prompt, and non-empty change list. Malformed output produces no ports, so a downstream generator is not launched during a full workflow run.
Multiple reconstructions and multiple instruction values form the normal Cartesian combinations, capped by the workflow batch limit. Use the single bundle port rather than connecting Prompt/Negative/Analysis/OCR in parallel.
Video Describer
Same as Image Describer but for video input. The vision LLM analyzes the video's motion, camera work, subjects, transitions, and mood.
Prompt Concatenator
Combines multiple text inputs into a single text output. Unlike the Prompt Composer (which uses a {token} template), the Concatenator is simpler:
- Fixed inputs: text A and text B (always present).
- Dynamic inputs: click + add input to add more text ports.
- Separator: choose what goes between each input (default:
,). - Reactive: the output updates instantly as inputs change. No Run button needed.
Example: Connect three Text nodes with "cinematic", "golden hour", "35mm film" to the three inputs with separator , . Output: cinematic, golden hour, 35mm film.
Output
Designates the final result of a workflow. It's a passthrough node — whatever connects to its input flows through unchanged.
- The connected result is considered the workflow's deliverable.
- Editable label to name the output (e.g., "Final Ad Creative", "Hero Image").
- Supports any kind: text, image, or video.
- Reactive: no Run button, just displays the connected result.
- It is a result marker only; it does not publish a separate app or change workflow permissions.
Example: In a complex graph with multiple branches, connect the final Creative node to an Output node. This marks it as the canonical result for sharing or exporting.