All nodes reference
This is the complete reference for the 77 node types available on the canvas. The server also accepts legacy aliases that are not available in the current editor; they are listed separately below. A port marked with * must have a value before the node can execute. Some generator requirements are conditional on the selected provider or model.
Cloud generator graph contract
For imageGenerator and videoGenerator, select provider/model ids and supported parameters from generation_catalog. Put provider-specific settings in data.params, using the catalog's exact parameter keys and value types. Aspect ratio is commonly data.params.aspect_ratio; do not invent a top-level data.aspectRatio field. Put an inline prompt in data.prompt, or connect a text node's out to the generator's prompt. Video duration in seconds may use data.duration; supported duration values come from the selected model.
For a reference imported with asset_import, create an importImage node with data.assetId, data.url and data.mime from the returned owned asset. Connect its out to the generator's image input. Video generators also accept endImage for a supported last-frame mode. Models may have additional conditional reference requirements; the server preflight validates them.
Create graph nodes with unique id, registered type, position: {x, y} and data. Edges need unique id, source, sourceHandle, target and targetHandle. For an existing project, use incremental operations with its latest updatedAt.
Value kinds
| Kind | Can connect to |
|---|---|
text | Text, prompt, number, toggle and seed inputs |
image | Image inputs |
video | Video inputs |
audio | Audio inputs |
3d | 3D results and downloads |
media | Image or video |
any | Any supported value kind |
One concrete input handle normally accepts one connection. Text {{token}} handles and dynamically added compositor or concatenator handles are separate inputs, not multiple wires into the same handle.
Text
| Node | Ports | Behaviour |
|---|---|---|
Text / Prompt (text) | Dynamic {{token}} inputs → out:text | Stores text or a system prompt. Connected token streams are substituted; multiple values produce a Cartesian batch. |
Variable (variable) | in:any → out:any | Emits its manual value, or passes through the connected stream. |
Concatenator (promptConcatenator) | a:text, b:text, dynamic inputs → out:text | Joins connected values with the configured separator. |
Output (output) | in:any* → no wire output | Final-result marker and passthrough display. It does not publish or export a workflow by itself. |
Iterators and lists
| Node | Ports | Behaviour |
|---|---|---|
Text Iterator (textIterator) | in:any → out:text | Merges manual rows, CSV/TXT imports and connected values into a batch. Supports reorder, sort and deduplication. |
Image Iterator (imageIterator) | in:image → out:image | Merges uploaded and connected images; every item continues separately downstream. |
Video Iterator (videoIterator) | in:video → out:video | Video equivalent of Image Iterator. |
Array (array) | in:any → out:text | Splits text by a delimiter; can trim, remove empty values, deduplicate and sort. |
List Selector (listSelector) | in:any* → out:any | Selects an item by index, first, last or random mode. |
A/B Matrix (abMatrix) | hooks:text, visuals:text, ctas:text → out:text | Produces every hook × visual × CTA combination, subject to the workflow batch limit. |
AI and text helpers
| Node | Ports | Behaviour |
|---|---|---|
Run any LLM (runLlm) | system:text, user:text, dynamic image inputs → out:text | Calls the chosen LLM. Multiref packs images into one vision request; standard mode fans them out. |
Prompt Enhancer (promptEnhancer) | subject:text → out:text | Expands a short subject using the selected style and detail level. |
Translate (promptTranslate) | in:text → out:text | Translates incoming text to English through an LLM. |
Prompt Rewriter (promptRewriter) | reconstruction:text*, instructions:text → out:text, negative:text, changes:text | Strictly rewrites a reconstruction bundle or raw prompt. Connected instructions override saved instructions; untrusted OCR cannot change the task. Invalid structured output fails closed. |
Prompt Composer (promptComposer) | subject, style, palette, camera, lighting → out:text | Reactively fills a single-brace {token} template. |
Image Describer (imageDescriber) | image:image* → out:text, negative:text, analysis:text, ocr:text, bundle:text | Reconstruct mode returns generator-ready prompts, structured visual analysis, OCR, and one linked bundle per image through strict vision calls. Legacy Describe/Prompt/Tags/Custom modes keep the original single text output. |
Visual Creative Editor (visualCreativeEditor) | reconstruction:text* → out:text, negative:text, layout:text | Opens a project-scoped, code-free scene from Image Describer.bundle. Moving or editing OCR layers deterministically rebuilds text-free safe-zone prompts and a validated HTML layout without another LLM call. A new source fingerprint remains stale until explicitly rebuilt. |
Video Describer (videoDescriber) | video:video* → out:text | Samples video frames and asks a vision-capable LLM to describe them. |
Transcribe Audio (transcriber) | audio:audio or video:video → out:text | Extracts speech from one connected audio or video source. |
Script → Storyboard (storyboard) | script:text → out:text | Returns a JSON shot list with prompts, durations, on-screen text and narration. |
Ad Copy Bundle (adCopyBundle) | topic:text → out:text | Returns structured ad copy: creative text, caption, hashtags and CTA. |
Video Hook (videoHook) | topic:text → out:text | Generates a short first-seconds hook for a video. |
Imports, search and image services
| Node | Ports | Behaviour |
|---|---|---|
Import Image (importImage) | no input → out:image | Emits the uploaded image asset. |
Import Video (importVideo) | no input → out:video | Emits the uploaded video asset. |
Import Audio (audioImport) | no input → out:audio | Emits the uploaded audio asset. |
Find Images (findImages) | query:text → out:image | Searches through Serper, safely downloads results and emits the valid images. Query may also be entered in the node. |
Upscaler (upscaler) | in:image* → out:image | Uses Atlas image upscaling when configured, otherwise bicubic resizing. |
Photo Restore (photoRestore) | in:image* → out:image | Uses Atlas photo cleanup when configured; the local fallback only applies light denoising. |
Face Swap (faceSwap) | target:image*, source:image* → out:image | Sends the target and source images to the Atlas face-swap service. |
Generators and creatives
| Node | Ports | Behaviour |
|---|---|---|
Text Generator (textGenerator) | system:text, user:text, dynamic image inputs → out:text | Generator-labelled form of the LLM node. |
Image Generator (imageGenerator) | prompt:text, negative:text, image references → out:image | Runs a Nova GPU workflow or a configured hosted image provider. Input requirements depend on the chosen model. |
Video Generator (videoGenerator) | prompt, negative, image, end image, audio, video → out:video | Supports t2v, i2v, first/last-frame, reference and provider-specific video modes. Also runs the local avatar models — see Talking avatars below. |
Novastorm Avatar (novastormAvatar) | script:text*, character:image, voice:audio → out:video | Writes a talking UGC-creator video from a script on your own GPU (MiniMax H3 UGC). Long videos are split into segments that keep the same person and voice, each continuing from the end of the previous segment, and merged into one video with sound. * The script can also be typed on the node, and a voice sample can be uploaded on it. See Novastorm Avatar below. |
Video Editor (videoEditor) | video:video*, prompt:text → out:video | Sends an existing clip and edit prompt to a hosted video-edit model. |
Audio Generator (audioGenerator) | prompt:text → out:audio | Runs TTS, music or sound-effect models exposed by configured providers. |
3D Generator (generator3d) | prompt:text, image references → out:3d | Creates a downloadable 3D result with an Atlas-routed 3D provider. |
Creative (creative) | image plus template text inputs, or layout:text → out:image | Fills an HTML creative template and renders PNG variants in Chromium. A connected validated Visual Creative layout takes precedence over templateId and composites its exact text/shapes over the connected generated background. |
Video Creative (videoCreative) | image/video, template text, Overlay Motion (motion) → out:video | Renders an animated HTML creative and encodes it as MP4. |
Generator (generator) | prompt, negative, image → out:image | Compatibility node for older saved workflows. New workflows use the dedicated generator types. |
Talking avatars
Lip-syncing a portrait to an existing speech track is not a separate node type. It is the ordinary
Video Generator with provider set to nova-avatar, which runs an audio-driven
avatar model on one of your own connected GPUs. Wire a portrait into image and a
speech track into audio; prompt is optional and directs motion rather than
describing the scene. The result is a clip whose length matches the audio. To have a
person say a written script in a generated voice, use Novastorm Avatar instead.
model | Notes |
|---|---|
aptavatar | Portrait plus speech with gesture direction. Native 704x1280 at 25 fps. Accepts a structured prompt that assigns a motion to a frame range. |
longcat-avatar-15 | Portrait plus speech. 25 fps, frame size derived from the portrait's aspect ratio. Set params.segments to cover longer speech; it is otherwise computed from the audio. |
Node data: provider: "nova-avatar", model, gpuNodeId (required), seed
(below zero picks one), optional params. There is no templateId — these models
ship their own inference runtime rather than a workflow.
Three things must be true before a run succeeds, and each reports a specific error:
the GPU node must be connected, its weight bundle must be installed from the model
catalog (aptavatar-bf16 or longcat-avatar-15-bf16), and that model's Python
runtime must be provisioned on the node. Connecting several prompts fans out into
one take per prompt against the same face, voice and seed.
Generation on your own GPU spends no credits. Renting the card is billed separately by the minute.
Novastorm Avatar (talking UGC creators from a script)
Novastorm Avatar (novastormAvatar) turns a written script into a selfie-style video of a
person who says it, with a generated voice and the real background sound of the place. It runs
the MiniMax H3 UGC workflows on one of your connected GPUs, so it spends no credits. The node is
available to admin and designer accounts.
| Port | Kind | Meaning |
|---|---|---|
script | text | What the person says. Required unless data.script is set; a connected value wins. Several scripts produce one video each. |
character | image | Optional photo of the person. The video keeps this face; without it the person is generated from person. |
voice | audio | Optional voice sample (audio/*, at most 16 MB; 5–15 s of real speech). It sets how the generated voice sounds and works with or without a photo. It is not played back or lip-synced. Instead of connecting one, the user can upload a file on the node (voiceUrl); a connected sample wins, and the upload is used again after disconnecting. |
out | video | One finished video with sound per script × photo × voice combination. |
| Data field | Default | Values |
|---|---|---|
gpuNodeId | — | Required. An online GPU with catalog model minimax-h3-ugc-int8 installed (workflow_gpu_nodes). |
script | "" | Used when nothing is connected to script. |
durationMode | "auto" | "auto" sizes the video to the script (about 3.2 words per second, 5–60 s); "fixed" runs at most duration, and a script that doesn't fill it is planned like "auto". Either way, a segment its words don't fill at about 3.2 words per second is shortened to them. |
duration | 10 | Seconds, 5–60, when durationMode is "fixed". |
maxSegmentSeconds | 0 | Longest single take: 0 (Auto) and 10 use 10 s; 15 needs a 48 GB+ card, and H3 now and then cuts to another shot inside a 15 s take. |
scene | "street" | street, kitchen, livingRoom, hallway, park, cafe, car, office, gym, custom. |
sceneDetails | "" | Extra details for a preset; for custom, the whole scene, starting with a verb ("walks along a pier at sunset…"). |
sounds | "" | Background sounds; empty uses the scene's own soundscape. |
person | a woman in her late twenties with shoulder-length brown hair and a cream knit sweater | Age, hair and clothes in one line. Repeated in every segment so the outfit survives the cuts. With a photo, describe only clothes and props. |
speaker | "woman" | woman or man. |
voiceStyle | "relaxed" | relaxed, warm, upbeat, calm, custom (text in voiceCustom). Ignored when a voice sample is connected or uploaded. |
voiceCustom | "" | Voice description for voiceStyle: "custom". |
voiceUrl | "" | A voice sample uploaded on the node (/api/assets/<id>). Used only when nothing is connected to voice. Written by the canvas upload; agents cannot set it and should connect an audioImport node (url from workflow_upload_asset) to voice instead. |
voiceName | "" | The uploaded file's name (display only). |
voiceMime | "" | Its media type, audio/*. |
voiceSeconds | 0 | Its length in seconds as measured by the browser; 0 = unknown. Display only. |
expression | "calm" | calm, friendly, serious, amused. |
language | "en" | Spoken language. Only en is verified; es, fr, de, it, pt, pl, ru, zh are experimental. |
format | "9:16" | "9:16" (768×1344) or "16:9" (1344×768, not yet tested). |
seed | -1 | -1 picks a random seed per segment; a fixed seed repeats the result (segment k uses seed + k − 1). |
draft | false | true renders at half size (384×672 for 9:16, 672×384 for 16:9), about 5× faster on an RTX 5090, with the same segments, prompts and establishing shot — to check the script, timing and scene before the final video. |
How it works. One MiniMax H3 take is at most 10 s (15 s when chosen, on a 48 GB+ card).
Longer videos are split at sentence boundaries into segments of about equal words, each as long as its words need;
when that would crowd one segment with more words than the longest segment holds at 3.2 words per second, the split
that leaves the densest segment the fewest words is used instead (a segment of 35 words in 10 s dropped a clause).
Segment 1 is text-to-video,
or a reference clip when a photo or voice sample is connected or uploaded. Every later segment is a reference
clip that starts on the previous segment's last five frames and the sound under them (anchored as its
first frames; the last one is also <Picture 1>), so the motion, framing, pose, background and room
sound carry on across the cut instead of being re-imagined (from one still frame H3 now and then
swapped the background within the first quarter second); it keeps the photo as <Picture 2> and speaks with the voice sample (connected or uploaded) or segment 1's
own voice (<Audio 1>). The person description is restated in every segment so the clothes stay
the same. The segments are merged with each one's own sound: the repeated frames are dropped, and
colour and loudness are matched at each seam. A video whose first words come after 0.5 s (the
speaker sometimes waited up to 1 s) starts 0.25 s before them instead; pauses between segments stay. Finished segments are saved in the Library, so a
failed run does not lose them. With a photo, segment 1 would start straight from the portrait, and
H3 then sometimes cut to another shot or morphed a prop inside it; so the node first renders a short
establishing shot (5 s, not kept) in which the person says the first few words of the script and goes
quiet, and pins segment 1 to its last steady frame and the four before it (not their sound), the way
every later segment is pinned to the one before it. If that shot fails, or segment 1 has fewer than 13 words (under 2.4 words per second even
in a 5 s segment), segment 1 starts from the photo as before. A pinned segment has to talk from its
first frame: given more time than its words need, H3 padded them with repeated or made-up words. So
every segment is shortened to its words (about 3.2 per second, to the nearest length H3 can render:
rounding up left up to 0.7 s of slack, which a take filled with a made-up phrase), a sparse Fixed-length script makes a
shorter video, and every segment after the first gets at least 13 words (the split falls back to a
clause or word boundary when a sentence boundary would leave fewer). People who move
behind the speaker can pop in and out in any segment, so the café, office and gym presets ask for a
few people seated, still and out of focus, and the café cup is an opaque one that stays on the table
(a glass in the hand was bent out of shape). In a custom scene or scene details, don't ask for a busy
or crowded place, for people passing behind the speaker, or for glassware in the speaker's hand.
In a local test on an RTX 5090, a voice sample raised the similarity of H3's voice to the reference
from 0.77 to 0.90. A real recording works best; voices cloned with text-to-speech sounded artificial.
The prompt follows the tested "natural UGC" recipe: a handheld front-camera selfie, the person moving with the phone, a voice described in words, and a loud soundscape of the place. Write the script as a short first-person story with one everyday detail, about 3 words per second. Do not use stage directions: text in brackets is spoken aloud.
Timing on an RTX 5090 at 768×1344: 5 s ≈ 4 min, 10 s ≈ 10 min, 20 s (2 segments) ≈ 20 min,
30 s (3 segments) ≈ 31 min, plus about 4 min for the establishing shot when a photo is connected.
Each segment gets its own time budget, so long videos and batches
aren't stopped at the usual 30-minute limit; a run still stops after 24 h — split very large
batches. Common errors name their fix: the bundle is not installed (install minimax-h3-ugc-int8
from the model catalog), the GPU ran out of memory (set maxSegmentSeconds to 10 or free VRAM),
the GPU is offline, or the GPU agent is too old for videos longer than one segment (update it to
0.3.15 or newer).
Example node for the workflow MCP:
{ "type": "novastormAvatar", "data": { "gpuNodeId": "<id from workflow_gpu_nodes>",
"script": "I walk past this coffee place every single morning, and today I finally went in.",
"durationMode": "auto", "duration": 10, "maxSegmentSeconds": 0, "scene": "street",
"sceneDetails": "", "sounds": "", "person": "a woman in her late twenties with shoulder-length brown hair and a cream knit sweater",
"speaker": "woman", "voiceStyle": "relaxed", "voiceCustom": "", "expression": "calm",
"language": "en", "format": "9:16", "seed": -1 } }
Local GPU workflows (ComfyUI templates)
The Image Generator and Video Generator also run Nova's ComfyUI workflow
templates on one of your connected GPUs. Node data: provider: "comfyui",
templateId, gpuNodeId, optional params and, for video, duration in seconds.
The MCP tool workflow_gpu_templates lists the templates. For each one it gives the
catalog model to install, the minimum VRAM, the inputs it reads and its parameter
defaults.
promptandnegativeports (ordata.prompt) feed the template's text inputs. Image and audio inputs come from the port with the same key. Number inputs come fromparams.<key>; a seed below zero picks one.- The template's catalog model must be installed on the node. Otherwise the run fails before it starts and names the missing files.
- MiniMax H3 templates generate speech together with the picture; write the line
in the prompt as
(S1) says: <d>[English] …</d>. On first-frame templates (…-i2v) anendImageport adds a last frame. Reference templates (…-r2v) take up to six images onimage,image1…image5(<Picture 1>…<Picture 6>), and a speech track onaudiobecomes the voice reference<Audio 1>. The generated voice then matches that timbre, so one persona keeps its voice across clips. builtin-minimax-h3-ugc-*templates (bundleminimax-h3-ugc-int8) are tuned for talking UGC creators: vertical 768×1344, 8 fast steps and a realism LoRA.
Datatypes and canvas-only nodes
| Node | Ports | Behaviour |
|---|---|---|
Number (number) | no input → out:text | Emits a number clamped to the configured minimum and maximum. |
Toggle (toggle) | no input → out:text | Emits true or false. |
Seed (seed) | no input → out:text | Shows one stable random value. It changes only when the dice button is clicked; locking preserves that value. |
Sticky Note (sticky) | no ports | Canvas annotation; skipped during execution. |
Group (group) | no ports | Resizable canvas frame; skipped during execution. |
Image editing and compositing
| Node | Ports | Behaviour |
|---|---|---|
Crop (crop) | in:image* → out:image | Center-crops to a preset ratio or custom dimensions. |
Resize (resize) | in:image* → out:image | Resamples to exact width and height. |
Blur (blur) | in:image* → out:image | Applies Gaussian blur. |
Invert (invert) | in:image* → out:image | Inverts image colours. |
Grayscale (grayscale) | in:image* → out:image | Converts an image to grayscale. |
Compare (compare) | a:image*, b:image* → out:image | Slider/toggle preview; forwards the selected A or B image. |
Compositor (compositor) | layer1:image, layer2:image, dynamic layers → out:image | Composites visible image layers using position, scale, opacity and blend mode. The first visible layer defines canvas size. |
Mask by Color (maskByColor) | in:image* → out:image | Builds a black-and-white mask using the centre pixel and tolerance. |
Remove BG (removeBackground) | in:image* → out:image | Uses Atlas background removal when available, otherwise corner-colour detection. |
AI Fill (aiFill) | image:image*, mask:image*, prompt:text → out:image | Runs prompt-based image editing. The current provider request does not transmit the mask bytes; see Known limitations below. |
Outpaint (outpaint) | image:image*, prompt:text → out:image | Extends an image using direction and amount. The current server-generated instruction is used instead of the connected prompt. |
Object Eraser (objectEraser) | image:image*, mask:image* → out:image | Runs prompt-based removal. The mask is currently validated and read but is not transmitted to the provider. |
Relight (relight) | image:image*, prompt:text → out:image | Re-lights an image through an edit provider using the connected or configured prompt. |
Color Grade (colorGrade) | in:media* → out:media | Applies a preset and manual brightness/contrast/saturation; supports image and video. |
Multi-Platform Resize (multiResize) | image:image* → out:image | Emits padded variants for selected Facebook, Instagram and Google ad sizes. |
Video and audio processing
| Node | Ports | Behaviour |
|---|---|---|
Extract Last Frame (extractFrame) | in:video* → out:image | Extracts the final video frame with ffmpeg. |
Extract Audio (extractAudio) | in:video* → out:audio | Outputs a video's soundtrack as MP3, e.g. a generated clip's speech as the voice reference for the next MiniMax H3 …-r2v clip. Fails clearly on a silent video. |
Thumbnail (thumbnailExtract) | video:video* → out:image | Extracts a frame at a timestamp; automatic mode tries one second, then the first frame. |
Merge Videos (concatVideo) | video1, video2, video3 → out:video | Normalizes and concatenates at least two clips. seamMode: "cut" is the default; "continuity" removes each later clip's conditioning frame and locally retimes its first 0.5 s. targetDuration: 0 keeps automatic length. audioMode: "drop" (default) outputs the picture only; "keep" carries every clip's own sound under its own frames (silence for silent clips), so talking segments stay in sync. With "keep", continuity only drops the conditioning frame and skips the head retiming. |
Video Trim (videoTrim) | video:video* → out:video | Keeps the configured start/end interval. |
Attach Audio (attachAudio) | video:video*, audio:audio* → out:video | Replaces or attaches the audio track. |
Audio Mix (audioMix) | audio1, audio2, audio3 → out:audio | Mixes all connected audio inputs. |
Voice Ducking (voiceDucking) | voice:audio*, music:audio* → out:audio | Optionally delays voice, ducks music from the voice sidechain, mixes both tracks, and limits the result to -1 dB. |
Audio Enhance (audioEnhance) | audio:audio* → out:audio | Applies ffmpeg noise reduction with configurable strength and noise floor. |
Audio Trim (audioTrim) | audio:audio* → out:audio | Trims and optionally fades the audio. |
Speed Control (speedControl) | video:video* → out:video | Changes video and audio playback speed from 0.25× to 4×. |
Caption Burn-in (captionBurn) | video:video*, text:text → out:video | Burns connected or configured text into the video. |
Marketing nodes
Marketing nodes are visible to roles allowed to use the Marketing palette.
| Node | Ports | Behaviour |
|---|---|---|
Marketing Content (marketingContent) | user:text → out:text | Generates the selected content type: headline, hook, body, CTA, link description, website URL, image prompt, video script, or custom. It is local workflow logic and is never externally overridden. |
Marketing Variable (marketingVariable) | no input → out:any | Declares a typed text or brand-asset input that NovaStorm may supply. Its dropdown binds it to an actual Creative port or a generator prompt and its placeholder is development-only. |
Image Style (imageStyle) | subject:text → out:text | Reactively combines a subject with style and palette prompt presets. |
Brand Style (brandStyle) | no input → out:text | Emits a structured brand-style instruction from colours, tone and font settings. |
Text Motion (textMotion) | no input → out:text | Emits motion configuration for Video Creative text animation. |
Output Template (templateOutput) | selected Meta rendition inputs → no wire output | Marks the required final image or video renditions of a reusable Workflow Template. Every enabled profile must be connected. |
Backend-only legacy aliases
These types can still be executed by the server for compatibility with old or externally created graphs, but they have no current palette entry, node definition or renderer. Replace them with Text Generator or the corresponding current marketing node when editing a workflow.
| Type | Server ports | Current status |
|---|---|---|
Link Description (linkDescription) | user:text → out:text | Generates the Meta link description; Workflow Templates can replace it through bindingKey. |
Website URL (websiteUrl) | user:text → out:text | Produces the destination URL; normally supplied dynamically by meta-algorithm. |
Image Prompt (imagePrompt) | user:text → out:text | Produces an image-generation prompt or accepts the planned prompt dynamically. |
Video Script (videoScript) | user:text → out:text | Produces a video-generation script or accepts the planned script dynamically. |
Known current limitations
- Novastorm Avatar is verified only for English speech in 9:16. Other languages and 16:9 are experimental.
- Novastorm Avatar cannot record from the microphone yet: record a voice memo on a phone and upload or drop the file.
- The quick action on a video result currently creates an Image Iterator instead of a Video Iterator. Add a Video Iterator from the palette as a workaround.
- Four legacy LLM aliases are accepted by the backend but cannot be created or reliably edited in the current canvas.
- AI Fill and Object Eraser require a mask in graph preflight, but their current provider request performs prompt-based editing and does not send that mask.
- Outpaint exposes a prompt port, but its current backend builds the instruction only from direction and amount.
- Parameter controls constrain normal UI input, but image-processing parameters still need server-side dimension and resource caps for hostile imported graphs.