Nova Creative / User guide

All nodes reference

This is the complete reference for the 77 node types available on the canvas. The server also accepts legacy aliases that are not available in the current editor; they are listed separately below. A port marked with * must have a value before the node can execute. Some generator requirements are conditional on the selected provider or model.

Cloud generator graph contract

For imageGenerator and videoGenerator, select provider/model ids and supported parameters from generation_catalog. Put provider-specific settings in data.params, using the catalog's exact parameter keys and value types. Aspect ratio is commonly data.params.aspect_ratio; do not invent a top-level data.aspectRatio field. Put an inline prompt in data.prompt, or connect a text node's out to the generator's prompt. Video duration in seconds may use data.duration; supported duration values come from the selected model.

For a reference imported with asset_import, create an importImage node with data.assetId, data.url and data.mime from the returned owned asset. Connect its out to the generator's image input. Video generators also accept endImage for a supported last-frame mode. Models may have additional conditional reference requirements; the server preflight validates them.

Create graph nodes with unique id, registered type, position: {x, y} and data. Edges need unique id, source, sourceHandle, target and targetHandle. For an existing project, use incremental operations with its latest updatedAt.

Value kinds

One concrete input handle normally accepts one connection. Text {{token}} handles and dynamically added compositor or concatenator handles are separate inputs, not multiple wires into the same handle.

Text

Iterators and lists

AI and text helpers

Imports, search and image services

Generators and creatives

Talking avatars

Lip-syncing a portrait to an existing speech track is not a separate node type. It is the ordinary Video Generator with provider set to nova-avatar, which runs an audio-driven avatar model on one of your own connected GPUs. Wire a portrait into image and a speech track into audio; prompt is optional and directs motion rather than describing the scene. The result is a clip whose length matches the audio. To have a person say a written script in a generated voice, use Novastorm Avatar instead.

Node data: provider: "nova-avatar", model, gpuNodeId (required), seed (below zero picks one), optional params. There is no templateId — these models ship their own inference runtime rather than a workflow.

Three things must be true before a run succeeds, and each reports a specific error: the GPU node must be connected, its weight bundle must be installed from the model catalog (aptavatar-bf16 or longcat-avatar-15-bf16), and that model's Python runtime must be provisioned on the node. Connecting several prompts fans out into one take per prompt against the same face, voice and seed.

Generation on your own GPU spends no credits. Renting the card is billed separately by the minute.

Novastorm Avatar (talking UGC creators from a script)

Novastorm Avatar (novastormAvatar) turns a written script into a selfie-style video of a person who says it, with a generated voice and the real background sound of the place. It runs the MiniMax H3 UGC workflows on one of your connected GPUs, so it spends no credits. The node is available to admin and designer accounts.

How it works. One MiniMax H3 take is at most 10 s (15 s when chosen, on a 48 GB+ card). Longer videos are split at sentence boundaries into segments of about equal words, each as long as its words need; when that would crowd one segment with more words than the longest segment holds at 3.2 words per second, the split that leaves the densest segment the fewest words is used instead (a segment of 35 words in 10 s dropped a clause). Segment 1 is text-to-video, or a reference clip when a photo or voice sample is connected or uploaded. Every later segment is a reference clip that starts on the previous segment's last five frames and the sound under them (anchored as its first frames; the last one is also <Picture 1>), so the motion, framing, pose, background and room sound carry on across the cut instead of being re-imagined (from one still frame H3 now and then swapped the background within the first quarter second); it keeps the photo as <Picture 2> and speaks with the voice sample (connected or uploaded) or segment 1's own voice (<Audio 1>). The person description is restated in every segment so the clothes stay the same. The segments are merged with each one's own sound: the repeated frames are dropped, and colour and loudness are matched at each seam. A video whose first words come after 0.5 s (the speaker sometimes waited up to 1 s) starts 0.25 s before them instead; pauses between segments stay. Finished segments are saved in the Library, so a failed run does not lose them. With a photo, segment 1 would start straight from the portrait, and H3 then sometimes cut to another shot or morphed a prop inside it; so the node first renders a short establishing shot (5 s, not kept) in which the person says the first few words of the script and goes quiet, and pins segment 1 to its last steady frame and the four before it (not their sound), the way every later segment is pinned to the one before it. If that shot fails, or segment 1 has fewer than 13 words (under 2.4 words per second even in a 5 s segment), segment 1 starts from the photo as before. A pinned segment has to talk from its first frame: given more time than its words need, H3 padded them with repeated or made-up words. So every segment is shortened to its words (about 3.2 per second, to the nearest length H3 can render: rounding up left up to 0.7 s of slack, which a take filled with a made-up phrase), a sparse Fixed-length script makes a shorter video, and every segment after the first gets at least 13 words (the split falls back to a clause or word boundary when a sentence boundary would leave fewer). People who move behind the speaker can pop in and out in any segment, so the café, office and gym presets ask for a few people seated, still and out of focus, and the café cup is an opaque one that stays on the table (a glass in the hand was bent out of shape). In a custom scene or scene details, don't ask for a busy or crowded place, for people passing behind the speaker, or for glassware in the speaker's hand. In a local test on an RTX 5090, a voice sample raised the similarity of H3's voice to the reference from 0.77 to 0.90. A real recording works best; voices cloned with text-to-speech sounded artificial.

The prompt follows the tested "natural UGC" recipe: a handheld front-camera selfie, the person moving with the phone, a voice described in words, and a loud soundscape of the place. Write the script as a short first-person story with one everyday detail, about 3 words per second. Do not use stage directions: text in brackets is spoken aloud.

Timing on an RTX 5090 at 768×1344: 5 s ≈ 4 min, 10 s ≈ 10 min, 20 s (2 segments) ≈ 20 min, 30 s (3 segments) ≈ 31 min, plus about 4 min for the establishing shot when a photo is connected. Each segment gets its own time budget, so long videos and batches aren't stopped at the usual 30-minute limit; a run still stops after 24 h — split very large batches. Common errors name their fix: the bundle is not installed (install minimax-h3-ugc-int8 from the model catalog), the GPU ran out of memory (set maxSegmentSeconds to 10 or free VRAM), the GPU is offline, or the GPU agent is too old for videos longer than one segment (update it to 0.3.15 or newer).

Example node for the workflow MCP:

{ "type": "novastormAvatar", "data": { "gpuNodeId": "<id from workflow_gpu_nodes>",
  "script": "I walk past this coffee place every single morning, and today I finally went in.",
  "durationMode": "auto", "duration": 10, "maxSegmentSeconds": 0, "scene": "street",
  "sceneDetails": "", "sounds": "", "person": "a woman in her late twenties with shoulder-length brown hair and a cream knit sweater",
  "speaker": "woman", "voiceStyle": "relaxed", "voiceCustom": "", "expression": "calm",
  "language": "en", "format": "9:16", "seed": -1 } }

Local GPU workflows (ComfyUI templates)

The Image Generator and Video Generator also run Nova's ComfyUI workflow templates on one of your connected GPUs. Node data: provider: "comfyui", templateId, gpuNodeId, optional params and, for video, duration in seconds. The MCP tool workflow_gpu_templates lists the templates. For each one it gives the catalog model to install, the minimum VRAM, the inputs it reads and its parameter defaults.

  • prompt and negative ports (or data.prompt) feed the template's text inputs. Image and audio inputs come from the port with the same key. Number inputs come from params.<key>; a seed below zero picks one.
  • The template's catalog model must be installed on the node. Otherwise the run fails before it starts and names the missing files.
  • MiniMax H3 templates generate speech together with the picture; write the line in the prompt as (S1) says: <d>[English] …</d>. On first-frame templates (…-i2v) an endImage port adds a last frame. Reference templates (…-r2v) take up to six images on image, image1 … image5 (<Picture 1> … <Picture 6>), and a speech track on audio becomes the voice reference <Audio 1>. The generated voice then matches that timbre, so one persona keeps its voice across clips.
  • builtin-minimax-h3-ugc-* templates (bundle minimax-h3-ugc-int8) are tuned for talking UGC creators: vertical 768×1344, 8 fast steps and a realism LoRA.

Datatypes and canvas-only nodes

Image editing and compositing

Video and audio processing

Marketing nodes

Marketing nodes are visible to roles allowed to use the Marketing palette.

Backend-only legacy aliases

These types can still be executed by the server for compatibility with old or externally created graphs, but they have no current palette entry, node definition or renderer. Replace them with Text Generator or the corresponding current marketing node when editing a workflow.

Known current limitations

  • Novastorm Avatar is verified only for English speech in 9:16. Other languages and 16:9 are experimental.
  • Novastorm Avatar cannot record from the microphone yet: record a voice memo on a phone and upload or drop the file.
  • The quick action on a video result currently creates an Image Iterator instead of a Video Iterator. Add a Video Iterator from the palette as a workaround.
  • Four legacy LLM aliases are accepted by the backend but cannot be created or reliably edited in the current canvas.
  • AI Fill and Object Eraser require a mask in graph preflight, but their current provider request performs prompt-based editing and does not send that mask.
  • Outpaint exposes a prompt port, but its current backend builds the instruction only from direction and amount.
  • Parameter controls constrain normal UI input, but image-processing parameters still need server-side dimension and resource caps for hostile imported graphs.