Eight qualified image models. One prompt-to-download workspace.

PRACTICAL GUIDE · PROMPTTOVISUAL

Kandinsky 6.0 Video: Lite vs Pro, Audio, Weights and API Pricing

Kandinsky 6.0 offers Lite and Pro models for five-second video with synchronized audio. Choose between self-hosted weights and a hosted API, then check the exact settings and cost. This independent guide is not a PromptToVisual integration or model test.

Open official Kandinsky repository →

What was released, and when?

The official repository records its open-source release on October 6, 2026; the paper was submitted October 4. These are separate dates. The family includes Lite (3B parameters) and Pro (29B). The publisher describes text-to-audio-video and image-to-audio-video, with five-second clips, 44 kHz sound and lip synchronization. These are release claims, not our output measurements.

Source for this section →

Lite or Pro: choose a trial around your task

For an inexpensive first draft, inspect the Lite offer first. For a shot where you can justify a higher experiment budget, evaluate Pro against a written acceptance checklist. This is a cost-based starting point, not a quality ranking: check subject continuity, motion, sound timing and unwanted changes in the actual output. More parameters do not guarantee that your particular shot will work.

Audio, five seconds and 1080p mean different things

The repository describes synchronized audio and video, plus a separate super-resolution stage reaching 1920×1080. Do not interpret Full HD as the native resolution of every base generation. Five seconds describes clip duration, not request latency. A lip-sync claim is not a guarantee for every voice, language or face.

Source for this section →

Self-hosting: select the checkpoint before setup

The repository requires an NVIDIA GPU and Python 3.13 or 3.14, with uv and just in its quick start. Its catalog distinguishes ordinary, pretrained and distilled checkpoints. Review the device configuration and storage requirements before downloading. Self-hosting means you operate the environment and pay its compute costs; a weight download is not a hosted generation service.

Source for this section →

Weights and the scope of MIT

The linked Pro-distill Diffusers model card declares MIT and specifies 10 steps with PiFlow and guidance 1.0. It explicitly warns against substituting the other checkpoints’ 50-step, guidance 5.0 settings. The repository code has an MIT license, and the paper states MIT release of code, checkpoints and Diffusers integration. This does not establish that every dependency has the same license. Check notices and the exact artifact you use.

Source for this section →

fal Lite API pricing: text or first image

Checked October 6, 2026: Lite text-to-video is USD 0.16 per unit; Lite image-to-video is USD 0.17. Each listed unit is one five-second, 480p video at 10 inference steps. These are external API prices, not PromptToVisual credits, and not a quote for arbitrary resolutions or step counts.

Source for this section →

fal Pro API pricing: different settings

The Pro text-to-video page lists USD 1.35 per unit; image-to-video lists USD 1.40. Each unit is one five-second, 480p video at 50 inference steps, checked October 6, 2026. Comparing these prices with Lite at 10 steps is a comparison of listed offers, not a controlled quality or speed benchmark.

Source for this section →

Budget super-resolution separately

The separate fal VSR listing quotes USD 0.00036 per output megapixel × frame at 5 inference steps. Confirm the actual frame count, dimensions and current billing before using it. Do not assume upscaling is included in a base-generation unit. Budget for every requested attempt, including outputs you choose not to keep.

Source for this section →

Before opening the official API

Choose text-to-video if starting from a scene description, or image-to-video if starting from a licensed image. Prepare the intended action, camera movement and sound, and identify what must remain unchanged. Read the chosen endpoint’s schema, account access and total quote. Keep the first experiment narrow enough that its result is easy to assess. This guide has not executed the setup or any API request.

Original prompt idea — not run

A ceramic cup rests on a wooden table. The camera moves slowly closer while steam rises. Soft room ambience; no music, speech, text or new objects. This original example is untested. Review the returned video for cup geometry, flicker and whether the requested audio is present; the instruction itself does not guarantee preservation.

Where to go next

Use the official repository, weight card and fal endpoints below for Kandinsky itself. It is not available in our model selector. PromptToVisual’s existing Seedance 2.5 workflow is a different alternative: one first image, four seconds and 40 paid credits under the current configuration. It is not equivalent access to Kandinsky.

Sources and next steps

Sources checked October 6, 2026. Independent documentation review; no integration or model generation test.

Official repository and setup →

Repository MIT license →

Paper and submission date →

Pro-distill Diffusers weights and license →

kandinsky6-lite/text-to-video →

kandinsky6-lite/image-to-video →

kandinsky6-pro/text-to-video →

kandinsky6-pro/image-to-video →

kandinsky6-vsr →

Utopai X access guide · Keyframe control guide · FLUX 3 Video guide · Seedance limits · PTV single-first-frame alternative · All guides

Checking image studio availability…