Upload & Setup
The first step of every project is uploading your audio and configuring how Filmgenic should generate your video. This page covers all the options available at project setup.
Audio Upload
Drag and drop or click to select your audio file. Filmgenic accepts the following formats:
| Format | Extension |
|---|---|
| MPEG Audio Layer 3 | .mp3 |
| Waveform Audio | .wav |
| MPEG-4 Audio | .m4a |
| Free Lossless Audio Codec | .flac |
Maximum file size is 50 MB. Upload progress is shown in real time with a progress bar.
Audio or script
A project is made either from an audio track (the original path: the track is analysed and cut into scene windows) or from a brief and script with no mandatory audio. Choose the mode when you create the project. A script-first project needs a pasted script and a target duration on its brief; its runs skip analysis, plan the scenes from the script, and render silent footage you can score later by adding audio.
Reference Images
Upload up to 14 reference images to guide visual consistency across your project. Reference images are used by the image generation model to maintain character, product, or brand identity across all scenes. How many of them a model reads differs: the GPT Image models - 2 and both 2.5 variants - accept all 14 (their edits endpoint takes up to 16), Nano Banana 2 uses up to 14, and Nano Banana Pro uses the first 5. Put your most important references first either way.
Each slot has its own control: Add reference image for the next empty slot, Replace reference N on a filled one, and Remove reference N to take one out. Removing a reference is saved immediately and survives a reload; the image file itself is kept because scenes already generated from it record it as a source.
For a recurring character, three clear shots of the same person (front, three-quarter and full body, same wardrobe) are enough to keep identity stable across every scene and every generation mode; a project with no references will drift between scenes.
Use reference images when you need:
- Recurring talent or character appearance consistency
- Product shots for advertising or branded content
- A specific visual world or brand identity
- Style guidance beyond what text prompts can convey
Image Model
The Image Model card chooses which AI model draws the still frame for every scene in this project. It is a per-project setting, not an account setting: two projects in the same account can use two different image models, and changing it here does not touch any other project.
Each button shows the model name and the exact model or deployment it will call. Which buttons appear depends on what this deployment is set up to run, so you may see fewer than the six below:
Image Model
Azure GPT Image 2
Uses up to 14 reference images.
Image Model
Azure GPT Image 2.5 - highest quality
The slower, higher-quality GPT Image 2.5 deployment. Uses up to 14 reference images.
Image Model
Azure GPT Image 2.5 - fast
The faster GPT Image 2.5 deployment, for quick passes over a whole film. Uses up to 14 reference images.
Image Model
Azure GPT Image 1.5
The earlier GPT Image model. Uses up to 14 reference images.
Image Model
Nano Banana Pro
Google's highest-quality image model. Uses the first 5 references.
Image Model
Nano Banana 2
Google's faster image model. Uses up to 14 reference images.
A project you have already started keeps the image model it is already set to. A model appearing in this list for the first time never changes an existing project's selection and never causes stills you already have to be redrawn — switching models is something you do deliberately, here or on a scene retry.
This setting controls normal scene image generation. A scene retry or a recovery action can still override the model for that one scene, as described on the Scene Generation page.
Video Model
Choose which AI video provider generates your scene clips. Each provider has different strengths, duration options, and visual characteristics. See the Video Providers page for a detailed comparison.
Video Model
Sora 2
Clip lengths: 4s, 8s, 12s. Strong at cinematic motion and human subjects.
Video Model
Veo 3.1
Clip lengths: 4s, 6s, 8s. Excellent at natural scenes and landscapes.
Video Model
Gemini Omni 1.1 Flash
Clip lengths: any whole second from 3s to 10s. Generates audio with the picture.
Video Model
Runway Gen4.5
Clip lengths: 5s, 10s. Great for stylized and abstract visuals.
Video Model
Luma Ray-2
Clip lengths: 5s, 10s. Strong at dynamic camera movement.
Video Model
Grok Imagine
Text-to-video with distinctive visual style.
Generation Mode
The generation mode shapes how the AI plans your video. See Generation Modes for a full breakdown.
Visual Style
Choose from 8 visual styles that influence the art direction across all scenes. See Visual Styles for examples and descriptions.
Planner Model
The planner model is the AI used for creative reasoning and planning. See Planner Models for selection guidance.
Script Input
Optionally paste lyrics or a script. When provided, the platform runs speech alignment to map words to timestamps, enabling word-aligned scene planning in performance mode. This is especially useful when you want scene cuts and performance direction to follow specific lyrics or dialogue.
