Does generative AI change your footage? Full-frame vs plate-preserving edits
Full-frame generative video editors re-render every pixel, not just your edit. What current tools document, how to test for plate changes, and when to use them.
Usually, yes. Full-frame generative video editors such as Runway Aleph 2, Kling O3 edit and Google’s Gemini Omni Flash take your clip and a prompt and generate a new video. Every pixel passes through the model, including the ones you didn’t ask to change, so grain, fine texture, faces and small text can shift across the whole frame. Plate-preserving work does the opposite: it changes only a masked region and composites that region back over the original, so everything else stays exactly as shot.
These tools are impressive, and for some jobs they’re the right choice. This post covers how diffusion video editing works at a high level, what current tools document about their limits as of September 2026, why the difference matters on client work, how to test any tool yourself in a compositor, and when generative editing is a good fit.
How a generative video editor processes your clip
Most published video generation models, such as Wan, are latent diffusion models. Commercial vendors don’t all publish their architectures, but the general approach is the same. Rather than working on your pixels directly, they:
- Encode the clip into a compact latent representation using a learned autoencoder, often called the VAE.
- Generate a new latent video, guided by your prompt, reference images and the encoded source.
- Decode that latent back into pixels, then encode it to a delivery file, usually an MP4.
The key paper behind this design, Rombach et al., describes running diffusion “in the latent space of powerful pretrained autoencoders” to make it computationally practical, and frames the autoencoder as a balance between reducing complexity and preserving detail. How much compression that means: the Wan video model paper says its VAE compresses a clip to one-quarter of the frames and one-eighth of the width and height, with 16 channels. By our arithmetic, each latent position stands in for an 8×8-pixel block across about four frames: 768 RGB values summarized as 16 numbers.
A decoder working from a summary that compact can’t give back your exact pixels. It gives back plausible ones.
Why grain, faces and text drift
The detail most likely to change is the detail the latent can’t hold exactly:
- Film grain and sensor noise are random per frame. The decoder produces grain-like texture, but not your grain, and its character can differ from the shots on either side of the cut.
- Faces carry identity in small details: skin texture, eyes, the corners of the mouth. Regenerating them can subtly change the person on screen.
- Small text and logos are high-contrast fine detail. Letters can soften, warp or change.
- Motion over time. Long clips are generated in windows and must stay consistent from frame to frame. When they don’t, detail can shimmer or slide.
This follows from how the pipeline works; it isn’t a flaw of one vendor. Every tool that decodes a new frame from a latent will do some version of it. The questions are how much, and whether it matters for your shot.
What current tools document, as of September 2026
These specs come from each vendor’s own documentation, accessed on September 24, 2026. They change often.
| Tool | How you direct the edit | Input limits | Output |
|---|---|---|---|
| Runway Aleph 2.0 (API and Edit Studio) | Text prompt plus up to 5 timed keyframe images; optional edit window. No mask input | API: 2–30 s, 30 fps or lower. Edit Studio: 480p–1080p, 24–30 fps (higher rates are downsampled), 10 cuts or fewer | Keeps the input resolution up to 1080p. MP4, ProRes or PNG sequence; ProRes 4444’s alpha is “present but fully opaque” |
| Kling O3 edit (via fal) | Text and reference images | .mp4/.mov, 3–15 s, 720–3840 px, 200 MB max | MP4 |
| Gemini Omni Flash (Gemini API) | Text and reference images | Uploaded video for editing: 10 s or less | 720p by default; 1080p and 4K are “upscaled”. SynthID watermark on all outputs |
| Luma Ray 3.2 video edit | Prompt with structural conditioning | Source 18 s or shorter, 200 MB max | 360p–1080p MP4; optional HDR and EXR export |
| Wan VACE 14B inpainting (via fal) | Mask video plus prompt | 17–241 frames, 5–30 fps | Up to 720p MP4 |
| LTX-2.3 inpaint (via fal) | Mask video: “White regions are regenerated; black regions are preserved” | 9–481 frames, 1–60 fps | MP4 |
Sources: Runway’s API docs, API reference and Edit Studio guide; fal’s Kling O3, Wan VACE and LTX-2.3 pages; Google’s Gemini Omni docs; Luma’s video editing guide.
Two points stand out. First, the full-frame editors have no mask input: you describe the edit, and the model decides what else moves. Second, the “masked” models are more controlled, but they still return a newly decoded, re-encoded MP4 of the whole frame. The pixels outside the mask come back close to the source, not identical, and in Wan VACE’s case at up to 720p. Even a masked model needs a composite-back step to be truly plate-preserving.
Google also notes that editing uploaded videos with Gemini Omni isn’t available in the EEA, Switzerland or the UK, and that editing images of certain recognizable people isn’t supported.
Why this matters on client work
For a social post, a regenerated frame is often fine. On a commercial, a music video or a feature, the plate is the reference that everything downstream relies on.
- Plate integrity and conform. Editorial, color and finishing expect a VFX shot to drop back into the cut at the same resolution, frame rate and frame count, with nothing changed outside the fix. A 1080p re-render of a UHD plate fails that before anyone looks at the pixels.
- Grain and continuity. A shot whose grain and texture were regenerated can stand out against its neighbors after the grade, even when no one can say why.
- Color and bit depth. Most outputs are 8-bit Rec.709 MP4. Runway offers a 10-bit Rec.709 option and Luma offers HDR and EXR, but a log or high-bit-depth plate still has to be converted for the model and back.
- People and approvals. If an actor’s face was regenerated, the performance on screen is technically new pixels. That can matter for talent approvals, and for brand approvals when a product or logo passed through the model.
- Provenance. Google says every Gemini Omni output carries an invisible SynthID watermark that can be detected programmatically. Know what’s embedded in your deliverable before a client asks.
If footage security is part of the question, our security checklist for client footage and AI tools covers the upload side.
How to test any tool for plate changes
You don’t have to trust anyone’s description, ours included. A difference matte shows exactly what changed. In Nuke, a Merge set to difference computes abs(A-B), which Foundry’s documentation describes as useful for comparing two very similar images. After Effects has a Difference blending mode that does the same job.
- Pick a test shot with visible grain, a face, some small text and a large area you won’t touch.
- Ask for a small, local edit in one corner, such as removing a sign or changing a cup’s color.
- Conform the result to the source: same resolution, frame rate and first frame. Confirm the frame counts match before you compare anything.
- Stack source and result in your compositor and set the top layer to difference.
- Gain up the difference with an exposure or levels node, so small changes become visible.
- Look outside the edit. Unchanged pixels stay black. A lossy re-encode shows up as faint, even noise. Regeneration shows up as structure: edges, facial features and letters appearing in the difference.
- Scrub the whole range. Check early, middle and late frames, and any chunk boundaries, for changes that come and go.
- Check the edit itself at 100% against the source, then in context after a rough grade.
Keep the test shot and the settings you used. When a tool updates, run it again.
When generative editing is the right call
Full-frame generation does things masking can’t: restyling, relighting a whole scene, changing weather, rebuilding a set. That’s valuable when the whole frame is allowed to change, or when you only need part of the output:
- Previz and pitchvis. Showing a director or client an idea quickly matters more than plate integrity.
- Temp comps for editorial. A stand-in effect that lets the cut move forward while the real shot is built.
- Social and short-form content where the regenerated look is acceptable and resolution caps don’t matter.
- Element generation. Generate a sky, a background plate or a set extension, then composite only that element through a matte over the real plate. The subject and the rest of the frame stay original. We cover this route in background replacement without a green screen.
The pattern that works for finals is the same in every case: use generation as a source of pixels, and decide in your compositor which of them reach the shot.
What “plate-preserving” means
A plate-preserving edit changes pixels only inside a defined edit region. Outside it, the output matches the source plate: identical decoded values, with the only differences coming from the delivery codec if you choose a lossy one. In practice that workflow looks like this:
- Build a matte for the region to change (see what rotoscoping is).
- Crop a tile around it at the model’s working resolution.
- Run the masked model on the tile.
- Composite only the feathered mask region back over the original, full-resolution plate.
- Regrain and color-match the patch.
That’s the same logic as a traditional paint fix; the model supplies the fill. The clean plate guide covers the regrain and QC side in more detail.
Where nolanlabs stands
Our principle is “only what you ask for changes”. For region jobs such as cleanup, object editing, environment, crowd and performance work, nolanlabs composites every result back onto the original plate, so everything outside the edit region stays as shot. Whole-frame jobs such as relighting change light across the frame; nolanlabs says so before the run and keeps the performance, framing and timing. It also states engine limits before a run, such as inpainting at up to 720p inside the edit region. The reasoning behind that is in the manifesto; the six use cases are on the home page.
Whichever tool you use, the test is the same. Put the result over the source in a compositor, set it to difference, and look at what changed outside your edit. If anything did, it’s up to you whether that’s acceptable for the shot.
Questions
- Does Runway Aleph change parts of the video I didn't ask to edit?
- It can. As of September 2026, Runway's API takes a video, a text prompt and up to five keyframe images for Aleph 2.0, with no mask input, and returns a newly generated video at up to 1080p. Pixels outside your edit are regenerated too, so check them with a difference matte, and composite the result back through a matte if the rest of the plate has to stay as shot.
- Is a masked AI inpainting model plate-preserving on its own?
- Not exactly. Masked models such as Wan VACE or LTX-2.3 inpaint on fal are told which region to change, but they return a newly encoded MP4 of the whole frame, and Wan VACE works at up to 720p. Pixels outside the mask come back close to the source, not identical. To keep the plate untouched, composite only the masked region of the result back over your original.
- Can I use a generative video edit in a final client deliverable?
- Sometimes, if you treat it as an element rather than a replacement plate: composite only the part you need back over the original through a matte, match grain and color, check it with a difference matte, and make sure the client has approved generated content in the shot. For previz, pitchvis and temp comps, full-frame output is often fine as it is.
Sources
- Runway API documentation (llms-full.txt): Aleph 2.0 input and output limits, Runway · accessed 2026-09-24
- Runway API reference: POST /v1/video_to_video (aleph2), Runway · accessed 2026-09-24
- Creating with Edit Studio, Runway Help Center · accessed 2026-09-24
- Kling O3 Edit Video [Pro], fal · accessed 2026-09-24
- Generate and edit videos with Gemini Omni Flash, Google AI for Developers · accessed 2026-09-24
- Video editing (Ray 3.2), Luma Agents documentation · accessed 2026-09-24
- Wan VACE 14B inpainting, fal · accessed 2026-09-24
- LTX-2.3 inpaint, fal · accessed 2026-09-24
- High-Resolution Image Synthesis with Latent Diffusion Models, by Rombach et al. (arXiv:2112.10752) · accessed 2026-09-24
- Wan: Open and Advanced Large-Scale Video Generative Models (arXiv:2503.20314) · accessed 2026-09-24
- Merge operations (Nuke 17.1), Foundry Learn · accessed 2026-09-24
- Blending modes and layer styles, Adobe After Effects Help · accessed 2026-09-24
Written by the nolanlabs team. nolanlabs does AI shot work for post-production; product names mentioned belong to their owners. Tool details were accurate on the date shown above. Check the vendor’s documentation before relying on them.
