Orply.

Character Reference Sheets Make AI Film Workflows Reusable

Alec WilcockElevenLabsSunday, August 9, 20269 min read

ElevenLabs’ Alec Wilcock argues that AI short films should be built around continuity rather than motion: establish the character, visual style and composition in still images before spending credits on video. His *Little Wanderer* workflow uses a reusable character reference sheet, narration-led story beats, and start and end frames to constrain each shot, then assembles and trims the results in Studio. The node-based Flows canvas, he says, makes the project reusable by allowing an upstream character change to regenerate the same film structure with a new protagonist.

The first constraint is not motion but continuity

The costly failure mode in AI filmmaking is generating video before deciding what the film should look like. A short, simple prompt—“hedgehog,” for example—can produce plausible images, but different models interpret it differently and impose different visual styles. The result may be attractive in isolation while being unusable as part of a sequence.

The workflow demonstrated in Little Wanderer begins with world-building in still images: the character, environment, objects, and visual language. Its protagonist is a tiny hedgehog in a needle-felted, stop-motion-inspired world of stitched autumn leaves, quilted hills, felt flowers, and embroidered stars. That choice creates a useful test case for the central production problem: getting the same character and material style to persist from one shot to the next.

Alec Wilcock advises starting with image generation before committing to video. Images are faster and cheaper to generate, while video takes more time and credits. A still also establishes the visual facts a video model will inherit: what the character looks like, where it is positioned, how the scene is lit, and which details need to remain stable.

Generating images is much cheaper and faster than actually generating videos. And the biggest mistake is actually trying to generate the video first because you'll end up burning through your credits way too quickly without even knowing what that generation is going to look like.
Alec Wilcock · Source

Nano Banana 2 Lite is presented as a low-cost exploration model: fast, inexpensive, and limited to 1K resolution. GPT Image 2 and Seedream 5 Pro produce different results from the same prompt. That variation comes from three sources shown in the workflow: each model interprets prompts differently, each has its own output style, and vague prompts leave more decisions to the model.

Specificity narrows that variation. The felt-hedgehog prompt is structured around four elements: the style of the scene; the subject and its physical details; the objects and setting; and exclusions such as “No text.” Framing and lighting details—macro close-up, shallow depth of field, warm morning light, or a cinematic wide shot—further specify the shot rather than merely naming its subject.

The practical sequence is to test models and prompts while exploration is inexpensive, then choose the model whose output best matches the intended film. For this felt-animation style, Wilcock prefers GPT Image 2 over Nano Banana because its material treatment better fits the handmade look.

A reference sheet turns a generated character into a reusable asset

A favorable character image is not enough to establish a character. The workflow turns the selected hedgehog into a high-resolution character reference sheet: a 16:9 image showing front, side, back, top, bottom, and three-quarter views, with close-ups, colors, scale, and construction details.

The displayed sheet identifies the hedgehog’s face-and-body color, spikes, nose, eyes, and inner ears. It also specifies the material cues that make the animal recognizable: needle-felted wool, soft felt spikes, plastic safety eyes and nose, an embroidered mouth, round ears, and stitched paw pads. It is generated at 4K and high quality because it will be reused as an input throughout the project.

That sheet is the mechanism for character consistency. Each storyboard scene receives the same reference, giving the image model a shared description of the hedgehog across different poses, camera angles, and environments. Scene prompts can still determine composition, action, and setting; the reference sheet keeps them from having to reinvent the protagonist every time.

If the character needs revision, the reference sheet can be routed into an image-editing node and changed upstream—blue spikes, for example—before subsequent scenes are generated. ElevenLabs presents Flows, its node-based canvas, as the system for retaining these relationships: a source image feeds a reference sheet; the sheet feeds storyboard images; those storyboard frames feed video nodes.

The demonstrated node links preserve reusable inputs rather than merely arranging a canvas. A creator can duplicate a connected image node for the next storyboard scene, replace its text prompt while retaining the character-reference connection and existing settings, then run that node. Generated assets can be saved in a shared folder and used across Flows and Studio without downloading and re-uploading them.

Narration establishes the film’s beats before the storyboard does

Generate narration before finishing the storyboard, because the voiceover determines how many visual beats the film needs and how long the final edit must run.

For Little Wanderer, the narration is brief:

“In a quiet corner of the world, a tiny adventurer takes her first step. One step. Then another. Until the whole sky opens up. Little Wanderer.”

That line becomes a five-shot sequence: the hedgehog peeks out from autumn leaves; walks through a meadow; meets a glowing firefly near a flower; watches it rise into the dusk sky; and arrives at a felt title card.

ShotStoryboard beatVisual purpose
1Hedgehog peeks from stitched autumn leavesIntroduces the character and felt material style
2Hedgehog toddles through a quilted meadowEstablishes movement and a wider environment
3Hedgehog meets a glowing firefly by a flowerIntroduces the story’s point of curiosity
4Hedgehog watches the firefly rise into duskCarries the action toward the night sky
5Embroidered stars and title cardResolves the short film as “Little Wanderer”
The five visual beats generated around the narration in the demonstrated storyboard

The voice is selected from a library according to its role. The example uses a narrator-style voice because it describes the hedgehog’s action rather than portraying the hedgehog as a speaker. If the result is unsuitable, the text, selected voice, or generation can be changed.

Eleven v3 adds more granular control through audio tags: bracketed instructions placed around specific lines. Tags shown include [Whisper], [Calm], [Laughing], [Shouting], [Surprised], and [Dreamy]. Applying [Whisper] to the opening changes the delivery accordingly. The use of tags is not simply to select a voice but to direct how that voice performs individual moments.

Once the voiceover is satisfactory, it is saved as an audio asset and brought into Flows. A text-to-speech node is available on the canvas, but the dedicated text-to-speech interface is presented as more useful for longer narration.

Solve composition in stills, then ask video to create change

The storyboard pairs a scene prompt with the recurring character reference sheet. In the example, each image node uses GPT Image 2 at 16:9, 4K, and high quality; what changes is the shot prompt.

This is where composition problems should be addressed. If the hedgehog needs to sit on the other side of the frame, a flower should change color, or an additional insect should appear, resolve that in the still image before generating motion. There are two routes: alter the scene prompt and regenerate within the image node, or feed the existing frame into an image-editing node for a targeted change such as “Add 2 bees.”

For precise image edits, GPT Image 2 and Nano Banana 2 are preferred over Nano Banana 2 Lite, whose 1K resolution limit makes it less appropriate when the edited frame will become a video reference.

Multiple image references require explicit tagging in the prompt. When a portrait of the presenter is introduced into a hedgehog scene, one reference is tagged as the source for a felt version of the person, while another is tagged as the scene that should remain intact. The resulting edit places a felt presenter at the right of the hedgehog.

That second still has an additional purpose: it can become the end frame for a video shot. The original hedgehog-only image provides the start frame; the edited image provides the destination. Rather than asking the video model to invent an ending, the creator gives it two defined visual states to bridge.

Once the stills are approved, video prompts should describe motion rather than repeat the whole composition already present in the frame. The first shot asks the hedgehog to blink, sniff the air, and take one gentle step forward while the macro camera remains mostly still. Other prompts specify a slow camera drift, a firefly’s pulsing glow, a head tilt, stars twinkling, or title lettering appearing thread by thread.

Small, controlled actions are a consistent feature of the examples. More detailed motion prompts can provide more control, but the image itself supplies much of the scene context.

Generate enough footage to cut, not footage that must be extended

Different video models are suited to different outputs. Kling models are presented as useful for animation-style work. Seedance 2.0 is selected for the example and characterized as realistic, but expensive and slow; Gemini Omni Flash is described as capable but sometimes difficult to make consistent.

The selected Seedance 2.0 settings are 1080p, 16:9, and seven seconds. The workflow notes that Seedance 2.0 can generate between four and 15 seconds, and Alec Wilcock suggests that a shot in this story may need about five seconds before choosing seven to leave room for the edit.

4–15 seconds
Seedance 2.0 duration range shown in the workflow

A lower-cost path is to test video at a lower resolution and upscale only the take worth keeping. The source shows a 4K-image-to-1080p-video workflow, followed by a Topaz upscale node using its “precise” mode. The example generates directly at 1080p, but the operational choice is to reserve upscaling for footage whose motion has already been accepted.

Extra material can be removed during the edit; asking a model to extend a shot later can undermine consistency.

Extending a scene is never as good as trimming down the original one.
Alec Wilcock

Start frames provide a defined opening state, and end frames can provide a defined destination. In the presenter-entry shot, the prompt adds “man walks in from the right” to explain the difference between the start and end images. Audio is disabled for those video generations because voiceover, music, and effects are handled separately.

Flows can preview the sequence, but Studio is where timing is resolved

Background music can be generated while video clips render. The example uses Music v2 with lyrics turned off and a detailed prompt for a warm, whimsical storybook score: music box, plucked strings, wooden flute, pizzicato strings, soft percussion, cello, harp, tambourine, and hand chimes, with no big drums, cinematic hits, or tension.

Flows can assemble a rapid preview through a composition node. Video clips are connected in sequence, then the saved voiceover and generated music are added as audio tracks. That creates a way to hear the narration against the clips and assess the order of the story before editing.

The limitation is that the composition node cannot trim. The clips therefore move to Studio, the timeline-based interface. The voiceover is placed on the timeline before the clips because it is the editorial timing reference; the visual material is then moved and trimmed to its cadence.

The music track is shortened to the voiceover’s duration, reduced to 15% volume, and given a roughly three-second fade-out toward the title screen. The generated video clips are cut down to the moments the narration needs.

Studio remains connected to the same generative tools. Music can be regenerated in the editor, speech can be changed, and previously generated image and video assets remain available. Sound effects are the final layer: SFX v2 can generate described sounds such as rustling leaves, footsteps, or rain, and those clips can be added directly to the timeline. A sound-effects library is also available for browsing existing material.

The reusable film is the graph rather than the export

Alec Wilcock presents the node graph as the reusable asset. The character reference sheet sits upstream of storyboard and video nodes. Regenerate that sheet, run the flow from that point, and the same prompts, structure, and story beats can be generated around a different character.

The source demonstrates the change by replacing the felt hedgehog with a small felt astronaut. In the resulting montage, the astronaut appears in the same autumn leaves, meadow, flower-and-firefly scene, and night-sky shot before the same Little Wanderer title card. The character changes across every scene while the visual progression remains recognizably the same.

That is different from replacing one clip in a conventional timeline. The reusable object is the set of links among the character reference, scene prompts, image nodes, video nodes, and audio assets. An upstream character change can be carried through the downstream scenes, while individual nodes remain available for local corrections.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free