AI Agents Assemble Documentary Motion Graphics From a Structured Brief
ElevenLabs presents a workflow for producing scrapbook-style documentary videos in which a tested visual prompt, persistent background and repeated reference assets establish continuity before scene generation begins. Claude converts a script and creative direction into a production brief, while ElevenCreative Flows builds the canvas, generates storyboard options and animates selected scenes. The company’s argument is that agent-assisted production can reduce manual setup without removing the creator’s role in supplying assets, choosing variants and revising the edit.

The style prompt is the production system
The workflow treats the visual style prompt as a reusable production specification: a common set of rules for every background, scene, asset, and animation. It should be established before generating individual images, then repeated throughout the build so separately generated pieces do not drift into different visual languages.
The demonstrated prompt defines an editorial scrapbook collage: aged archival paper, muted map texture, print grain, halftone dots, black-and-white photo cutouts, torn-paper scraps, masking tape, and sharp headline typography. It also sets constraints rather than just mood: a limited tan, black, gray, red, and mustard palette; red reserved for strokes, underlines, and arrows; a matte print finish; and no glossy 3D, lens flares, watermarks, or logos.
The palette itself is not the point. The operational distinction is between style rules, which recur in every generation, and scene prompts, which specify what each individual image needs to depict. Settling the former early means later prompts can focus on the story beat without re-solving the visual identity.
A low-cost test with Nano Banana 2 Lite lets the creator inspect what the style block actually produces before committing to the rest of the project. If the paper texture, cutout treatment, or color balance does not work, revise the block and rerun it before building scenes around it.
The same principle applies to the persistent background. Rather than asking every scene to invent its own paper stage, the workflow generates one background and reuses it as a reference in every clip. In the example, that background and the scene images are generated with GPT Image 2 at 4K quality, alongside personal references: a childhood photo, a current photo, and a camcorder resembling the original. Repeated assets and repeated style instructions are the mechanism for continuity.
The voiceover defines what each visual beat must support
The images are meant to accompany a spoken story, not simply decorate a mood. The script and voiceover therefore come before the Flows agent is asked to build the production canvas.
The example is a compact personal narrative: a child borrows a cousin’s camcorder in the early 2000s and films sketches in the back garden; nobody watches, but he keeps filming; 20 years later, video is his career. Those short beats give the generation system distinct moments to illustrate.
Eleven v3 generates the narration. Audio tags shape delivery: storytelling establishes the overall tone, while a pause at the end leaves room for trimming. Tags can also control individual lines. Adding whisper to “In the early 2000s,” produces a quieter, lower delivery for that phrase.
A narrator-style voice is presented as a good fit for this kind of storytelling. The example uses Noah, but the workflow is to preview alternatives and regenerate until the voice and pacing fit the intended film.
The practical implication for editing is to generate clips longer than the final cut requires. Extra footage permits a scene to begin earlier or later against the narration without forcing a full regeneration.
Claude structures the plan; the agent builds the canvas
Building manually requires creating nodes, attaching assets, writing prompts, generating variations, and arranging the sequence one piece at a time. The proposed shortcut does not remove those decisions. It divides them among three roles: Claude turns creative direction into a structured brief; the Flows agent creates the canvas, nodes, and generation variants; the creator provides the source assets, selects outputs, and corrects weak boards.
The supplied Claude prompt casts Claude as a video-production planner. Given a short documentary script and a list of reference images, it returns a ready-to-paste build brief for the ElevenLabs Flows agent, covering voiceover, style frame, collage boards, and animated clips.
For this planning input, the script is pasted without its audio tags. Reference assets can be described loosely—such as a photo of the creator as a child, as an adult, and a camcorder—rather than uploaded to Claude. The prompt directs the eventual Flows agent to request the actual files.
Two parts of the brief carry the continuity rules forward:
- The style block states the colors, asset treatment, and animation vocabulary that should govern the build.
- The character lock is intended to keep people in reference images visually consistent. The source says Gemini Omni Flash can otherwise change a character’s appearance during animation; movement is wanted, but a different-looking subject is not.
Once the resulting brief is pasted into a new flow, the agent asks for reference assets, places them on the canvas, and presents generation prompts. It can pause for approval or run automatically if the user changes permissions.
The agent produces alternatives for the persistent background and storyboards. The creator chooses a preferred result in chat, or edits and reruns a prompt when a scene is too similar to another or includes an unwanted asset. Automation constructs the sequence; it does not decide which visual interpretation best serves the story.
Animation is treated as a composition, not a moving still
The source describes Gemini Omni Flash as reading the complete storyboard frame as context, separating its elements, and animating them into place. That differs from the behavior attributed to many other video models, which are said to use the supplied image as a starting frame and animate forward from it.
In the first completed scene, the child cutout enters from the left, the camcorder moves in, and a red arrow and labels appear around them. The desired effect is not motion applied uniformly to a flat image. It is a collage whose individual objects arrive and move as separate elements.
Storyboard iteration remains available while scenes render. When a second beat looks too similar to the first, the workflow changes its prompt and considers removing the camcorder so the frame can focus on the child. That board is regenerated while the first scene continues rendering.
Studio turns generated clips into a timed film
Completed animations are saved to an asset folder, then imported into an ElevenLabs Studio video project with the generated voiceover. The clips are ordered to match the narrative beats and aligned against the narration.
Music is generated separately in Studio. The example requests instrumental, marimba-style documentary music and sets the generation length to 38 seconds for a roughly 30-second film, leaving room to trim. Three music variations are returned; one is added to the timeline, shortened, and lowered in volume so it functions as background.
The result is presented as a usable first pass rather than a finished edit. Timing can be adjusted, weaker scenes can be regenerated, more speech can be created, and Studio can also generate video directly using references.
