A Red-Line Guide Turns Aerial Images Into Virtual Drone Flights
ElevenLabs presents a workflow for generating drone-style video from a single AI-made aerial image, using a red line drawn over the image as a virtual camera route. The guide line tells the video model where to move through the scene, while the prompt specifies motion, pacing and visual action; the tutorial positions Gemini Omni Flash as the faster, cheaper option and Seedance 2.0 for longer, higher-resolution shots with native audio.

The red line turns an image into a camera route
A generated aerial image can serve as more than a visual reference: it can become a terrain map for a virtual camera. The controlling device is a red line drawn directly over that map. The video model is instructed to treat the line as navigation guidance, follow its route through the scene, and omit it from the final video.
The demonstrated palace image is downloaded, opened in Mac Preview, and marked with the sketch tool. The practical requirement is simple: draw one continuous red route through the locations the camera should traverse, then replace the original reference image in ElevenCreative with that marked version. The source shows the red line starting from the coastal approach, crossing the grounds and courtyard, and continuing into the palace hall; the resulting Seedance video is shown flying through the corresponding environment.
“The red line is a navigation guide only and must not appear in the final generated video.”
Without that mark, the video prompt asks the model to infer a natural route from the geography and layout of the scene, moving from harbor to palace interior. The line replaces that inference with a specific path. The text prompt still matters, but for a different job: it defines the flight’s behavior and dramatic sequencing—continuous movement, no cuts or teleportation, banking, changing altitude, foreground parallax, and acceleration through the action.
In the example, the route becomes a sequence of camera beats: a smooth approach over the coastline and courtyard; a low, fast pass by a central pillar as a spear strikes it; a bank toward a paneled door as another spear lands; a low advance along the wall; and a brief slowdown beside Telemachus at the final strike. The guide supplies where the camera goes. The prompt supplies how the journey should move, look, and build tension.
Build the world as an aerial map first
The workflow begins in ElevenCreative’s Image and Video area. A creator who already has a suitable aerial image can proceed directly to video. Otherwise, the first task is to generate an image that gives the camera a coherent place to travel through.
The central instruction is to make an aerial view of the scenery or landscape, not merely an attractive establishing shot. In the example, that means a photorealistic, 16:9 aerial map of the Palace of Odysseus in ancient Ithaca: a Mycenaean stone complex centered on a long great hall, with wooden pillars, a paneled door, thick surrounding walls, nearby coastline, and open terrain that will accommodate a later route.
The image prompt makes spatial legibility an explicit design constraint. It asks for the terrain and hall interior to remain “clean and open so a route can be drawn in afterward,” while prohibiting red lines, route markers, labels, borders, and a compass rose. The image needs enough detail to establish the world—a battle in progress, thrown spears, armed suitors, torchlight, dust, haze, and late-afternoon Mediterranean light—but not so much visual obstruction that a route becomes ambiguous.
The demonstrated setup uses GPT Image 2, at least 2K resolution, high quality, and a 16:9 aspect ratio. The aspect ratio is not presented as universally required; it must be compatible with the video model selected later. Generating several candidates creates room to choose a reference image whose layout is both visually strong and navigable.
Once a preferred image is selected, it is downloaded and marked up externally. On a Mac, the demonstrated method uses Preview’s sketch tool to draw the red line. The original, unmarked reference image is then removed from the video workflow and replaced with the marked version.
Choose the model for the output you need
ElevenCreative offers Gemini Omni Flash and Seedance 2.0 for the video-generation stage. The source frames the choice as a trade-off rather than a universal recommendation.
Gemini Omni Flash is described as cheaper and faster, making it the lower-cost, quicker option when its constraints are acceptable. Those constraints are a 720p ceiling, a 10-second maximum duration, and generations the tutorial describes as somewhat inconsistent. The source shows a Gemini result moving through the palace before ending on a warrior’s face.
Seedance 2.0 is the demonstrated choice for the marked palace route because the workflow calls for greater output flexibility. It can generate at 4K, offers more duration control, is not capped at 10 seconds, and can produce audio with the shot when sound is enabled. The source’s Seedance example moves through the palace combat while displaying the marked reference image and its red route alongside the output.
| Model | Advantages described | Constraints described |
|---|---|---|
| Gemini Omni Flash | Cheaper and faster | Limited to 720p and 10 seconds; described as less consistent |
| Seedance 2.0 | Up to 4K, more duration control, native audio | No specific drawback stated in the demonstration |
The generation settings need not use a model’s maximum capability. Although Seedance 2.0 can generate at 4K, the demonstration uses 1080p while testing, with the option to upscale afterward. It sets the shot to roughly eight seconds and produces two versions for comparison. That comparison remains useful even with an explicit route: the red line directs the flight, but multiple generations still provide a choice among the resulting interpretations of the scene.