Skip to article
david sheffer.
← All notes

Local AI · Building in public

I gave an 8 GB GPU a film crew.

A local director, a video model, and an 8 GB GPU. Why my first action shot failed visually—and why a quieter second attempt taught me more about directing the model.

A quieter second attempt

My first test was a golden retriever jumping over a wooden hurdle. It produced a playable file, but the anatomy and motion were unconvincing. More detail would only have made those problems easier to see. I needed to change the brief.

For the second attempt, I chose a mountain lake at dawn: a fixed camera, a thin layer of drifting mist, and gentle ripples. I wrote this shot prompt by hand and ran it through the local Wan 2.2 5B workflow. The aim was a scene with fewer moving parts and a composition worth holding on screen.

The sample below is five seconds at 960 × 544 and 24 frames per second, generated on the same NVIDIA RTX 5050 with about 8 GB of video memory. It is silent AI-generated footage. This is a single curated shot, not a demonstration of a finished automated film director.

Local Wan 2.2 5B render: 5 seconds, 960 × 544, 24 fps. A hand-written prompt, a fixed-camera composition, and restrained motion, with a light brightness and color adjustment. AI-generated footage; no audio.
Open sample video

The useful unit was a shot

I built the pipeline around a small division of labor. A local language model turns the request into a shot plan. A video-generation workflow renders each shot. A conventional media tool normalizes the clips and assembles the final MP4. The language model is the director in this analogy, not the source of the pixels.

The default shot length is five seconds. A sixty-second request therefore becomes twelve shot jobs in the current planner. That is the decomposition implemented in the code, not a claim that I produced a coherent one-minute film. The original dog test exercised that pipeline; the lake sample above is a separate, manually directed render through its ComfyUI video backend.

The current loop renders shots sequentially. Independent jobs create a possible path to parallel workers later, but they do not make this version parallel. Keeping those two statements separate matters when a promising architecture starts sounding like a finished product.

StageResponsibility
Request → planLocal language model produces structured shot descriptions
Plan → validated jobsApplication checks shot count, assigns durations, and applies the requested style
Shot → generated footageComfyUI runs the selected video workflow; LTX in the original pipeline, Wan for the lake sample
Clips → final fileFFmpeg trims, normalizes, and joins the output
The prototype separates creative decisions from validation and media processing.

Eight gigabytes changed the design

The original workflow uses the LTX-Video 2B 0.9.8 distilled checkpoint, an FP8 text encoder, and tiled VAE decoding. The runtime log records low-VRAM operation and weight offloading. Those are specific choices in this experiment, not a claim that any video model will fit on an 8 GB card. The lake experiment uses the installed Wan 2.2 5B workflow instead, with 24 sampling steps and the same low-VRAM approach.

Model files on disk and the memory needed while executing a stage are different quantities. The question was which weights and intermediate data had to be resident together. Offloading and tiled decoding let the workflow manage that pressure, with time and transfer costs that do not disappear just because the job completes.

I kept planning and rendering as separate steps. The original LTX workflow also uses a fixed eight-step sampling schedule. Exposing a setting named steps in an application would be misleading if changing it did not change that schedule. The interface and the engine have to agree on what a control actually controls.

The director needed a contract

Asking a language model to think like a director was not enough. The application validates the number of shots and rejects blank descriptions. It sets the shot indices and durations itself, then appends the requested visual style. These are decisions the program can enforce instead of hoping the model follows them.

The directing instructions became concrete: describe visible action before, during, and after the movement; keep the subject in frame; place an obstacle across the direction of travel; do not add another action after the requested one. The goal was a usable shot description, not increasingly cinematic prose.

If the model-backed planner fails, the application logs the failure and falls back to a deterministic planner. That keeps the workflow usable, but a fallback plan is a different kind of result. A system should preserve that distinction instead of presenting every completed job as equivalent.

A valid MP4 can still tell the wrong story

There are several independent ways for this project to succeed. The shot plan can match its schema. The renderer can return frames. The file can have the requested duration and frame rate. The final video can play in a browser. Each is worth checking, and none proves that the requested action looks convincing.

Increasing resolution gives the viewer more detail to inspect. It does not prove that the dog clears the obstacle naturally, keeps a consistent body, or moves plausibly between frames. The dog test made that limitation visible. Choosing a quiet landscape for the next attempt reduced the action the model had to coordinate; it did not solve anatomy or physical reasoning.

I also kept a mock rendering mode that makes simple title-card clips. It is useful for exercising job flow and file assembly without paying for video inference. It cannot test visual quality. A test earns its value by making a specific claim, not by standing in for everything downstream.

The next problem is continuity

Short independent shots reduce the size of each generation job. They also create another problem: making the next shot belong to the same scene. A consistent character, location, and direction of motion cannot be assumed just because two clips concatenate successfully.

Reference images, image-to-video generation, quality scoring, and selective retries are next steps in the project notes. So are sound and a richer editing workflow. They are not capabilities demonstrated by the silent sample above.

What I have now is a working slice of a local creative pipeline: a plan, a renderer, an assembly step, and an artifact I can inspect. The exciting part is how much becomes possible when the job is divided carefully. The unfinished part is teaching the pipeline to recognize when its output deserves to survive the edit.

NEXT NOTEThe long way into engineering.