AI Video

Self-Hosted AI Video Production: An Experiment — Generating Talking-Head Videos Entirely on Your Own Hardware

Self-Hosted AI Video Production: An Experiment — Generating Talking-Head Videos Entirely on Your Own Hardware

Self-Hosted AI Video Production: An Experiment — Generating Talking-Head Videos Entirely on Your Own Hardware

An experiment report: we set out to find how far local hardware gets you when generating talking-head videos entirely on your own — from ComfyUI and MiniMax-H3 to lip-syncing a cloned voice. The result: surprisingly close, but not production-ready. All numbers, settings and dead ends.

What this article is about

AI-generated videos with speaking people are everywhere: LinkedIn clips, product explainers, social media formats. The common tools run in the cloud — you upload photos, pick a voice, and someone else's server produces the video. That is convenient, but it has three drawbacks: clips cost several euros each, customer recordings leave your own infrastructure, and the available controls are limited. Anyone who wants precise, repeatable control — the same person, the same room, the same camera aesthetic — hits a wall quickly.

So we ran the counter-experiment at Context Studios: a pipeline where every component runs on our own hardware — as an experiment to understand what a single desktop device can do today. This article explains, in plain language, how the pipeline is built, which decisions we made, what measured well and what didn't — including the numbers that such reports usually leave out.

The building blocks of a self-hosted video pipeline

Before the deep end, some orientation. Four components work together:

The hardware. An ASUS Ascent GX10 — a flat, compact mini PC based on the NVIDIA DGX Spark platform. Inside the GB10 chip, an Arm CPU and a Blackwell GPU share 128 GB of unified memory.

The image layer. Before any video happens, the person has to be right — the way a director knows an actor's photos before the camera rolls. An image model (in our case Qwen-Image-2.1) generates or validates reference images of the key person: face, hair, beard, glasses, clothing.

The video model. MiniMax-H3 takes the reference images and a script-style prompt and generates a video clip — including mouth movement and even generated speech.

The voice. A voice clone (ours via ElevenLabs) speaks the German script in the client's brand voice. The catch: that voice and the mouth movement in the video must be synchronized — a processing step of its own, with its own tools and its own pitfalls.

text
The data flow in one sentence:
Reference photos → image model validates identity → video model
renders a speaking avatar → voice clone produces the audio →
lip-sync aligns mouth movement to the audio → post look → QC

The setup in detail

Our pipeline runs in ComfyUI. Think of it as a patchboard for AI models: you wire building blocks into a workflow with the mouse, and the whole thing can additionally be driven automatically by programs (via a so-called HTTP API) — exactly what automation needs. On the DGX Spark you'll find:

  • MiniMax-H3, quantized to int8 (more on that shortly), as the video model — rendering vertical 768×1344, the classic 9:16 social format
  • Qwen3-VL-32B as the text brain: a language model that reads the script prompt and translates it into a form the video model can work with
  • Two separate VAE decoders — translator components that turn the model's internal data back into real images and sound: one for video frames, one for audio. H3 generates both together
  • Dynamic memory management: the video model preloads only 19.9 GB and fetches the remaining parts when they are actually needed — like taking tools out of the cupboard only when you really need them
text
Render times (5-second clip, 24 fps):
864×480 baseline:            280 seconds
768×1344, 20 sampling steps: ~23 minutes including decode

One finding we didn't expect: the GB10 chip throttles thermally. Above 85 °C die temperature it lowers its clock by about ten percent. Over a five-hour render session that summed to 81 minutes of active throttling — and explains why identical renders took anywhere from 53 to 105 seconds per step. Anyone comparing benchmarks needs to think about cooling.

Step 1: Getting the person right — the image layer

The video model is only as good as its reference material. We used three photos of a real person and systematically tested which settings maximize identity (resemblance to the real person) and realism (natural skin, natural micro-expressions).

Quantization: int8 instead of fp8. Both variants produced identical skin-texture metrics. int8 wins because it loads faster — a free speed gain.

Fewer sampling steps can be more. 20 steps with the euler sampler and the beta57 scheduler produced clearly more natural micro-expressions (eyebrows, eye creases, blinking) than the much longer original render we had for comparison.

The realism-LoRA compromise. A LoRA is a small add-on model that pushes a base model in a particular direction — here, photorealistic skin. Our test confirmed the effect: pores, beard filaments, and eye creases became visibly more natural. But there was a catch you only notice on the second look: the LoRA also changed identity features — the hairline became more pronounced, the beard grayer and denser, the glasses different. Our fix: reduce LoRA strength from 1.0 to 0.7. The texture gain stays, the deviation drops to acceptable levels.

The trap nobody has on their radar. Our prompt deliberately contained many iPhone-realism attributes: slight hand tremor, hesitant autofocus, exposure breathing. The effect was striking — and counterproductive: the model transferred the iPhone aesthetic onto the scene and reinterpreted the setting. The sage-green interior wall became a concrete wall with bricks, outdoors. The lesson: realism attributes and setting must be anchored separately and explicitly. Our setting prompt now reads, in essence: "INDOOR: exactly the sage-green wall from reference picture 1, no outdoors, no concrete, no bricks."

Step 2: The video model — dialog as part of the render

MiniMax-H3 has a property that simplifies the pipeline decisively: it generates video and audio together. The dialog goes into a special tag inside the script prompt, and H3 produces an avatar speaking that text — with matching mouth movement and audible speech.

In our test, the model spoke the German sentence "Die meisten Leute machen KI-Video viel zu kompliziert. Drei lokale Modelle reichen vollkommen aus" intelligibly. We didn't just listen — we measured: a speech-to-text model transcribed the generated audio word for word.

The big advantage of this co-generation: because mouth movement and sound originate in the same process, they are mechanically perfectly synchronized. The beard stays untouched, the expressions natural. The downside: the generated voice sounds generic — it is not the client's brand voice. For brand voice, steps 3 and 4 are needed.

Step 3: The voice — and the hallucination trap

For the brand voice we use a voice clone at ElevenLabs: a model that forms a voice from reference recordings of a real speaker. Our clone (internally "CS-Klon") was supposed to speak the German test sentence. On first listen the result seemed unremarkable. Then we ran the decisive test: transcribing the audio word for word, without smoothing anything:

text
Intended text:  Die meisten Leute machen KI-Video viel zu kompliziert.
                Drei lokale Modelle reichen vollkommen aus.

Actually spoken: Ähm, und die meisten Leute machen KI-Video viel zu
                kompliziert und drei lokale Modelle, ähm, reichen
                vollkommen aus. Wir machen das, ähm

Three "ähm"s, two inserted "und"s, and a completely invented, cut-off extra sentence — none of it was in the prompt. We also measured the melody of the voice via pitch tracking: the fundamental frequency fluctuated uncontrollably (standard deviation 64 Hz around a 143 Hz mean — a coefficient of variation of almost 45 percent, compared to just over 28 percent after the fix). And the pacing ranged from 5.6 to 10.4 seconds for the same text with identical settings.

Why this happens — according to the official docs

The explanation was sitting in the ElevenLabs documentation; you just had to assemble it:

First: Instant Voice Cloning is a zero-shot procedure. The system mixes the reference material with generic data from a foundation model — and imitates everything it hears in that material. If the material contains breathing, throat-clearing, or filler sounds, the clone produces exactly those. The docs put it plainly: "The AI will attempt to mimic everything it hears in the audio."

Second: the voice's style parameter is more dangerous than its name suggests. The troubleshooting page recommends keeping it at 0 at all times, because it causes "inconsistent speed, mispronunciation and the addition of extra sounds." We had values up to 0.45 — the prime suspect for the extra sounds.

Third: low stability is not an expressiveness knob, it is a variance knob. The docs warn about "overly random performances that cause the character to speak too quickly." And without a set seed (a start value for the random generator), every generation is different — "determinism is not guaranteed."

The fix stack and its metrics

Our final settings: stability fixed at 0.55, similarity 0.75, style 0.0, speaker boost on, seed fixed per clip, numbers and acronyms written out, no ellipses. Plus a quality gate: every generation gets transcribed via speech-to-text — if even one filler appears, it gets regenerated.

text
                         before              after
Fillers in STT:          3× "ähm" + sentence →  0 (6 of 6 generations)
Pitch variance (CV):     44.8 %              →  28.0 %
Duration variance:       5.6–10.4 s          →  ±0.1 s (with seed)
Dynamics (LRA):          0.3–3.4 LU random   →  stable 1.4 LU
Loudness:                −33 LUFS            →  −16 LUFS (production level)

In a listening comparison of two model generations, multilingual_v2 won (short, crisp, calm prosody) against the newer v4 (more natural pacing, more voice presence). And the cleanest fix is still ahead: rebuilding the clone from 60–120 seconds of carefully cut reference material without breaths — because the clone imitates whatever it hears.

Step 4: Lip-sync — three approaches compared

When video and audio come from different sources, the mouth has to be recomputed. A market of tools exists for this — we tested three approaches.

Approach 1: Everything from one hand (H3-native)

Because H3 generates sound and image together, synchronization is perfect by construction. Intelligible German, flawless beard, natural expressions. For fast iteration, this is the best route. The price: the voice is the model's, not the client's.

Approach 2: Foreign voice plus an external lip-sync pass

For real brand voice, a sync pass is unavoidable: a specialized model recomputes the mouth movement to match the audio. We tested the lipsync-2 family from sync.so via the fal.ai infrastructure:

The standard model synced cleanly — but the beard around the mouth was almost completely washed away. The "freshly shaved" look. The pro model (with an explicit beard-and-teeth guarantee) kept considerably more beard structure, though a slight thinning remained. The reason is conceptual: these models re-render the entire lower face.

A second lever was the sync mode: remap instead of cut_off. In the default mode the audio is hard-cut to the video length; in remap mode it is mapped onto the mouth movement. That measurably improved both sync quality and beard preservation.

Even more important is the insight on the render side: lip-sync models of this family need an actively speaking mouth as input. Our first attempt — rendering the video with closed lips and letting the mouth be animated — produces, per the docs, only "generic results." Render with open, speaking performance; the model's (wrong) audio gets discarded in the sync pass anyway.

Approach 3: Open source on your own box

For the long-term no-API-cost solution, the open-source landscape of 2026 looks like this: KeySync from Imperial College is the strongest freely available model — two-stage diffusion, explicit occlusion handling (crucial for glasses!), and measurably better than the well-known ByteDance alternative LatentSync in the scientific comparison; LatentSync has not seen a release since mid-2025. LTX-LipDub is intriguing but text-driven rather than audio-driven — the wrong approach for voiceover files. MuseTalk is out for beards.

KeySync is on our roadmap: the ComfyUI integration exists, the checkpoints weigh about 24 GB. The open question is the build on the GB10 chip — there is no documented experience with that yet; we will report.

The experiment in 7 steps

text
1. Validate reference photos (frontal, details, setting)
2. Image model: identity verification of the person
3. H3 render: 768×1344, 20 steps, euler/beta57, fixed seed,
   realism LoRA 0.7, dialog in the language tag, open speaking
   performance, setting explicitly anchored
4. Voice clone with fix stack + STT gate against fillers
5. lipsync-2-pro with remap mode
6. Post look: subtle contrast, light grain, warm tones —
   iPhone aesthetic without HDR gloss
7. QC: frame inspection plus STT verification of the final mix

About 30 minutes per clip, 23 of which are the H3 render. The chain is fully scriptable and worked reproducibly in our test — as an experimental setup, not a production line.

What we learned

What works: co-generated audio for fast iteration. External brand voice with pro lip-sync and remap mode. The documented fix stack against voice-clone hallucinations. Word-accurate STT analysis as an automated quality gate — it surfaced problems that were inconspicuous by ear. Seed fixation for reproducible takes. int8 quantization as a free speed gain.

What doesn't work: style exaggeration on voice clones (documented, but easy to miss). Low stability as an "expressiveness" control. Closed lips as a sync target. The default sync mode with mismatched audio/video lengths. Hoping the beard and glasses frame survive standard lip-sync.

The honest overall impression: the experiment works. On a single desktop device, a clip with synchronized lips and a foreign voice emerges in half an hour. For production use, however, we judge it too early — the quality is good, not good enough to simply generate content from it and post it. The last residual artificiality sits not in the synchronization but in the micro-expressions: a trained eye recognizes the limit of the current model generation. That residual not-quite-real quality is exactly why we documented the project as an experiment rather than taking it into production.

Watch the results

So the report doesn't just claim what it explains: here are the two final clips from the test. The first shows the brand-voice variant — German voiceover from the voice clone, lip-sync via the pro pass, post look in iPhone aesthetics. The second shows the H3-native variant, where the video model generates speech and mouth movement together and synchronization is therefore mechanically perfect.

Variant A: CS clone voice with lipsync-2-pro

Variant A: CS clone voice (ElevenLabs) with lipsync-2-pro and remap mode — the final workflow.

Variant B: H3-native, co-generated audio

Variant B: H3-native — speech and mouth from one diffusion run, beard and expressions untouched.

Both clips come from the same render (identical seed, identical scene prompt) — the difference lies entirely in the audio and sync chain. That comparison taught us where each method's strengths lie.

The pipeline as a diagram

The complete data flow at a glance — from the reference photos through the H3 render to the two result variants:

Data flow of the self-hosted AI video pipeline: reference photos, Qwen-Image identity check, MiniMax-H3 render, ElevenLabs voice clone with STT gate, lipsync-2-pro remap, post look and quality control
Data flow of the self-hosted AI video pipeline: reference photos, Qwen-Image identity check, MiniMax-H3 render, ElevenLabs voice clone with STT gate, lipsync-2-pro remap, post look and quality control

Sources

Frequently asked questions

Why self-host at all when lip-sync and video models exist as APIs?

Three reasons. First, cost: a 5-second clip quickly costs several euros via cloud APIs, and iterative work multiplies that. Second, data control: customer recordings and reference photos never leave your own infrastructure. Third, control over every dial: seed, sampler, scheduler, LoRA strength, and render path are only fully accessible self-hosted — exactly the grips we needed to fix identity drift and setting errors. The API remains in our pipeline where it is better: the lip-sync pass.

Is lip-syncing a foreign voiceover file really feasible cleanly?

Yes, but not by simply overlaying the audio. We render the video with an open, speaking performance, let a specialized model recompute the mouth movement to exactly match the audio — in remap mode, which maps the audio onto the movement instead of cutting it — and verify the result by frame inspection. The documented compromise: the beard in the mouth area thins slightly, because the model re-renders the lower face. The pro model mitigates that considerably; open-source models with occlusion masks are the next step.

Why does a voice clone hallucinate filler words like "ähm" that were never in the text?

Instant voice cloning is a zero-shot procedure: the model mixes the reference material with generic foundation data and imitates everything it hears — including breathing and filler sounds. If the prompt additionally contains acronyms, numbers, or ellipses, the model improvises even more. The countermeasures are officially documented and worked in our test: clean reference material without breathing, style parameter at 0, stability between 0.5 and 0.65, a set seed, writing out everything that can be written out — and verifying the result via speech-to-text before it enters the pipeline.

How close is the result to a real iPhone video?

Closer than we expected, with a clear limit. The combination of handheld-camera attributes in the prompt, subtle grain, warm color mood, and the absence of HDR gloss holds up well in still-frame comparison. The trained eye stays on the micro-expressions: blinking and eyebrows feel natural, but the range of facial expressions is narrower than in a real recording. For social clips of 5 to 15 seconds with a calm camera, the difference is nearly invisible; for long formats, our assessment is that another model generation is needed.

Is such a device worth it for video experiments?

For our experiment: yes — with open eyes. The 128 GB of shared memory allow the video model and image model to stay loaded simultaneously, and about 25 minutes per clip is workable for iterative production. Against that stand thermal throttling under sustained load and the ecosystem disadvantage: newer open-source models have no documented build for this chip, and graphics libraries sometimes need patching first. The GX10 is not a render-farm replacement but a fascinating experimental device for a person or a small team — and that is exactly how we used it.

What is the next step to remove the last bit of artificiality?

Three levers in priority order. First, KeySync locally on the box: beard-friendly lip-sync without API costs, though the build effort on the GB10 architecture is real. Second, rebuilding the voice clone from 60–120 seconds of filler-free, breath-free reference material — the documented root-cause fix for the hallucinations. Third, best-of-N selection: generate several takes per clip and automatically pick the best via speech-to-text plus image metrics, instead of betting on a single generation.

Relevant for your team? Let's talk for 30 minutes.

We sort out what of this actually works in your company — concrete, no slide marathon.

No commitment · 30 minutes · Proposal within 48 h