Eros Investments · Track B mask-free dubbing · internal review

Phase 1 acceptance: Tamil dubs, side by side with their source

Every clip below is one frozen test case run through the locked public X-Dub baseline on Eros footage. Left half is the untouched original; right half is the model's output. Only the mouth region is meant to change — everything else is the measure of whether it behaved.

Runbaseline-20260803t072705z-f4b61142
Statusaccepted · 5/5 cases
Suitetrack-b-phase1-acceptance-ta-v1
Suite sha256809ea01a1c00ab6d…
Review sha256ccfda7d584dd9936…
Settings512² · 25fps · 77f · 30 steps

What this is — and what it is not

These are outputs of the public X-Dub baseline, the reference implementation this project must reproduce before changing anything. They are not from an Eros-trained model — that model has not been trained yet, and cannot be until the Phase 2 dataset exists.

Their purpose is evidence: the baseline runs correctly on our own footage, at the locked settings, with every output hashed into an immutable run record. That is what Phase 1 acceptance certifies.

The cases

6 clips · Tamil · frontal · shot-local

Conditions were measured from the frames themselves — brightness and head yaw — never guessed. Every clip is verified free of shot cuts, because generating across a cut is prohibited. The first three come from the accepted acceptance run; 0608 are a supplementary frontal set, found by measuring yaw across 29 candidate windows and keeping only those under 0.35. Two further accepted cases, ta-eros-dub-02 and ta-eros-dub-03, are profile and rapid-motion shots and stay in the run record rather than here.

01

ta-eros-dub-01

df059e371b · speaker SPEAKER_01 · in at 00:00:11.880

frontallow-light
luma
22.9
yaw
0.26
motion
0.8px/f
frames
108
gen
278s

The darkest clip in the set — mean luma 22.9 of 255. Watch whether the mouth interior stays consistent when there is almost no fill light.

04

ta-eros-dub-04

57afc7013a · speaker s001_SPEAKER_01 · in at 00:10:53.634

frontallow-light
luma
33.5
yaw
0.27
motion
3.7px/f
frames
107
gen
265s

Soft, low-contrast frame (Laplacian variance 60 — the flattest in the set). Tests whether the model invents detail that was never in the source.

05

ta-eros-dub-05

57afc7013a · speaker s002_SPEAKER_01 · in at 00:15:07.666

frontal
luma
73.7
yaw
0.32
motion
2.1px/f
frames
92
gen
328s

The cleanest case: well-lit, frontal, near-static. This is the upper bound — if lip-sync is weak here it is weak everywhere.

06

ta-eros-dub-06

57afc7013a · speaker s000_SPEAKER_02 · in at 00:00:15.793

frontallow-light
luma
33.4
yaw
0.32
frames
114
gen
377s

First of the supplementary frontal set. Found by measuring head yaw across 29 candidate windows and keeping only those under 0.35.

07

ta-eros-dub-07

57afc7013a · speaker s000_SPEAKER_02 · in at 00:02:27.853

frontallow-light
luma
31.6
yaw
0.11
frames
95
gen
345s

The most squarely frontal clip found anywhere in the library — yaw 0.11 means the nose sits almost exactly between the eyes.

08

ta-eros-dub-08

57afc7013a · speaker s001_SPEAKER_01 · in at 00:11:59.353

frontallow-light
luma
36.4
yaw
0.27
frames
89
gen
364s

A different speaker from the other two supplementary clips, so the set is not one face repeated.

Reference case

upstream example

The first run on this host, 30 July, using X-Dub's own Apache-2.0 example clip. Kept as the control: it proves the pipeline itself, independent of Eros material.

Run baseline-20260730t181359z-4db37b75 · 444 s on one GPU · peak VRAM 20,739 MiB

Recorded limitations

carried in the acceptance record

Written into config/phase1/status.json so this run can never be mistaken for full coverage.

Tamil only

The suite has no Hindi or Malayalam case, though both are required for complete Phase 1 coverage. Accepted as partial by owner decision.

No occlusion or grain

Measured coverage is frontal, profile, low-light and rapid-motion. Neither an occlusion case nor a film-grain case was found in the cut-free windows.

24 → 25 fps

Source films run at 24 fps; the locked contract is 25. Clips were conformed at extraction, which is recorded with the rights declaration.

14 of 19 windows discarded

These films cut fast. Only five candidate windows were free of a shot boundary; generating across a cut is prohibited, so the rest were rejected.