Eros Investments · Track B mask-free dubbing · internal review
Phase 1 acceptance: Tamil dubs, side by side with their source
Every clip below is one frozen test case run through the locked public X-Dub baseline
on Eros footage. Left half is the untouched original; right half is the model's output.
Only the mouth region is meant to change — everything else is the measure of whether it behaved.
Runbaseline-20260803t072705z-f4b61142
Statusaccepted · 5/5 cases
Suitetrack-b-phase1-acceptance-ta-v1
Suite sha256809ea01a1c00ab6d…
Review sha256ccfda7d584dd9936…
Settings512² · 25fps · 77f · 30 steps
What this is — and what it is not
These are outputs of the public X-Dub baseline, the reference implementation
this project must reproduce before changing anything. They are not from an
Eros-trained model — that model has not been trained yet, and cannot be until the Phase 2
dataset exists.
Their purpose is evidence: the baseline runs correctly on our own footage, at the locked
settings, with every output hashed into an immutable run record. That is what Phase 1
acceptance certifies.
The cases
7 clips · frontal · shot-local
The first clip is the quality benchmark: best master, largest face, most denoising steps.
Conditions were measured from the frames themselves — brightness and head yaw — never guessed.
Every clip is verified free of shot cuts, because generating across a cut is prohibited.
The first three come from the accepted acceptance run;
06–08 are a supplementary frontal set,
found by measuring yaw across 29 candidate windows and keeping only those under 0.35.
Two further accepted cases, ta-eros-dub-02 and
ta-eros-dub-03, are profile and rapid-motion shots and stay in the
run record rather than here.
best
best-dub-01
SabKushalMangal ·
speaker ECAPA-matched 0.657 ·
in at 00:03:20.000
frontalbest-master50-step
OriginalDubbed
luma
88.0
yaw
0.26
frames
100
gen
399s
The highest-quality run in this set, and the one to judge the baseline by. Source is the archive's best master (1920x1080, 26 Mbps, the only title scoring 4.28 sync confidence at a perfect 1.0 hit rate). The face measures 881 px, so the 512 crop DOWNSCALES - the model works from real detail instead of upscaled interpolation, unlike every other clip here. Well lit at luma 88 and generated at 50 denoising steps rather than 30. Comparison is 3840x1080: two full-HD frames.
01
ta-eros-dub-01
df059e371b ·
speaker SPEAKER_01 ·
in at 00:00:11.880
frontallow-light
OriginalDubbed
luma
22.9
yaw
0.26
motion
0.8px/f
frames
108
gen
278s
The darkest clip in the set — mean luma 22.9 of 255. Watch whether the mouth interior stays consistent when there is almost no fill light.
04
ta-eros-dub-04
57afc7013a ·
speaker s001_SPEAKER_01 ·
in at 00:10:53.634
frontallow-light
OriginalDubbed
luma
33.5
yaw
0.27
motion
3.7px/f
frames
107
gen
265s
Soft, low-contrast frame (Laplacian variance 60 — the flattest in the set). Tests whether the model invents detail that was never in the source.
05
ta-eros-dub-05
57afc7013a ·
speaker s002_SPEAKER_01 ·
in at 00:15:07.666
frontal
OriginalDubbed
luma
73.7
yaw
0.32
motion
2.1px/f
frames
92
gen
328s
The cleanest case: well-lit, frontal, near-static. This is the upper bound — if lip-sync is weak here it is weak everywhere.
06
ta-eros-dub-06
57afc7013a ·
speaker s000_SPEAKER_02 ·
in at 00:00:15.793
frontallow-light
OriginalDubbed
luma
33.4
yaw
0.32
frames
114
gen
377s
First of the supplementary frontal set. Found by measuring head yaw across 29 candidate windows and keeping only those under 0.35.
07
ta-eros-dub-07
57afc7013a ·
speaker s000_SPEAKER_02 ·
in at 00:02:27.853
frontallow-light
OriginalDubbed
luma
31.6
yaw
0.11
frames
95
gen
345s
The most squarely frontal clip found anywhere in the library — yaw 0.11 means the nose sits almost exactly between the eyes.
08
ta-eros-dub-08
57afc7013a ·
speaker s001_SPEAKER_01 ·
in at 00:11:59.353
frontallow-light
OriginalDubbed
luma
36.4
yaw
0.27
frames
89
gen
364s
A different speaker from the other two supplementary clips, so the set is not one face repeated.
Reference case
upstream example
The first run on this host, 30 July, using X-Dub's own Apache-2.0 example clip. Kept as the
control: it proves the pipeline itself, independent of Eros material.
OriginalDubbed
Run baseline-20260730t181359z-4db37b75 · 444 s on one GPU ·
peak VRAM 20,739 MiB
Recorded limitations
carried in the acceptance record
Written into config/phase1/status.json so this run can never be
mistaken for full coverage.
Tamil only
The suite has no Hindi or Malayalam case, though both are required for complete Phase 1
coverage. Accepted as partial by owner decision.
No occlusion or grain
Measured coverage is frontal, profile, low-light and rapid-motion. Neither an occlusion
case nor a film-grain case was found in the cut-free windows.
24 → 25 fps
Source films run at 24 fps; the locked contract is 25. Clips were conformed at
extraction, which is recorded with the rights declaration.
14 of 19 windows discarded
These films cut fast. Only five candidate windows were free of a shot boundary; generating
across a cut is prohibited, so the rest were rejected.