ta-eros-dub-01
df059e371b · speaker SPEAKER_01 · in at 00:00:11.880
The darkest clip in the set — mean luma 22.9 of 255. Watch whether the mouth interior stays consistent when there is almost no fill light.
Eros Investments · Track B mask-free dubbing · internal review
Every clip below is one frozen test case run through the locked public X-Dub baseline on Eros footage. Left half is the untouched original; right half is the model's output. Only the mouth region is meant to change — everything else is the measure of whether it behaved.
These are outputs of the public X-Dub baseline, the reference implementation this project must reproduce before changing anything. They are not from an Eros-trained model — that model has not been trained yet, and cannot be until the Phase 2 dataset exists.
Their purpose is evidence: the baseline runs correctly on our own footage, at the locked settings, with every output hashed into an immutable run record. That is what Phase 1 acceptance certifies.
Conditions were measured from the frames themselves — brightness and head yaw — never guessed. Every clip is verified free of shot cuts, because generating across a cut is prohibited. The first three come from the accepted acceptance run; 06–08 are a supplementary frontal set, found by measuring yaw across 29 candidate windows and keeping only those under 0.35. Two further accepted cases, ta-eros-dub-02 and ta-eros-dub-03, are profile and rapid-motion shots and stay in the run record rather than here.
df059e371b · speaker SPEAKER_01 · in at 00:00:11.880
The darkest clip in the set — mean luma 22.9 of 255. Watch whether the mouth interior stays consistent when there is almost no fill light.
57afc7013a · speaker s001_SPEAKER_01 · in at 00:10:53.634
Soft, low-contrast frame (Laplacian variance 60 — the flattest in the set). Tests whether the model invents detail that was never in the source.
57afc7013a · speaker s002_SPEAKER_01 · in at 00:15:07.666
The cleanest case: well-lit, frontal, near-static. This is the upper bound — if lip-sync is weak here it is weak everywhere.
57afc7013a · speaker s000_SPEAKER_02 · in at 00:00:15.793
First of the supplementary frontal set. Found by measuring head yaw across 29 candidate windows and keeping only those under 0.35.
57afc7013a · speaker s000_SPEAKER_02 · in at 00:02:27.853
The most squarely frontal clip found anywhere in the library — yaw 0.11 means the nose sits almost exactly between the eyes.
57afc7013a · speaker s001_SPEAKER_01 · in at 00:11:59.353
A different speaker from the other two supplementary clips, so the set is not one face repeated.
The first run on this host, 30 July, using X-Dub's own Apache-2.0 example clip. Kept as the control: it proves the pipeline itself, independent of Eros material.
Run baseline-20260730t181359z-4db37b75 · 444 s on one GPU · peak VRAM 20,739 MiB
Written into config/phase1/status.json so this run can never be mistaken for full coverage.
The suite has no Hindi or Malayalam case, though both are required for complete Phase 1 coverage. Accepted as partial by owner decision.
Measured coverage is frontal, profile, low-light and rapid-motion. Neither an occlusion case nor a film-grain case was found in the cut-free windows.
Source films run at 24 fps; the locked contract is 25. Clips were conformed at extraction, which is recorded with the rights declaration.
These films cut fast. Only five candidate windows were free of a shot boundary; generating across a cut is prohibited, so the rest were rejected.