The assembled Tavus training file and the clips it was built from. Everything here is a compressed review copy; the masters are in the project folder.
2:07. Talking clip, settle blend, then sixty seconds of ping-ponged idle, as one continuous file. 1072×1904 at 30fps, continuous audio throughout. This is what goes into Tavus.
The only bit that really needs your eye. Eleven seconds spanning both cuts: the talking clip ends, the settle blend runs about three seconds, then the silent idle begins. Watch for a hitch, a jump in his width, or a shift in colour across either cut.
64 seconds from the still plus your ElevenLabs track, through Kling's Digital Human. The script is generic viseme coverage — every visually distinct mouth position, including /ʒ/.
Three of six survived. Two leaked at the mouth, one was your aesthetic reject. All three below hold a closed seam throughout and close on their own first frame.