Great Indian Summer is a 3:14 music video generated end to end — the city, the dancers, the rooftops, the track. The face is mine, and it is the only real thing in it. This page is the part that doesn't show: eight versions, and every way the film was wrong while it looked finished.
One frame from each window. Every one is a separate generation, carrying the same face across seven locations — a percussion pit, a festooned terrace, a rooftop above the city, a stepwell, the first monsoon, a mango market and a kulfi street at noon.
The first assembly played start to finish, 194.5 seconds, every cut landing on its boundary. It looked like a film. Three of its shots were from a completely different film — clips from an unrelated project that had leaked into the folder — and five more were playing in the wrong windows. Nothing flagged any of it.
The fix was provenance, not resemblance. Every window had been generated against its own eight-second slice of the song, and all twenty-five slices are different from one another. So each shot could be matched back against the slice it was actually submitted with. Its true window scores 1.00; the runner-up scores 0.27 to 0.65. The question that works is not which window does this look like — it is which window was this made for.
A second, independent signal agreed on 32 of 33 placed shots: the location named in the prompt is the location that window is supposed to be. The five misplaced shots turned out to be exactly the five scoring worst on lip-sync. The clips were never bad. The slots were wrong.
The same pass recovered 23 renders that had been generated and never downloaded. One window went from a single candidate to eight. Best take now wins each window instead of only take.
A lip-sync score asks: does this mouth follow the vocal? It never asks the question that actually broke the film on playback: is he mouthing when the song has no words? A shot can score perfectly well and still show a man rapping through a drum break.
The transcript could not find these either. Speech recognition merges and pads segments on a sung track, and it reported one silent window as 100% covered. It was not.
Measuring the signal found them. A lead vocal is mixed to the centre, so it survives in the mid channel (L+R) and cancels in the side channel (L−R). Comparing the two across 300–3400 Hz gives vocal presence moment by moment, and it agreed with what could be heard. The film's silent stretches are now covered by shots where the mouth is not readable at all — full body, overhead, back to camera.
The worst window in the film had never synced, through every re-selection of every available take. Pulling the original job's payload showed why: it had been submitted with an identity photograph and no vocal attached at all. It was never given the thing it was meant to follow. Re-sent with its vocal, same prompt, same photograph: 0.31 → 0.73.
It had looked exactly like a model limitation. It was a submission bug. Worth checking what you actually sent before concluding the model can't do it.
A related finding, from the same investigation: handing the vocal over as a black-frame video scores 0.91–0.95. Handing over the identical vocal as an audio file scores 0.10–0.21. Same audio, same prompt — the container it arrives in decides whether the model listens.
Right through version 7, the audio was clipping. The source track is mastered hot — it peaks above full scale — and every version passed it straight through, delivering at −0.002 dBFS.
Every automated probe said the file was correct, because by every structural measure it was: right container, right codec, right duration, right frame count. None of them listen to the audio. The defect only appears if you measure the encoded file itself.
Fixed with a limiter and loudness normalisation rather than a blunt gain cut — a first attempt at −3 dB cleared the peak but dropped the film four LUFS quieter. The shipped version sits at a true peak of −1.06 dBFS with the loudness unchanged. Nothing sounds quieter; the peaks just came off the ceiling.
Eight of the twenty-five windows are not lip-synced, and cannot be. They do not contain enough vocal to sync against — one is only 53% vocal, and all ten takes ever generated for it score between 0.00 and 0.21. Those windows are chosen as shots: wide, overhead, turned away, mouth unreadable. It is the right answer, but it is a workaround.
Two windows have no vocal anywhere on record, so nothing can ever sync there. They are filled with same-look takes used as cutaways.
Across the finished film: 11 windows locked at 0.85 or better, 5 good between 0.6 and 0.85, and 8 carried by the shot rather than the sync.
| Version | What changed |
|---|---|
| v1 | First assembly. Three shots from an unrelated film, five more in the wrong windows, mono audio downmix. |
| v2 | Every shot matched back to the window it was generated for. 23 never-downloaded renders recovered. Stereo restored from the original master. |
| v3 | Mouthing over instrumental stretches found by measuring vocal presence, and covered with unreadable-mouth shots. |
| v4 | The never-syncing window re-rendered with the vocal it had never been sent. Final cut moved onto the vocal-out so the silent tail sits on one non-mouthing shot. |
| v5 | That window swapped to the other take — chosen on the shot rather than the number. |
| v6 | Six re-renders. One window went 0.00 → 0.96, the best sync in the film. Another was proven unfixable and reclassified as shot-led. |
| v7 | A hero shot returned to the window it was originally generated for. |
| v8 | Audio de-clipped. True peak −0.002 → −1.06 dBFS, loudness unchanged. |