built with VideoClaw

Twenty-five shots.
A song that doesn't exist.
One real face.

Great Indian Summer is a 3:14 music video generated end to end — the city, the dancers, the rooftops, the track. The face is mine, and it is the only real thing in it. This page is the part that doesn't show: eight versions, and every way the film was wrong while it looked finished.

GREAT INDIAN SUMMER3:141080×1920 · 24fps 25 windows8 versions
25
generated windows
8
master versions
3
shots from another film
1
real person

The film in twenty-five frames

One frame from each window. Every one is a separate generation, carrying the same face across seven locations — a percussion pit, a festooned terrace, a rooftop above the city, a stepwell, the first monsoon, a mango market and a kulfi street at noon.

lip-synced to the vocal chosen as a shot — too little vocal in the window to sync to

Version 1 looked finished. It wasn't.

The first assembly played start to finish, 194.5 seconds, every cut landing on its boundary. It looked like a film. Three of its shots were from a completely different film — clips from an unrelated project that had leaked into the folder — and five more were playing in the wrong windows. Nothing flagged any of it.

The fix was provenance, not resemblance. Every window had been generated against its own eight-second slice of the song, and all twenty-five slices are different from one another. So each shot could be matched back against the slice it was actually submitted with. Its true window scores 1.00; the runner-up scores 0.27 to 0.65. The question that works is not which window does this look like — it is which window was this made for.

A second, independent signal agreed on 32 of 33 placed shots: the location named in the prompt is the location that window is supposed to be. The five misplaced shots turned out to be exactly the five scoring worst on lip-sync. The clips were never bad. The slots were wrong.

The same pass recovered 23 renders that had been generated and never downloaded. One window went from a single candidate to eight. Best take now wins each window instead of only take.

The score answered the wrong question

A lip-sync score asks: does this mouth follow the vocal? It never asks the question that actually broke the film on playback: is he mouthing when the song has no words? A shot can score perfectly well and still show a man rapping through a drum break.

The transcript could not find these either. Speech recognition merges and pads segments on a sung track, and it reported one silent window as 100% covered. It was not.

Measuring the signal found them. A lead vocal is mixed to the centre, so it survives in the mid channel (L+R) and cancels in the side channel (L−R). Comparing the two across 300–3400 Hz gives vocal presence moment by moment, and it agreed with what could be heard. The film's silent stretches are now covered by shots where the mouth is not readable at all — full body, overhead, back to camera.

One failure wasn't the model's fault

The worst window in the film had never synced, through every re-selection of every available take. Pulling the original job's payload showed why: it had been submitted with an identity photograph and no vocal attached at all. It was never given the thing it was meant to follow. Re-sent with its vocal, same prompt, same photograph: 0.31 → 0.73.

It had looked exactly like a model limitation. It was a submission bug. Worth checking what you actually sent before concluding the model can't do it.

A related finding, from the same investigation: handing the vocal over as a black-frame video scores 0.91–0.95. Handing over the identical vocal as an audio file scores 0.10–0.21. Same audio, same prompt — the container it arrives in decides whether the model listens.

The defect no tool could see

Right through version 7, the audio was clipping. The source track is mastered hot — it peaks above full scale — and every version passed it straight through, delivering at −0.002 dBFS.

Every automated probe said the file was correct, because by every structural measure it was: right container, right codec, right duration, right frame count. None of them listen to the audio. The defect only appears if you measure the encoded file itself.

Fixed with a limiter and loudness normalisation rather than a blunt gain cut — a first attempt at −3 dB cleared the peak but dropped the film four LUFS quieter. The shipped version sits at a true peak of −1.06 dBFS with the loudness unchanged. Nothing sounds quieter; the peaks just came off the ceiling.

What is still wrong with it

Eight of the twenty-five windows are not lip-synced, and cannot be. They do not contain enough vocal to sync against — one is only 53% vocal, and all ten takes ever generated for it score between 0.00 and 0.21. Those windows are chosen as shots: wide, overhead, turned away, mouth unreadable. It is the right answer, but it is a workaround.

Two windows have no vocal anywhere on record, so nothing can ever sync there. They are filled with same-look takes used as cutaways.

Across the finished film: 11 windows locked at 0.85 or better, 5 good between 0.6 and 0.85, and 8 carried by the shot rather than the sync.

The honest summary: this is not a film where the technology disappeared. It is a film where the failures were found, named and either fixed or worked around — and three of the four worst ones had already passed every automated check.

Eight versions

VersionWhat changed
v1First assembly. Three shots from an unrelated film, five more in the wrong windows, mono audio downmix.
v2Every shot matched back to the window it was generated for. 23 never-downloaded renders recovered. Stereo restored from the original master.
v3Mouthing over instrumental stretches found by measuring vocal presence, and covered with unreadable-mouth shots.
v4The never-syncing window re-rendered with the vocal it had never been sent. Final cut moved onto the vocal-out so the silent tail sits on one non-mouthing shot.
v5That window swapped to the other take — chosen on the shot rather than the number.
v6Six re-renders. One window went 0.00 → 0.96, the best sync in the film. Another was proven unfixable and reclassified as shot-led.
v7A hero shot returned to the window it was originally generated for.
v8Audio de-clipped. True peak −0.002 → −1.06 dBFS, loudness unchanged.

How it was built

01
Write the song
Lyrics first, then a generated 3:14 track. Everything downstream is timed to it.
02
Cut the song into windows
Twenty-five eight-second slices. Each one becomes a shot's brief and its sync reference.
03
Lock the face
One photograph, carried into every prompt as an identity reference — face only, wardrobe set per scene.
04
Generate to the vocal
Each window rendered against its own slice, handed over as video rather than audio so the model actually follows it.
05
Verify, don't assume
Provenance matching, vocal-presence measurement, frame counts, freeze and black-frame scans, and a true-peak reading of the finished file.
06
Version, never overwrite
Every master kept. Each version records what changed and why, so a regression is visible instead of mysterious.
Everything in this film is generated except one thing. The dancers, the crowds, the city, the rooftops and the song do not exist and depict no real person. The face is mine, generated from a photograph of me, with my consent.