Public benchmark

AI video continuity benchmark

Live. Updates as runs come in.

Every generator gets the same briefs, the same prompts and the same check your own clips get. Then we see who holds it together.

Loading scores…

How this is done

Every generator gets the same briefs: the same people coming back, more than one location, and some changes the story explains and some it does not. Each clip is then watched the way the analysis page describes, and the score is worked out in code from what came back. Read how.

Run by VideoConsistency. If you built one of these and think a score is wrong, ask for a re-run with the published prompts, or point at the finding and the frames and argue with me.