I built a pipeline that reads a live site and returns a finished video on the other end. It pulls the real copy, the CSS palette, the computed fonts, the logo and the background pattern, then writes the shot list in numbers instead of prose: frame counts, camera positions, easing curves. Remotion renders it, so the motion graphics are React code and every scene is reproducible. First full run was a 15 second vertical reel for a bug reporting tool, every word on screen lifted verbatim off the site. Because the scenes are numbers, the canvas is a parameter, and that one master spits out 9:16, 16:9, 4:5, 1:1 and as many crops as I need without opening After Effects. The rule I learned the hard way is that fitting and working are different results. When the dominant axis flips, a stack has to become a row, and 2 of those 5 scenes needed real recomposing for 16:9. Under a cent per video, and I still look at every format one by one.