I hsve things text to video or image to video that needs more input than is currently provided. My personal computer isn't strong enough to render on its own. Alternately allow multiple image upload to allow for better video generation that can pull from available images and create scene more intimately to scale based on available inputs.