Text can now be turned into a video. AI learned this from thousands of hours of recording

A team of machine learning engineers from Facebook's parent company Meta (identified as extremist)

organization, activities are prohibited forterritory of the Russian Federation) introduced a new system called Make-A-Video. As the name suggests, this AI model makes videos. It works simply: the user enters a rough description of the scene, and the system generates a short video corresponding to the text.

"teddy bear painting a portrait"

In a message announcing Make-a-Video, the companynotes that video creation tools are invaluable “for content creators and artists.” But as with text-to-image models, there are troubling prospects. The results of these tools can be used for disinformation and propaganda.

Top left:a dog in a superhero cape flies through the sky. Top right: The spacecraft lands on Mars. Bottom left: The artist's brush is painting on the canvas in close-up, in great detail. Bottom right: horse drinking water.

In a document that describes the technical detailsmodels, the authors of the development tell how it works. Make-A-Video is trained on image-caption pairs as well as unlabeled video footage. The training content was obtained from two datasets (WebVid-10M and HD-VILA-100M). They contain millions of videos with hundreds of thousands of hours of footage. There are also stock videos created by sites like Shutterstock and random videos from the Internet.

So far, Make-A-Video outputs 16 frames of video at 64 by 64 pixels, which are then scaled up to 768 by 768 using a separate AI model.

Read more:

It turned out what happens to the human brain after one hour in the forest

It became known which tea destroys protein in the brain

Strange sea creatures in the depths of the ocean turned out to be similar to humans