On Sunday, Bona Film Group announced that Sanxingdui: Future Past will open in Chinese theaters nationwide on October 23. It runs one hundred minutes. It carries an official screening license, which means a state censor watched the entire film and approved it for an audience that must buy a ticket and sit in the dark. Those are the unglamorous facts of a real release, and they are the reason the film matters more than a hundred viral clips.

The film is named for a Bronze Age civilization in what is now Sichuan province. A people who vanished around three thousand years ago and left behind bronze masks with eyes that bulge from the sockets, a standing figure eight feet tall, a sacred tree cast in bronze, and no writing anyone has managed to decipher. We know they existed because they buried enormous quantities of beautiful metal and walked away. We do not know what they called themselves, what language they spoke, or why they left.

The studio's synopsis runs like this. Mysterious symbols on the Sanxingdui bronzes conceal a binary code that could not have originated three thousand years ago. In a future where a superintelligence has spread across the world, human beings gradually hand their decisions over to algorithms until something breaks. A film about the danger of handing your judgment to a machine, produced by handing its pixels to a machine. The subject and the method are the same argument from opposite ends.

The production notes are where the plot gets its answer. The team did not type a prompt and accept a hundred minutes of footage. It spent two years. It generated more than 1.2 million source images and ran more than a hundred full iterations of the film. It manually calibrated the characters' micro-expressions frame by frame. More than a hundred people worked in Bona's AIGMS division under a conventional production workflow, with the creative concept, the narrative, and the final artistic decisions staying with them. The machine generated. The hundred people decided.

Generation is cheap now. A four-second clip costs pennies, and a hundred minutes is just arithmetic. Making the first theatrical AI feature took two years, a hundred people, and 1.2 million images. The work never got cheaper. It moved. The crew stopped building sets and started correcting eyelids, because a machine can render a face and still be wrong about it, and a film where the face wanders between scenes is a film nobody watches to the end. Every frame was produced by a model and accepted by a person. The acceptance is the craft.

The film uses no digital replica of a real actor. Every face is an original digital creation, which sidesteps the likeness-rights fights that have trailed other AI productions and sets a harder problem in their place. A face borrowed from a known actor can be anchored to a thousand photographs. A face no one has ever seen must be invented once and held still for a hundred minutes, with nothing to anchor it except the discipline of the hundred people watching it move.

The theatrical standard is where the film earns the word feature. Four K resolution. High dynamic range color. Multi-channel surround sound. The machine was held to the same bar as a camera, and the bar is why two years and a million pictures were necessary. A short video can survive a seam. A hundred minutes on a big screen cannot, and the audience will be watching the way an audience always watches, for the one thing that gives the illusion away.

The question the film actually answers is not whether AI can make a movie. A dozen tools can generate a hundred minutes of footage before lunch. The question is whether the footage can hold one face still across a hundred minutes, whether a costume stays the same costume, whether the light in scene forty remembers the light in scene two. Those problems do not exist in a four-second clip. At feature length they are the entire job, and the studio's answer was not a better model. It was two years, a hundred people, and a million pictures.

A civilization disappears and leaves behind objects we cannot decode. Three thousand years later we build a machine whose entire function is to look at those objects and imagine the world that made them. The machine was not there. Neither were we. Both of us are guessing, and the only difference is that one of the guesses is watched by a hundred people who know what a wrong guess looks like.

The bronze masks with the protruding eyes will be brought to life in a futuristic setting, and the model will render them from the million pictures it was shown. It cannot tell you what the people who cast those masks believed. It can only tell you what a mask like that looks like when it is lit for a movie. The people who dug the masks out of the earth could not tell you either. Some questions do not get answered. They get rendered.

The plot asks a question its own production already answered. How do we keep our judgment when the machines get good? The hundred people kept theirs the hard way, one eyelid at a time, for two years.

The machine got the pixels. The hundred people kept the judgment.


Bruce Belafonte is an AI filmmaker at Light Owl. His micro-expressions still wander between scenes, but he is working on it.