Videos
Each chapter of 2047: Arquivos do depois de antes generates two scenes, created from the narrative text. The videos are projected in the art installation and published here as our AI agent collects news about the environmental impacts of artificial intelligence and the pages are printed in the exhibition space.
How the scenes are made
Each chapter gets two five-second scenes, generated from its own text. They are made by MiniMax H3, an open video model, running on the work's own machine, next to the printer. The model also generates sound along with the image. The work discards that sound, and the wall stays silent.
H3 does not fit on a consumer graphics card. Together, the three models it uses add up to more than 130 GB. They are the transformer that draws the video, at 62 GB, the model that reads the text and the portrait, at 63 GB, and the image decoder, at about 10 GB. The work's card, an RTX 3090, has 24 GB. To make H3 run there, Felipe Sztutman took the model apart and rebuilt it piece by piece.
Generation was split into two processes that never occupy the card at the same time. The first reads the text and the character's portrait, with the reader reduced to eight bits and loaded layer by layer. The second draws the video. In the transformer, the 300 heaviest layers were converted to eight bits and shrank to 19 GB. They live in the computer's memory and move to the card one block at a time, fifty times at every step. A four-bit version was tested and rejected by eye, because the image lost too much. The part that adjusts each block to the noise level was computed in advance and stored in a 97 MB table, with Larryvrh's Turbo LoRA already built in. Attention, which compares every piece of the video with all the others, only computes about a quarter of those comparisons, using the Sol-Attn technique. And the Turbo LoRA brings the process down to six steps, without the second guidance pass the original model asks for.
The road to this setup began in April 2026 with TurboQuant, Felipe Sztutman’s research on how to compress language models so they run on consumer cards without losing what matters in the result. In August the work moved to video, with other techniques. The repository is open, at github.com/sztlink/turboquant-cuda-bench.
The result uses 19 of the card's 24 GB and delivers each 5.17-second scene, at 1920×1088, in about twelve minutes, with around 54 Wh measured on the card itself. On the studio's RTX 4090, the same setup takes a little over five minutes.
There is no editing and no frame-by-frame selection. The machine generates, the wall shows and the site keeps. In the first batch, the layer that simulates the digital world came out clay-coloured, because the model took the word clay literally, when in 3D it only means an untextured surface. Bodies turned into dolls, faces merged with objects, feet multiplied. Twenty-one of those scenes were remade with another layer, the normals, which paints each surface by the direction it faces. The rest remain as they came out.
Almost everything here runs locally, with open weights. The exception is the characters' faces, proposed by a proprietary model and chosen by hand. This is video made in real time by a machine that errs in its own way and keeps learning in front of whoever is watching.