Taste without a verifier
Coding agents somehow learned to build good-looking websites, although nothing checks whether a design is beautiful. What video editing can borrow from that: preference models, critics and loops.
I’ve been trying to understand something I’ve noticed this year. Ask a current coding agent to build a landing page, and the result usually looks good: sensible spacing, a coherent type scale, restrained color, decent hierarchy. Two years ago the same request produced something that worked and looked like 2009. Nothing in software checks whether a design is beautiful. Tests check that the button submits the form. So where did the taste come from?
In 2024, writing about AlphaProof, I drew a line between domains with a verifier, like formal math, code and games, where self-play wins, and domains without one, like aesthetics, where “correct” doesn’t exist. I said taste would be the harder, less settled problem. Coding agents’ design sense is evidence about how it gets partly solved, and I think the lessons carry over to video editing, which is what I spend my time on now.
The following is my best guess from public research and the outputs I’ve seen; I don’t know how any particular lab trained its system.
Learned preference models. People compare two outputs and pick the better one, at scale, and a reward model learns to predict their choice. This is RLHF, but aimed at visual output. The NIMA paper I wrote about in 2018 already showed that predicting the distribution of human ratings, not one score, captures disagreement. The reward doesn’t have to be right in any absolute sense. It has to be closer to what people prefer than the model’s default is.
Rendering in the loop. Agents that can take a screenshot of what they built and look at it do much better than ones that only read their own code. The critic judges the output as a person would see it, not the source.
Critique and revise. The agent proposes, renders, critiques against explicit principles (hierarchy, contrast, alignment, whitespace) and against a learned sense of what’s good, then revises. Several passes. It’s how human designers work too.
Strong priors from good examples. Models have seen a lot of well-designed websites and design system documentation. A lot of taste is knowing the conventions that work and applying them consistently.
Now try the same for editing. A preference model for video edits: show people two cuts of the same footage, ask which is better, collect it at scale across styles and genres, and model disagreement between audiences explicitly. Rendering in the loop: the editing agent renders its timeline and watches it, with audio, because pacing only exists in playback. Critique and revise against explicit principles, like cutting on action, holding for reactions, hitting the beat, and not crossing the line, plus the learned preference. Style conditioning: “edit this like a trailer,” “like a memory,” “like a documentary,” each with its own preferences, since a good documentary cut is a bad trailer cut.
The failure mode to watch is the one from every reward model: optimize against it hard enough and the agent finds what the model likes that people don’t. For websites that looks like every page converging on the same tasteful gray minimalism. For video it would be every edit converging on the same punchy, beat-synced template. The defense is keeping people in the loop and treating the critic as advice. The final call about what’s good stays with a person who has taste.