Vinson·Li

Essay No. 136

Consistent characters are the unlock

Google's new image model, known everywhere as Nano Banana, keeps a person looking like themselves across edits and scenes. For AI storytelling, identity preservation matters more than image quality.


For the last few weeks, the image model everyone is talking about has a ridiculous name. Gemini 2.5 Flash Image appeared anonymously on a public model comparison site as “nano-banana,” topped the leaderboard, and kept the nickname after Google announced it at the end of August. I work at Google, but not on this, and I’m going by public information and my own use of it.

The headline feature isn’t image quality, although that’s very good. It’s that it keeps things the same. Upload a photo of a person and ask for them in a different outfit, a different place, a different pose, or ten years older, and the result still looks like that specific person. Edit an image in several steps in a conversation, and the parts you didn’t ask to change stay put. Combine a photo of a person with a photo of a room and they end up in the room, with the right lighting, still looking like themselves.

For a story told across multiple shots, that consistency matters more to me than another jump in image quality.

In my July post about microdramas I said the bottleneck for AI-made series is consistency: the same characters across a hundred episodes, the same wardrobe, the same sets. And in 2023 I wrote about how generated shots didn’t cut together because the character, the jacket and the props changed between them. Every AI filmmaker I talk to spends a large part of their time fighting drift. Faces shift slightly between shots, costumes change color, a character’s age wobbles. Audiences notice immediately, even when they can’t say what’s wrong.

The workflow that’s emerging looks like traditional production design. You create a character sheet for each character: front, profile and three-quarter views, a neutral expression, the costume. You create reference images for each location. Then you generate every shot’s first frame by combining those references with a description of the shot, and you animate that frame with a video model. The image model’s job is to keep identity and design locked across dozens or hundreds of first frames. Until now it couldn’t do that reliably. Now it mostly can.

It’s the same problem I spent my master’s on, from the other direction. My thesis was about separating a face’s identity, its conformation, from its expression, so one could change without disturbing the other. These models do something like that implicitly, and far better, from a photo.

The concern is the obvious one. A tool that makes it easy to put a real person convincingly into any scene makes it easy to do that to people who didn’t agree. Google adds an invisible SynthID watermark to outputs, which helps with detection. It doesn’t solve consent. I’ve been writing about that since deepfakes appeared in 2017, and each jump in capability makes it more pressing.

For storytelling, though, this is the piece that was missing. Consistent characters, then consistent locations, then consistent motion across shots. The first one just got a lot easier.

Fin.

Add a comment

Comments

Plain text

  • Loading comments…