All guides

AI Video line · stop 07 of 14 · 30 min · members

Training a face you can reuse, and the dataset choices that decide whether it holds

Building an identity you can call up months later, with the dataset choices that decide whether it holds.

Free with an account

Sign in to read.

Membership is free: an account opens all 86 script pages. The Lab, Studio Canvas and the paid guides need the $99 pass, paid once. Already signed in on this browser? The page opens by itself.

01

When it is worth it

Long sequences, over weeks, across situations a single reference cannot reach.

Below that, a locked reference frame is faster and good enough.

Training an identity has a fixed cost: assembling a dataset, running the training, testing it. For a six-shot job that cost is not repaid, and a single well-chosen master image carried as a reference does the work.

It becomes worthwhile when you need the same person across dozens of shots, in angles and expressions a single frame cannot support, over a period long enough that you will come back to it.

The decision is about coverage more than volume: if the shots you need stay within what one reference shows, you do not need training.

02

Consent first

A face belongs to the person whose face it is.

This is the part to settle before any of the technical work.

Training on your own likeness is straightforward. Training on anyone else's requires their explicit agreement, in writing, for the specific use — and that agreement should say what the identity will be used for and for how long.

For client work involving a presenter or model, this belongs in the contract alongside usage rights for the images themselves. It is a normal conversation and it is much easier to have at the start than afterwards.

Training on a face without permission is not something to do, however easy the tooling has become. Treat the requirement as fixed rather than as a formality.

03

The dataset

Vary everything except the person.

This single principle determines almost all of the result quality.

The common instinct — supply the best photographs — produces an identity that reproduces those photographs. What you want is the subject constant and the circumstances varied.

  • Multiple angles including both profiles and three-quarters.
  • Multiple lighting conditions, including unflattering ones.
  • Varied backgrounds, ideally uninteresting.
  • Mostly neutral expressions with a few others.
  • A range of distances, from close to full figure.

Twenty varied images reliably beat sixty similar ones. Similar images teach the model a pose rather than a person.

04

Captioning

Describe the circumstances, not the face.

Anything you describe becomes conditional on that description.

Caption the background, the lighting, the framing and the clothing — the things that vary. Leave the person's permanent features undescribed.

The reasoning: described attributes become associated with their words rather than with the identity itself. Describe the face carefully and it becomes dependent on those words appearing; leave it alone and it becomes the thing the identity carries.

Use a distinctive trigger word consistently across all captions, so you can invoke the identity deliberately rather than having it influence everything you generate.

05

Testing it

Ask for something the dataset never showed.

The only test that separates learning from memorisation.

Generate the person in a location, pose and light that appear nowhere in the training images. If the identity holds, the training worked. If it drags a familiar background along, refuses an angle, or reverts to a training pose, it memorised.

The remedy for memorisation is always the dataset — more variety, fewer near-duplicates — not more training steps. Additional training on a narrow set makes the problem worse.

Test at several strengths as well. An identity that only works at maximum strength is one that will fight every other instruction in your prompts.

06

Living with it

Store the dataset, the trigger and the working strength together.

A trained identity is a long-lived asset and needs its context.

Keep the training images, the trigger word, the strength you typically use and a few reference outputs in one folder with the model file.

You will want to retrain eventually — with better images, on a newer base model — and reconstructing which images were used a year later is not realistic.

Record the consent as well, alongside the dataset. Knowing what you were permitted to do, and when that permission was given, is part of the asset.