LoRA captioning for a character: what to write and what to leave out

LoRA captioning for a character comes down to 6 parts in one order: trigger word, shot and pose, outfit, place, light, framing, and never a word about her face, hair, eyes, skin or body. What the captions leave out, the trainer ties to the trigger word, so her face and hair become “her”. Get this wrong and the LoRA behind your consistent AI character learns a face that drifts.

On this page
  1. What does a caption teach a character LoRA?
  2. What goes in each of the 6 caption parts, and in what order?
  3. Why do captions leave out hair and eyes when prompts include them?
  4. What does a full set of captions look like?
  5. Should you caption by hand or automatically?
  6. What does caption dropout do?
  7. What can go wrong?

What does a caption teach a character LoRA?

A caption tells the trainer which parts of an image are allowed to change. Everything you name, like the outfit, place and light, becomes a separate detail you can swap in a prompt later. Everything you never name stays tied to the trigger word. For a character, the thing you never name is her.

The basics, one .txt per image with a matching file name and a short example, are in our guide to how many images you need and how to caption them. AI Toolkit reads captions the same way: “named the same as the images but with a .txt extension” (AI Toolkit README, checked 2 Oct 2026). This page goes one level deeper: each part, the order, a full set, and why prompts break the hair rule on purpose.

What goes in each of the 6 caption parts, and in what order?

Six parts, always in the same order, separated by commas: trigger word, shot and pose, outfit, place, light, framing. It’s the order our Dataset Maker writes, and a fixed order makes captions easy to check. Every caption should still say what the light is and how she’s framed, even in a plain shot.

Part What to write Examples
Trigger word The exact same string every time zvx woman
Shot and pose One phrase: the shot type, her pose and where she looks. An expression, if she has one, follows as its own short part close-up portrait looking up with her chin raised; mirror selfie kneeling on a bedroom carpet; standing with a hand on her hip, pouting
Outfit What she wears, with colours and accessories black cropped t-shirt, gold necklace
Place Where she is plain grey wall, garden with white roses
Light What the light is like soft even light, golden hour, bright sunlight
Framing Last: how much of her is in the shot upper body, cowboy shot, three-quarter body, full body

Close-up portraits are the exception: the shot type already says it (“close-up portrait”), so nothing goes at the end. Here is the example from the Dataset Maker’s README, a close-up with bare shoulders and no outfit part:

zvx woman, close-up portrait looking over her bare shoulder, head turned, plain grey wall, soft even light

Where she looks follows the course’s six angle prompts: looking slightly left, slightly right, down, up, and both side profiles. The order is the same one our prompts use for the scene, so a caption and a prompt for the same shot share their words.

Use the same trigger word everywhere, character for character. The Dataset Maker puts it at the start of every caption. Keep it first in every file, and type the same string into the Trigger Word field of your training job too. How to pick one, and what AI Toolkit does with it, is in our LoRA trigger word guide.

Why do captions leave out hair and eyes when prompts include them?

Training and generating are different jobs. The course’s rule: captions never name her face or hair, so the LoRA ties them to the trigger word, and prompts add one short hair-and-eyes line. The way we think about why: at generation, that line matches what the LoRA learned, so it backs the LoRA up.

In captions. This part is the course’s own reason. If every caption said “long wavy dark brown hair”, the trainer would learn hair as a separate, changeable detail, like the outfit. Prompts that leave it out would then be free to change it. Leave it out of captions and the hair is learned as part of zvx woman.

In prompts. The course’s format puts the hair-and-eyes line right after the trigger: zvx woman, long wavy dark brown hair, middle part, hazel eyes, then the scene. The way we think about it: neon light, wind and pool water all pull on how hair looks in an image, and a line that says what her hair already is helps keep those scene words from restyling it. Either way, it has to match the hair she was trained with. Write a different colour and you’re fighting your own LoRA.

Never in either. Face shape, makeup, skin and body words go nowhere. In a caption they get split off from the trigger. In a prompt they pull the base model toward its own idea of a woman, which is a different woman. The Dataset Maker has a hair_and_eyes field in the format <length> <texture> <colour> hair, <eye colour> eyes, with one optional style detail after the hair, and still writes every caption without it.

Captions (training) Prompts (generating)
Trigger word First First
Hair and eyes Never One short line, right after the trigger
Face, makeup, skin, body Never Never
Shot and pose Yes, right after the trigger Yes, right after the hair-and-eyes line
Outfit, place, light Yes Yes
Framing Yes, last Yes, last before the ending
Ending Nothing extra “candid smartphone photo, natural skin texture”

What does a full set of captions look like?

Here are 10 example captions we wrote for this page in the course format, mixed the way the course mixes framings: a third close-up or upper body, a third cowboy or three-quarter, a third full body. No outfit, place or light is used twice. A real set runs to about 50.

zvx woman, close-up selfie sitting in a parked car, soft smile, white tank top, warm afternoon sun through the window, upper body
zvx woman, close-up portrait looking slightly to the left, grey hoodie, kitchen at home, soft morning window light
zvx woman, sitting in right side profile, laughing, black turtleneck, cafe table by the window, overcast daylight, upper body
zvx woman, leaning on a desk looking up, calm expression, navy blazer over a striped shirt, office with glass walls, cool overhead office light, upper body
zvx woman, standing in three-quarter view, smiling, red satin slip dress, rooftop bar, neon light at night, cowboy shot
zvx woman, leaning on a railing looking down, olive cargo pants and a cropped black tee, city sidewalk, bright midday sun, cowboy shot
zvx woman, walking in left side profile, serious expression, beige trench coat, train platform, flat grey daylight, three-quarter body
zvx woman, standing facing the camera, smiling, light blue sundress and sandals, beach boardwalk, golden hour, full body
zvx woman, walking toward the camera, laughing, sage green workout set and white sneakers, gym floor, bright overhead gym lights, full body
zvx woman, mirror selfie standing by the bed, cream knit set, bedroom, warm lamp light at night, full body

Read them as a set, not one by one. If “white” shows up in half the outfits or “golden hour” in half the light, the LoRA starts treating it as part of her. That’s one of the causes in our LoRA overfitting guide.

The free LoRA dataset planner builds a shot list for a whole set, with a ready caption for each image.

The demo AI character on a pier at dusk, blue long-sleeve top and grey sweatpants, made by the Dataset Maker.
A Dataset Maker image. How we'd caption it: zvx woman, standing on a pier looking away to the left, blue off-shoulder long-sleeve top, grey sweatpants, sea at dusk, overcast light, cowboy shot.

Should you caption by hand or automatically?

Automatically, then check every caption by hand. The Dataset Maker writes a .txt caption in this order for every image: the trigger word, then the template photo’s own caption. Our older course lesson captioned with Gemini in batches of 10. Either way, a model wrote them, so read each one for face, hair and body words before you train.

The fastest check is the free LoRA caption checker. Paste your captions or pick your .txt files, and it marks every line that names hair, eyes, face, skin or body, has the trigger in the wrong place, or misses the framing or light. It runs in your browser, so the files stay on your device.

The checker can’t see your images, so read for these yourself:

  1. The framing matches the image. An auto-captioner may call a cowboy shot “full body”.
  2. The outfit, place and light match the image. A caption that names the wrong outfit teaches the wrong thing.
  3. No outfit, place or light repeats across many captions, unless the images really repeat it, in which case fix the images.

What does caption dropout do?

Caption dropout trains a small share of steps with the caption removed, so the trainer sees some images with no text at all. Our template keeps AI Toolkit’s default, 0.05. AI Toolkit’s example config describes that value as “will drop out the caption 5% of time”.

The value is in AI Toolkit’s new-job defaults and its example config (both checked 2 Oct 2026).

What can go wrong?

  • Some captions name her hair. The auto-captioner slipped “brown hair” into a few files. The LoRA can learn hair as changeable, and it may drift in some images. Search every file for hair and eye words before training.
  • The trigger differs between files. zvx woman in most, Zvx woman or zvx woman in a few. Those images train toward a different token.
  • A caption file doesn’t match its image. image_01.png with img_01.txt trains that image with no caption. Keep names identical.
  • Every caption says the same light. “Golden hour” in 30 of 50 captions means 30 images share it. Either the dataset needs more light variety, or the captions copy-pasted a default.
  • The prompt hair line doesn’t match the training hair. Captions left hair out, so the LoRA learned her real hair. A prompt that says “blonde” fights it. Write the hair she has.
  • Captions describe the body. Body shape comes from the images; the Dataset Maker keeps one body type across a whole set. Body words in captions split it off from her.

Questions people ask

Is a LoRA caption the same as a prompt?
No. Both start with the trigger word and name the shot, outfit, place and light. A prompt adds one short hair-and-eyes line after the trigger and ends with candid smartphone photo, natural skin texture. A caption has neither, and never mentions hair or eyes.
Can AI caption my training images?
Yes. The Dataset Maker writes a caption for every image it makes, and an older course lesson used Gemini in batches of 10. Either way, read every caption and delete any face, hair or body words before you train.
Should the outfit's colour go in the caption?
Yes. The outfit changes from image to image, so name it with its colour: navy blazer over a striped shirt, red satin slip dress. Colour words stay out only for her hair and eyes, which are part of her.

Read next