Free LoRA caption checker
A good LoRA training caption has 6 parts in a fixed order: trigger word, shot and pose, outfit, place, light, and the framing last. It never describes her face, hair, eyes, skin or body. Paste your captions or pick your .txt files, and the checker marks every line that breaks a rule and writes a fixed copy.
The rules are the ones from our guide to how many images you need for LoRA training, and how to caption them.
Results
5 captions: 1 OK, 1 to check, 3 to fix. 3 of 5 start with your exact trigger word.
- Hair words: 1
- Trigger word not first: 1
- Eye words: 1
- Face or makeup words: 1
- Trigger capitals differ: 1
- Odd characters: 1
- Extra spaces: 1
- Too short: 1
- No light described: 1
| Line | Caption | Result and fix |
|---|---|---|
| Line 1 | zvx woman, standing and looking slightly to the left, soft smile, navy blazer over a white top, train station, overcast daylight, upper body | OK |
| Line 2 | zvx woman, walking toward the camera, laughing, long wavy dark brown hair, white linen sundress, beach boardwalk, golden hour, full bodyFixed: zvx woman, walking toward the camera, laughing, white linen sundress, beach boardwalk, golden hour, full body | Fix
|
| Line 3 | close-up portrait, zvx woman with green eyes and freckles, neutral expression, black turtleneck, cafe, soft window lightFixed: zvx woman, close-up portrait, neutral expression, black turtleneck, cafe, soft window light | Fix
|
| Line 4 | Zvx woman, standing in three-quarter view, smiling, red satin slip dress, at a friend’s rooftop bar, neon light at night, cowboy shot Fixed: zvx woman, standing in three-quarter view, smiling, red satin slip dress, at a friend's rooftop bar, neon light at night, cowboy shot | Fix
|
| Line 5 | zvx woman, mirror selfie, gym | Check
|
Fixed captions
Trigger word first and exact; hair, eye, face and skin words removed; odd characters and spaces cleaned. Body, ethnicity and age words stay for you to check by hand. Read it before you use it.
zvx woman, standing and looking slightly to the left, soft smile, navy blazer over a white top, train station, overcast daylight, upper body zvx woman, walking toward the camera, laughing, white linen sundress, beach boardwalk, golden hour, full body zvx woman, close-up portrait, neutral expression, black turtleneck, cafe, soft window light zvx woman, standing in three-quarter view, smiling, red satin slip dress, at a friend's rooftop bar, neon light at night, cowboy shot zvx woman, mirror selfie, gym
What does the LoRA caption checker check?
Nine things on every line: the exact trigger word at the start, words about her hair and eyes, her face and skin, her body, age or ethnicity, duplicates, length, a framing word, light, and odd characters or stray spaces. It also checks the trigger word itself. Rule breaks are marked Fix; judgement calls are marked Check.
| Check | What it flags | In the fixed version |
|---|---|---|
| Trigger word | Missing, not first, different capitals or spacing | Put first, spelled exactly |
| Hair and eyes | “long wavy brown hair”, “blonde”, “ponytail”, “green eyes” | Removed |
| Face and skin | “freckles”, “full lips”, makeup, “pale skin” | Removed |
| Body, age, ethnicity | “curvy”, “25-year-old”, “asian woman”; anything that suggests a minor is always a Fix | Left for you: some of these words describe clothes or places |
| Duplicates | The same caption twice | Left for you |
| Length | Under 8 words or over 50 | Left for you |
| Framing | No framing word, or framing stuck in the middle (tips, not errors) | Left for you |
| Light | No light described (a tip, not an error) | Left for you |
| Characters and spaces | Curly quotes, emoji, tabs, underscores, trailing or double spaces, stray commas | Cleaned |
Framing words it knows: close-up, upper body, cowboy shot, three-quarter body, full body, plus selfie and mirror selfie. The framing goes last, or first when it’s part of the shot, as in “close-up portrait looking up” or “mirror selfie in a bedroom”. Both are what our Dataset Maker writes. A caption with several lines (common in files from an AI captioner) is joined into one line in the fixed version.
Why leave her face and hair out of LoRA training captions?
Because the trainer ties whatever the captions don’t mention to the trigger word. Leave out her face, hair and eye colour, and they become part of “zvx woman”. Write “long brown hair” in every caption, and the LoRA learns hair as a separate, changeable detail, so it can drift from image to image.
A caption that passes every check:
zvx woman, standing and looking slightly to the left, soft smile, navy blazer over a white top, train station, overcast daylight, upper bodyOur Dataset Maker writes one .txt caption per image in this order when it builds a 50-image datasetfrom one face: the trigger word, then the template photo’s own caption. In an older lesson the captions came from Gemini, in batches of 10. Whichever captioner you use, read what it gives you: an AI captioner usually describes hair and eye colour unless it’s told not to, and that’s the first thing this checker catches.

Why do prompts get hair and eyes when captions don’t?
Captions teach; prompts ask. When you train, leaving hair and eyes out ties them to the trigger word. When you generate, our videos add one short hair-and-eyes line after the trigger word, which keeps those two details steady from image to image. Her face shape, makeup, skin and body stay out of both.
The same scene as a training caption, then as a prompt:
zvx woman, standing and looking slightly to the left, soft smile, navy blazer over a white top, train station, overcast daylight, upper bodyzvx woman, long wavy dark brown hair, middle part, hazel eyes, standing and looking slightly to the left, soft smile, navy blazer over a white top, train station, overcast daylight, upper body, candid smartphone photo, natural skin textureSo don’t run your prompts through this checker: it will flag the hair-and-eyes line, and that line belongs there. The freeAI influencer prompt generator builds prompts in the right format.
What makes a good LoRA trigger word?
A made-up token the model has never seen, plus one plain word for what she is. Our default is “zvx woman”. The token carries no old meaning, so nothing leaks into her face, and “woman” tells the model what it’s looking at. Use the exact same spelling in every caption, in training and in every prompt. More on picking one, and what AI Toolkit does with it, in how to pick and use a LoRA trigger word.
- Not a real word or name. “Emma” or “velvet” already mean something to the model, and that meaning pulls on her face.
- Lowercase, one space, no commas. A comma splits it in two. Capitals are one more thing to mistype across 50 files.
- Short. Three or four letters is plenty. You’ll type it hundreds of times.
- “woman”, not “girl”. Characters on this site are always clearly adult, and the trigger word should say so.
How do you fix captions the checker flags?
Fix the trigger word first, because it touches every line, then work down the Fix rows and judge the Check rows yourself. The fixed version handles the mechanical part: trigger first, hair, eye, face and skin words removed, characters cleaned. You still read it before training.
- Type your trigger word exactly as you set it for training. In the Dataset Maker that’s the trigger_word field.
- Paste your captions, one per line, or pick all the .txt files from your dataset folder.
- Read the trigger word notes. A weak trigger is worth fixing before anything else.
- Go through every Fix row, then decide on each Check row: “petite” might be her body or a clothing size.
- Copy or download the fixed captions. With files, the download is one .txt with each caption under its file name: paste each one back into its own file.
- Keep each caption file named like its image (image_01.png and image_01.txt), then run the check again until it’s clean.
What can go wrong?
Most failures come from one bad file in a set of 50, or from trusting a word list too far. Here are the usual ones, and what to do about each, before you spend GPU time and money on a training run.
- One file has a typo in the trigger word. That caption doesn’t match the rest, so that image teaches the trigger word less. The checker names the line or file.
- The captioner described her hair in half the files. Her hair starts changing between images. Strip it from every caption, not just most.
- The checker can’t see your images. It can’t tell whether image_01.txt describes image_01.png, or whether the outfit is right. Read a few captions against their photos.
- It’s a word list, not a reader. It may flag a word used another way, or miss a description phrased oddly. That’s why body, age and ethnicity words are Check, not Fix.
- A fixed caption lost a whole part. A part that only describes her hair, like “long wavy brown hair”, is dropped whole. A pose like “hand in her hair” stays, because it doesn’t say what her hair looks like. If a fix leaves the caption short, describe a new pose.
- Hidden characters from copy and paste. Word processors and chat apps add curly quotes and non-breaking spaces. The fixed version swaps them for plain ones.
Questions people ask
Should LoRA captions include the trigger word?
Yes. It goes first in every caption, spelled exactly the same, because it’s the word the LoRA learns her under. Then you start every prompt with it. One capital letter or a missing space, and that caption no longer matches the rest.
Should I describe hair colour in LoRA captions?
Not for a character LoRA. Hair, eyes and face stay out of the captions, so the LoRA ties them to the trigger word. Hair and eyes go in your prompts instead, as one short line right after the trigger word.
Are my caption files uploaded anywhere?
No. The checker runs in your browser. It reads the .txt files you pick on your own device, and nothing is sent anywhere. Once the page has loaded, it works offline.
How long should a LoRA caption be?
Long enough for every part and no longer. Our example caption is 23 words. The checker flags captions under 8 words, which usually skip parts, and over 50, which usually means piled-up adjectives. Those two limits are the checker’s guard rails, not a trainer rule.
Why caption a character LoRA at all?
Captions tell the trainer what can change. Name the outfit, place and light in every caption, and the LoRA learns them as details you can swap in a prompt. What the captions never name, her face, ends up in the trigger word.
Read next
- How many images do you need to train a character LoRA?
A character LoRA needs 20 to 50 images. We train on 50. What they should show, how to caption them, how to pick a trigger word, and what goes wrong.
- Free AI influencer prompt generator
Pick outfit, location, pose, light and camera style and get a ready-to-copy AI influencer prompt in the format that works with a character LoRA. Free.
- How to keep your AI character's face the same in every photo
Prompts, reference images, a character LoRA, or a LoRA trained on a dataset you generate. What each holds, what it costs, and when you need a LoRA.