IMAGE DATASET CHECKLIST

A better dataset starts before labeling.

A practical checklist for teams preparing their first computer vision dataset. Focus on coverage, clear labels, and an evaluation set you can trust.

Audit coverage and collection conditions.

List the situations the model will encounter: lighting, viewpoints, object sizes, devices, and uncommon cases. Compare that list with your image collection. Thousands of nearly identical images may cover fewer useful situations than a smaller, deliberately varied sample.

  • Record source and capture conditions where available.
  • Check image quality and remove unusable files.
  • Identify missing classes and rare conditions before labeling at scale.

Write the annotation rules first.

Give reviewers class definitions, examples, and a way to flag uncertainty. Label a small common batch, compare interpretations, and refine the rules before expanding the task. Disagreement is useful evidence that a class or boundary needs clarification.

  • Include edge cases and examples of what not to label.
  • Specify rules for overlap, truncation, and empty images.
  • Keep ambiguous cases available for expert review.

Protect the evaluation set.

Separate training, validation, and final evaluation data before producing augmented versions. Keep variations and near-duplicates of the same source image in the same split. For video frames or production batches, split by sequence or batch when that better represents future use.

  • Keep the final test set out of model and threshold selection.
  • Evaluate synthetic enrichment on held-out real images.
  • Document the split so experiments remain comparable.

Use model errors to plan the next collection.

After a baseline run, group errors by condition and class. Collect or annotate examples that address a demonstrated weakness. Change one major factor at a time where practical, then compare on the same validation set. Repeatedly tuning to a final test set turns it into another validation set.

  • Inspect errors instead of collecting more images indiscriminately.
  • Track dataset changes alongside experiment results.
  • Refresh evaluation coverage when deployment conditions change.
A CLOSER LOOK

Should synthetic images replace real evaluation data?

No. Synthetic variation may help explore specific gaps in training data, but it can also introduce unrealistic features. Review generated examples and evaluate their contribution on real images representative of the intended application.

YOUR NEXT STEP

Put it into practice.

Discuss your use case