A dataset can look correct in an annotation tool and still arrive incorrectly in a training pipeline. The boxes may be attached to the right objects, but an export or conversion step can change how their coordinates, class numbers, or image references are interpreted.
Annotation export validation checks that boundary. It asks a narrower question than annotation quality assurance: did the approved labels survive the handoff with their meaning intact? This guide focuses on bounding-box detection datasets, with a small illustrative example rather than a claimed project result.
Start with the receiving pipeline
Before choosing an export option, write down what the training loader expects. A file extension is not a complete specification. Record the coordinate order, whether values are pixels or normalized fractions, the class-ID mapping, and which image files the coordinates refer to.
Also agree on how the delivery distinguishes reviewed images with no target objects from images that were never reviewed. Both might have no boxes, but they do not represent the same annotation decision.
Keep the original export. Run conversion on a separate copy so a failed check can be traced back to the source without reconstructing it from the converted files.
Four numbers can describe different boxes
Torchvision's box conversion documentation distinguishes corner coordinates, top-left plus width and height, and center plus width and height. These representations are not interchangeable. Renaming a file does not convert its geometry.
For example, the Ultralytics YOLO detection format uses a class number followed by normalized center coordinates, width, and height. Horizontal values are divided by image width; vertical values by image height. Its class numbering starts at zero.
Image: 1000 x 500 pixels. Green box: top-left (100, 50), bottom-right (300, 150). Blue dot: center (200, 100).
[100, 50, 300, 150][0.2, 0.2, 0.2, 0.2]Illustrative geometry only. No customer image, prediction, or measured result is shown.
In this example, width is 300 minus 100: 200 pixels. Height is 150 minus 50: 100 pixels. The center is (200, 100). Dividing the horizontal quantities by 1000 and the vertical quantities by 500 produces the four normalized values shown above.
If a converter instead treats the top-left corner as the center, the numbers can still be within the allowed range. A range check alone would not reveal the misplaced box. Compare a rendered overlay with the source annotation to check the interpretation.
Check identities, not just class counts
Suppose an export assigns class 0 to a carton and class 1 to a pouch. If the receiving configuration lists those names in the opposite order, every ID is valid but each object is interpreted as the wrong category. This is an illustrative failure, not a report about an N-Base dataset.
Compare the source label name, exported ID, and receiving label name as a three-way mapping. Do not rebuild the mapping by independently sorting names in two different tools. Save the exact mapping alongside the delivery.
Counts remain useful: compare instances per class before and after conversion. Investigate every difference, including deliberate filtering. Equal totals are a useful check, but they do not prove that individual labels kept the correct identity.
Do not confuse an empty image with missing work
Ultralytics documents that an image with no objects does not require a label text file. That convention makes a separate completion record important: the absence of a file cannot, by itself, tell you whether the image was approved as a negative or its labels were lost.
Use a manifest to record each expected image and its review status. Reconcile reviewed negatives, images with annotations, exclusions, and unresolved items against the delivery. Avoid a cleanup rule that removes every image without a label file; it could remove intentional negative examples.
The acceptance decision belongs in that record, not in a guess made by the converter. This also helps keep incomplete work out of a training release.
Verify the image coordinate frame
Check the dimensions of the decoded image that the loader actually uses. If labels were created for one size but attached to a resized copy, pixel coordinates need the corresponding transformation. A crop or rotation also changes the relationship between the image and its boxes.
Inspect image orientation handling at the boundary between tools. Keep a few deliberately chosen fixtures, including a non-square image and a box close to an edge. These make swapped dimensions and incorrect transforms easier to spot than a centered box on a square image.
Do not silently clip every out-of-bounds box to make validation pass. First determine whether the source convention permits it, a transform is incorrect, or the label itself needs review. Any approved correction should be recorded.
A compact export acceptance checklist
Use automated checks for structural properties and visual inspection for interpretation. Neither replaces the other. Define tolerances for rounding before comparing converted coordinates.
| Check | Compare | Investigate before acceptance |
|---|---|---|
| Image references | Manifest entries against decoded files | Missing, duplicate, unreadable, or incorrectly paired images |
| Class identity | Source name, export ID, receiving name | Unknown IDs or changed class meanings |
| Box geometry | Coordinate order, units, dimensions, overlays | Non-finite values, invalid sizes, unexpected shifts |
| Completeness | Per-image and per-class counts, approved negatives | Unexplained losses or additions |
| Preserved fields | Required attributes against destination support | Discarded attributes that downstream decisions need |
| Split membership | Exported image IDs against the agreed split lists | Files moved, duplicated, or omitted during packaging |
| Loader behavior | Loaded samples against the accepted export | Silent filtering, misread categories, incorrect transforms |
Finish with the actual training loader
Render a small selection of loaded samples with random augmentation disabled. Include each class and the difficult conversion fixtures, not just a convenient first batch. Check labels after loading, rather than trusting only the annotation tool's preview.
A round-trip test can help: convert a known box to the destination format and back, then compare it with the starting coordinates. But a converter and its inverse can share the same mistaken assumption. Keep independently calculated fixtures, such as the example above, and check the receiving loader too.
Package the original export, converted labels, class mapping, image manifest, converter version, validation results, and documented exclusions together. Your dataset release process can then identify exactly which files were accepted.
Passing these checks does not establish annotation accuracy or model performance. It establishes something more specific: the receiving pipeline can interpret the delivered labels as intended. That is worth confirming before spending time on training.
Frequently asked questions
Does a valid annotation file mean the dataset is ready for training?
No. A file can be structurally valid while using the wrong class mapping or coordinate convention. Validate its interpretation in the receiving loader and assess annotation quality separately.
Should images without boxes be removed?
Not automatically. They may be reviewed negative examples or unfinished work. Use the agreed review manifest and the receiving format's rules to distinguish them.
Is a round-trip conversion test enough?
No. Matching forward and reverse mistakes can cancel each other out. Add independently calculated test cases and inspect samples produced by the actual training loader.
Planning an annotation handoff?
Bring the target format, class mapping, and acceptance requirements into the discussion before production begins.