AI Training Data Resource

AI Training Data Glossary

Clear definitions of the terms enterprise teams use when planning data annotation, AI training datasets, human-in-the-loop review and model quality programs.

Core AI Data Terms

01

Active learning

A training strategy where uncertain or high-value examples are routed for human review so annotation effort is focused where it can improve model performance most.

02

Annotation guideline

A written instruction set that defines labels, edge cases, examples, counterexamples and acceptance rules for reviewers.

04

Classification

Assigning a category or label to an item such as an image, document, utterance, product or content item.

05

Computer vision

AI systems that interpret visual data such as images, video, medical scans, retail shelves, roads, products and industrial scenes.

06

Data curation

Selecting, cleaning, organizing and preparing datasets so they represent the model objective and production environment.

07

Data labeling

Assigning structured labels to data so a machine learning model can learn from examples.

08

Dataset validation

Checking training data for label quality, completeness, consistency, class balance and suitability for model training or evaluation.

09

Ground truth

The validated reference label or answer used to train, test or evaluate an AI model.

10

Human-in-the-loop

A workflow where human reviewers guide, validate, audit or correct AI outputs. It is central to high-risk, ambiguous and quality-sensitive AI systems.

11

Instance segmentation

A pixel-level annotation method that separates each object instance in an image, even when objects belong to the same class.

12

Keypoint annotation

Marking specific points on an object or body, often used for pose estimation, product landmarking and visual measurement tasks.

13

LiDAR annotation

Labeling 3D point cloud data for autonomous vehicles, robotics, mapping and spatial AI. See LiDAR annotation services.

14

Named entity recognition

An NLP annotation method that identifies entities such as people, companies, products, dates, locations and domain-specific terms in text.

15

OCR annotation

Preparing text extraction, document labeling and field validation data for optical character recognition and document AI systems.

16

Polygon annotation

A visual annotation method that traces object boundaries more precisely than a bounding box.

17

RLHF

Reinforcement learning from human feedback, where human preferences or evaluations help improve model behavior, especially in LLM workflows.

18

Semantic segmentation

Pixel-level labeling where each pixel is assigned to a class such as road, vehicle, product, tumor, shelf or background.

19

Synthetic data

Artificially generated training data used to supplement real datasets, stress-test models or fill data gaps.

20

Training dataset

The collection of labeled examples used to train or fine-tune a machine learning model.

Need Help Planning Training Data?

Northern Base AI Labs helps enterprise teams scope, annotate, validate and improve AI training datasets.

Contact Us