--- title: "Training Data Management" slug: "clone-training-data-management" updated: 2026-06-22T10:05:52Z published: 2026-06-22T10:05:52Z canonical: "help.hyperscience.ai/clone-training-data-management" --- > ## Documentation Index > Fetch the complete documentation index at: https://help.hyperscience.ai/llms.txt > Use this file to discover all available pages before exploring further. # Training Data Management > [!WARNING] > **Accessing this feature** > > Your access to the feature described in this article depends on your license package and pricing plan. > > To learn which features are available to your organization and how to add more, contact your Hyperscience representative. **Training Data Management** allows you to improve and supervise models by working directly with the training data (“Ground truth”) obtained from each document in the Training Set. You can group documents, see incompatible ones, annotate representative parts of them, and detect potential inconsistencies. - To learn more about ground truth for Identification models, see Step 3 and Step 4 of our [Training an Identification Model](/v42/docs/v43-training-an-identification-model) article. - Learn more about the ground truth of Classification models in [TDM for Classification Models](/v42/docs/v42-tdm-for-classification-models). > [!NOTE] > The performance of your models depends on the quality of the pages, the diversity of the documents, and the consistency of the annotations. For more information on model-training results, see [Evaluating Model Training Results](/v42/docs/evaluating-model-training-results). TDM includes tools for controlling and managing the Identification and Classification models’ performance. Learn more about model performance in [Monitoring Model Performance](/v42/docs/monitoring-model-performance) and [Improving Model Performance](/v42/docs/improving-model-performance). ## TDM for Identification models TDM for Identification models includes the following features: - **Training Data Analysis** — allows you to use grouping for easier and faster annotations. It also provides insights into the training dataset. - Learn more about grouping in [Training Data Analysis](/v42/docs/training-data-analysis). - **Document Eligibility Filtering** — indicates whether a document is eligible for training, based on internal checks in the application and our machine learning logic. It provides additional information about documents that were excluded from the training set. - To learn more, see [Document Eligibility Filtering](/v42/docs/document-eligibility-filtering). - **Training Data Curator** — labels each training document as having high or low importance. The importance is calculated by determining which data would best contribute to the model’s performance. - For more information, see [Training Data Curator](/v42/docs/training-data-curator) - **Labeling Anomaly Detection for Fields and Tables** — identifies potential discrepancies in the training datasets before running model training. Once the annotations are ready, the user can analyze the data to find inconsistencies and ensure a top-performing locator model. - See [Labeling Anomaly Detection](/v42/docs/labeling-anomaly-detection) for more details. - **Search by Text Segment** — search for fields or cells by specific text segments directly within the Training Data Management interface. This feature allows you to locate, review, and annotate data faster during the model-training process. - **Tagging documents** — Starting v42.3, you can organize and manage the training documents more efficiently in TDM by adding tags. - Starting v43, you can tag documents directly from the Annotation experience. Learn more in [TDM for Identification Models](/v42/docs/tdm-for-identification-models). - This feature allows you to add, filter, import, and export tags for documents, making it easier to categorize and find the information you need. It provides the following key capabilities: - **Manual** **tagging** — Hover over the **Tags** cell in the Training Data table to reveal a **+** button. Click it to open the drop-down list with all existing tags. From there, you can select an existing tag or create a new tag. - **Tag** **filtering** — Filter documents in the Training Data table by tag to find relevant items quickly. - **Import** **tags** — If the training data contains tags, they will be automatically imported. - **Special**-**character** handling — Tags cannot contain “;” or spaces (spaces are replaced with underscores). - **Unused** **tags** — Unassigned tags are automatically deleted. > [!NOTE] > Starting in v43.1, tags are case-insensitive > > Tags with the same name but different capitalization are automatically merged into a single tag. This helps prevent duplicate tags and improves consistency when organizing training documents. Learn how to use these features to maximize the performance of your identification model in our [Training an Identification Model](/v42/docs/v43-training-an-identification-model) article. ## TDM for Classification TDM for Classification models allows you to add, remove, and update training pages for Classification models. Learn more in [TDM for Classification](/v42/docs/v42-tdm-for-classification-models). ## TDM for VLM Extraction > [!NOTE] > Available in v42.3 and later Training Data Management (TDM) for VLM Extraction models is where you prepare and manage the data used to train a specialized models on top of the ORCA base model. Learn more in [TDM for ORCA VLMs](/v42/docs/tdm-for-orca-vlms). ## Accessing Training Data Management tools If you have the View Training Data permission (given to System Admin and Business Admin permission groups by default), you can access the Training Data Management tools for a model. Learn more in [Permission Groups](/v42/docs/permission-groups). 1. Go to the **Models** section. Learn more about the models table in [Models Page](/v42/docs/models-page). 2. Click on the specific tab to view the models you need: - **Classification** - **VLM Field Extraction** — Learn more in [TDM for ORCA VLMs](/v42/docs/v43-tdm-for-orca-vlms). - **Identification** - **Text Classification** - **Transcription** ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/models(2).jpg) 1. Click on the name of the model you would like to view training data for: - For ID models: - Click the **Field Identification** or the **Table Identification** tab, depending on the type of training data you would like to view. - The Training Data Management tools are located on the Training Data Health card. - For Classification models: - Click on the **Training Data** tab to edit the documents used for training. ## Continuous Model Training When you import an ID or a Classification model from another instance while **Continuous Field Locator model improvement** and/or **Continuous Classification model improvement** are enabled, the model’s automation rates may decrease. Models only learn from the training data available in their **current instance**. If the new instance contains limited or no training data, the imported model may be replaced by a lower-performing version. To maintain optimal performance: - Train models manually after import. - Keep **Continuous Field Locator model improvement** and **Continuous Classification model improvement** *disabled*, unless specifically advised otherwise by a Hyperscience representative. Manually annotated data used to train our machine learning models. We use a subset of this data to assess the performance of your models A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples. A foundational model that provides core general-purpose capabilities and is not directly trained on customer-specific examples. Use-case specialization is achieved through additional training on top of the base model using customer-specific data.