TDM for Identification Models

Prev Next

In this article, you’ll learn how to navigate through and use Training Data Management for Identification models. Learn more about each feature in Training Data Management. To learn more about Identification models, see Training an Identification Model.

Accessing this feature

Your access to the feature described in this article depends on your license package and pricing plan.

To learn which features are available to your organization and how to add more, contact your Hyperscience representative.

Accessing TDM for Identification models

Each Identification model trained for the specific Semi-structured Layout has its own tab on the Model Management Page (i.e., Field Identification or Table Identification).

  • You can access Layout Management from the Model Management page, as shown in the image below:

Navigating to a layout's Model Management page from the Models section

In the Models section, go to the Identification tab and click the name of your layout to access its Model Management page.

The Actions menu in the page’s upper-right corner allows you to:

  • Train new model

  • Upload model

  • Upload training documents

  • Download training documents

The Actions menu on the Model Management page

Importing duplicate models

In v43.2 and later, Hyperscience rejects an imported Identification model when its UUID already exists. The error message identifies the duplicate UUID and indicates that the existing model may be in the Archive view.

Model Summary card

The Model Summary Card displays the Live and Candidate models. You can see the following information:

Field

Description

Status

The current status of the model:

  • Live — The model is deployed.

  • Trained — You have a candidate model ready to be deployed.

  • Uploaded — The model was uploaded.

  • Failed — The model training has failed.

  • Training — The model training is still in progress.

  • Reqs not met — The requirements for training a model are not met. Either you don’t have enough training documents, or you haven’t run Training Data Analysis on the documents you have. Learn more in Requirements for Training a New Model.

  • Analysis runningTraining Data Analysis is in progress.

  • Ready to train — Both requirements are met, and you can start a training.

  • Queued — The model is scheduled for training in the queue.

Projected Automation

Displays the predicted Automation based on the Test Target Accuracy. Learn more in our Evaluating Model Training Results article.

Projected Cell Automation

Displays the projected automation per cell, based on the Test Target Accuracy.

Test Target Accuracy

The accuracy percentage used to calculate the projected automation. You can adjust it from the Model History card.

Documents Trained

The number of documents used for training the model.

Fields Layout / Trained or Columns Table / Trained

The number of fields or columns in the current live version of your layout, and the number of fields or columns used for training the model.

  • For example, if you train a model with 5 fields or columns but remove a field or a Column from the current Layout Version, the numbers will be 4 / 5 (i.e., the layout has 4 fields or columns, but the model was trained on 5).

Trained

Date the model was trained.

Deployed

Date the model was deployed.

The Model Summary card showing the live model's details

If the requirements for training a model are not met, then the model summary card will display the status “Reqs not met,” and a View training data button will appear on the card. It will redirect you to the Training Data table.

Two requirements must be met before training

The status shows “Reqs not met” until both of the following are true: you have at least the required number of training documents, and a Training Data Analysis has completed on them. A model with enough documents still shows “Reqs not met” if the analysis hasn’t been run, so check the Training Data Health card if the count looks sufficient.

The Model Summary card showing the Reqs not met status and the View training data button

If your model is ready for training, the status “Ready to train” will appear on the model summary card. You’ll be able to start training by clicking the Train button located on the right-hand side of the card.

The Model Summary card showing the Ready to train status and the Train button

If you want to cancel your training, you can click the Cancel training job button next to the status of your model in the model summary card.

The Model Summary card showing a queued training job and the Cancel training job button

Projected Automation chart

The Projected Automation chart displays the performance of the model that’s currently live. Expand it by clicking the arrow.

The chart displays how the target accuracy affects the automation. The lower the accuracy, the higher the automation, and vice versa.

Note that projected model performance (i.e., accuracy and automation) can increase by adding more QA records. You can also see the Margin of Error (MoE) for this model.

Margin of Error (MoE)

The Margin of Error (MoE) indicates the allowable range of inaccuracy in the system's results. It shows you how much the output can differ from the true value while still being acceptable. A smaller margin of error means the system is more accurate.

  • Adjust the Target Accuracy percentage by clicking the up and down arrows.

The Projected Automation chart with the Target Accuracy control

The chart will display the projected automation of your model with the target accuracy you specify. You can determine the target accuracy value that best meets your needs by entering test values.

Identification Report

The Identification Report displays the number of identified fields (whether the machine or a human identified them), their accuracy, and the field-level automation (i.e., the automation of the fields the model was trained on).

The Identification Report is available only for Field Identification models.

Select a specific date range for the report to see charts for the total number of identified fields (machine-identified and human-identified) and their respective accuracy values.

  • You can also see the Margin of Error, the calculation points, and the Automation Rate for the selected period. Hover over the Accuracy value to see the MoE and the Calculation points.

The Identification Report showing accuracy with the Margin of Error and calculation points

Calculation points are the total fields used to calculate accuracy. They represent the number of evaluated fields. For instance, an accuracy of 50% could come from 1/2 or 400/800 evaluated fields. Learn more in our Accuracy article.

Fields Identified chart

The Fields Identified chart displays the number of machine- and manually-identified entries for a specific period.

  • You can filter the Identification Entries by clicking the Total Fields, Manual, and Machine buttons located at the top of the chart, or Machine Field ID or Manual Field ID located below the chart.

  • Hover over the chart’s data to see the specific date and the number of fields identified by the machine or by a human.

  • Download the data as a CSV file by clicking the Download CSV button located on the right-hand side of each chart.

The Fields Identified chart with machine and manual filters

Field Identification Accuracy chart

The Field Identification Accuracy chart displays the percent accuracy for the selected time. You can see:

  • Field Identification Accuracy

  • Manual Accuracy

  • Target Accuracy

Field / Table Level Automation

The Field / Table Level Automation card displays the automation percentage of the fields or columns your model was trained on:

Field

Description

Field Name

Displays the name of the field in your layout

Machine Identified or Machine Identified Cells

Indicates the number of values identified by the machine for that field or for that column

Total or Total Cells

The total number of values identified for this field or column

Automation Rate

The percentage of automation for the specific field

  • You can sort each column and select the number of rows to display on the page.

The Field / Table Level Automation card listing per-field automation rates

Training Data Health Card

The Training Data Health card displays a breakdown of your Dataset. It shows the following insights on the uploaded documents:

  • The number of additional documents recommended for training. These documents do not include the minimum required for training to begin. Our recommendation is 100 documents.

  • Required — the number of documents required for training a model. The default number is 100. If you would like to change it, contact your Hyperscience representative.

  • Number of ineligible/eligible documents:

    • Ineligible documents — documents that do not meet the criteria for processing

    • Eligible documents — documents that meet the required criteria and can be used for model training

  • Number of added or removed documents since the last training data analysis

  • Number of Groups discovered during the training data analysis

  • Number of documents with potential anomalies

Learn more about eligibility in Document Eligibility Filtering.

The Training Data Health card showing document counts, groups, and anomalies

  • Analyze your data by clicking the Reanalyze data button.

Learn how to optimize your data and achieve better model performance by using Training Data Analysis, as described in Step 4 of Training a Semi-structured Model.

Training Data Table

The Training Data table allows you to review, organize, download and manage the documents used to train your model.

Filter the table’s contents by:

  • Status — training document status

  • Date Modified — set a specific date range to view the training documents

  • Group — filter by groups, formed after Training data analysis.

  • Training Eligibility — filter by document eligibility. To learn more, see Document Eligibility Filtering.

  • Importance — high or low importance documents. Learn more in Training Data Curator.

  • Has Anomalies — filter by anomalies detected. Learn more in Labeling Anomaly Detection.

  • Tags — you can use tags to organize and filter training documents. Learn more below.

Search and filtering

You can search for documents by their IDs or their file names.

The Group filter supports multi-select, making it easier to review and compare documents across several groups. Additionally, you can filter by multiple statuses to make the review process more flexible and efficient.

The filters panel above the Training Data table

Selecting at least one training document allows you to use the Actions drop-down menu. This drop-down menu has the following buttons:

  • Edit training status — change the training status of the selected training documents that have been annotated. All unannotated training documents that you’ve selected will keep their current status.

  • Edit tags — add or remove tags on the selected training documents.

  • Export — starting in v43, you can export the Identification model’s training data to a CSV file. This feature lets you download your dataset in bulk and reuse it in another instance or analyze it outside the system.

    • The export includes all relevant table columns and a link to each document’s ground truth item.

    • You can export the entire training data table or a selected subset of rows.

  • Delete rows

The menu next to the Actions drop-down list contains the Manage Tags option.

CSV exports in v43.2 and later

Hyperscience confirms that the CSV export is being prepared. When the export is ready, a download link is available in the Notifications menu.

A CSV export option is unavailable when the export would include documents that are still loading. You can export selected rows after those documents finish loading. To export all training data, wait for every document to finish loading. Hover over an unavailable option to see why it is unavailable.

The Actions drop-down menu above the Training Data table

The Training Data table contains the following columns:

Column

Description

ID

The ID number of your document

File Name

The original document name. Allows you to manage training data faster. Search for a document by its original name in the search bar.

  • Document names can come from submissions or from TDM uploads.

  • For older documents, names are only available if they were included in the submissions the documents were part of. If the system doesn’t have that information, the document names won’t be shown.

  • When importing older training data, the name will appear as “N/A,” since document names aren’t included in training data exports.

Group

The ID number of the document’s group. This ID is available after training analysis.

Priority

“High” or “Low,” depending on the results from the training data analysis

Pages

The number of pages in the specific document

Modified

The last date changes were made in the document

Status

The training status of your document:

  • Always — The document will always be used to train the model.

  • Auto — The document will be used to train the model until its scheduled deletion, based on the PII Data Deletion settings for your instance.

  • Never — The document will never be used in future model training, even if it hasn’t been deleted as part of the PII Data Deletion.

  • Loading — The document is currently being pre-processed and prepared for Annotation.

  • Ready to annotate — The document has been uploaded successfully, but has not been annotated yet.

Deletion

The time remaining until the document is scheduled for deletion, for example 3 months. It shows Today when the document is scheduled for deletion within the next 24 hours, and Never when no deletion is scheduled.

Tags

Labels used to organize training documents. Tags can be added manually and used to filter the Training Data table. See Working with tags below.

Actions

Each document has its own Actions link:

  • Annotate — available for documents that do not have any annotations.

  • Edit Annotations — available for documents that have already been annotated.

Learn how to annotate your documents to achieve a top-performing model in Training an Identification Model.

Working with tags

You can assign tags directly from the Training Data table by hovering over the Tags cell for a document and selecting an existing tag or creating a new one. Starting in v43.1, tags are case-insensitive. Tags that differ only by capitalization are automatically treated as the same tag and normalized to a single value. For example, TAG, Tag, tag, and taG are all recognized as the same tag, and each one is displayed as TAG.

Behavior

Description

Manual tagging

Hover over the Tags cell to reveal a + button. Use it to select an existing tag or create a new one.

Tag filtering

Filter documents in the Training Data table by tag to find relevant items quickly.

Tag import

Tags included in imported training data are automatically imported.

Tag naming rules

Tags cannot contain semicolons (;) or spaces. Both are replaced with underscores (_).

Tag scope

Tags are shared across models, so the same tag can apply to documents in more than one model.

Managing tags

To delete tags, click Manage Tags in the menu next to the Actions drop-down list. Search for the tags you want to remove, select them, and then click Delete. Deleting a tag removes it from every document and model that uses it. Click Undo in the message that appears to cancel the deletion.

Starting in v43.2, long tag names are shortened

Long tag names are shortened with an ellipsis so they do not obscure other columns in the Training Data table. Hover over a shortened tag to view its complete name.

Adding a tag to a document from the Tags cell in the Training Data table

Model History table

The Model History table, located at the bottom of the Model Management page, provides a comprehensive overview of your model's lifecycle. It displays the following columns:

Column

Description

Name

The name of the model version.

UUID

The model version’s unique identifier. This column is hidden by default. See View and copy a model UUID below.

Date Created

Date and time the model was created. Helps in tracking the model’s version history and ensures you’re working with the most recent model version.

Version

The specific version of the model that was trained on.

Source

Indicates where the model was trained—either within the current instance or externally and then uploaded to this instance.

Proj auto

Displays the predicted automation based on the Test Target Accuracy. Learn more in our Evaluating Model Training Results article.

Fields Layout / trained

Displays the number of fields in the current live layout vs. the number of fields the model was trained on. Numbers in parentheses show the difference between them.

Docs Trained

The total number of documents used for training the model.

Last Deploy

The last date and hour the model was deployed.

Actions

The options in this menu allow you to take the following actions on a model version:

  • Deploy

  • Undeploy

  • Reject Candidate

  • Download

  • Archive

Find specific records in the table in the following ways:

  • Filtering — Filter the contents of the Model History table by creation date, last-deploy date, source, and Trainer version. Click Filter and select the criteria that match what you’re looking for.

  • Searching — Search for a model version by entering a name in the search box.

  • Sorting — Sort the table's contents by clicking on the names of the Name, Date created, Version, Source, and Last deploy columns.

Additionally, you can choose which columns are included in the table by clicking the menu next to the Filter drop-down list and clicking the Manage columns option.

Archive and unarchive model versions

In v43.2 and later, you can archive Identification model versions that you no longer actively use. Archived models are removed from the active Model History view but remain available in the Archive view. Undeploy a live model or reject a candidate model before archiving it.

To archive a model version:

  1. In the Model History table, find the model version.

  2. Open the model version’s options menu, and then select Archive.

  3. Confirm the action in the message that appears.

To manage archived model versions, open the table options menu and select View Archive. To restore a model version, open its options menu and select Unarchive. Select View Active to return to active models.

Permanently deleting an archived model

You can permanently delete an archived model version. This action cannot be undone.

View and copy a model UUID

In v43.2 and later, the UUID column contains each model version’s unique identifier. The table shows a shortened value. Hover over it to view the complete UUID, or select it to copy the complete value to your clipboard.

The UUID column is hidden by default. To display it:

  1. In the Model History table, open the menu next to Filter.

  2. Select Manage columns.

  3. Select UUID, and then select Save.

To train an Identification model using TDM, see Training an Identification Model.