TDM for ORCA VLMs

Prev Next

Training Data Management (TDM) allows you to train and improve models by working directly with the training data, or Ground truth, obtained from each document in the Training Set.

For ORCA VLMs, use TDM to:

  • prepare and annotate documents,

  • train a specialized model on top of an ORCA base model,

  • evaluate the resulting candidate,

  • and manage the model versions associated with your use case.

TDM is where you add representative examples and confirm the correct field values that the specialized model should learn from. Training creates a candidate model without changing the base model; after you evaluate and deploy the candidate, it becomes the live specialized model used for document processing, depending on your flow settings. In this article, you’ll learn how to prepare and train a specialized model in TDM for ORCA VLMs.

Prerequisites

Before you begin, make sure that:

  • You have the View Training Data permission. Actions that add, annotate, or change training data also require the applicable edit permissions. See Permission Groups.

  • A locked Semi-structured Layout with fields and notes is associated with the flow that uses ORCA.

  • A Model Definition exists for your scope, task, and model type. The model definition links the layout with its training data and model versions. See Model Definitions for more information.

Base model requirement

  • You can prepare and manage training data before installing a base model.

  • To train a specialized model, at least one compatible ORCA base model must be installed.

  • In v43.2 and later, the ORCA 2 base model is available in addition to ORCA 1.0. Learn more in ORCA Vision Language Models and Installing ORCA VLMs.

Access TDM for ORCA VLMs

To access TDM for ORCA VLMs:

  1. Go to Models.

  2. Open the VLM Field Extraction tab.

  3. Select the model definition associated with your layout.

The model details page contains three tabs:

Tab

Use it to

Overview

Check the live or candidate models, training readiness, latest base model, training data health, and Projected Automation.

Training Data

Add, find, annotate, tag, export, and manage the documents used as ground truth.

History

Review base models or specialized model versions, copy UUIDs, deploy compatible versions, and manage archived models.

The page-level Actions menu contains model-definition actions such as Deploy candidate model, Reject candidate model, Train model, Upload Documents, Export, and Import. The available actions depend on the model state, your permissions, and whether training requirements are met.

TDM for VLM Extraction tabs

Do not confuse the page-level menu with the Actions menu in the Training Data table, which applies to selected training documents.

Opening a model definition for the first time

When you open a model definition, you land on the Overview tab. Before you add training data, it shows only the Live base model summary, labeled Pre-trained, and a Training summary of Reqs not met. This is expected — start in the Training Data tab to upload and annotate documents. The Overview tab fills in as your models training progress.

Specialization process

Training a specialized model follows five stages:

  • Prepare training data,

  • Annotate training data,

  • Train the model,

  • Evaluate the candidate,

  • and deploy the candidate.

The sections below describe where each stage happens in TDM. For the end-to-end process, including dataset guidance and validating a Deployed Model, see Training a Specialized Model.

Training Data tab

The Training Data tab lists the documents associated with the model definition. Use this tab to build a representative dataset and maintain the training data for future training runs. Learn how to upload and manage training data by following the steps described in the next sections or see the interactive walkthrough below.

Upload documents

  1. Open the page-level Actions menu.

  2. Select Upload Documents.

  3. Add the documents you want to use as training examples and complete the upload.

  4. Wait for document loading to finish before annotating or exporting those documents.

Page-level actions in TDM for VLM Extraction

For guidance on choosing representative documents and the dataset requirements (page limits, uniqueness, and minimum counts), see Training a Specialized Model.

Training data table

Find and filter training data

Once uploaded, the training data can be filtered by Usage Rule, Source, Sched. deletion, or Tags.

Training data table filters

The Usage rule indicates how the system uses each document for training:

Value

Description

Auto

The system uses the document to train the model until it is scheduled for deletion, based on the PII Data Deletion settings for your instance.

Always

The system always uses the document to train the model.

Never

The system never uses the document in future model training, even if it has not been deleted as part of the PII Data Deletion.

Loading

The system is pre-processing the document and preparing it for annotation.

Ready to annotate

The document has been uploaded successfully, but has not been annotated yet. The system does not include it in training until you annotate it.

Open the table-options menu next to Filter to manage tags or choose which columns appear.

Additional options - manage tags and manage columns in the Training Data table

Tag training documents

Tags help you categorize, filter, and find training documents:

  • To tag one document, hover over its Tags cell and select +.

  • To tag multiple documents, select them and choose Actions > Edit tags.

  • To create, rename, or delete available tags, open the table-options menu and select Manage tags....

The system automatically adds tags that are included in imported training data. Tag names cannot contain spaces or semicolons. The system replaces both with underscores. The system also automatically deletes unassigned tags. Tags are case-insensitive, and the system normalizes them to uppercase.

In v43.2 and later, the Tags column shortens long tag names with an ellipsis.

Hover over a shortened tag to see its full name.

Tags

Export training data as CSV

  1. To export a subset of the training data, select its rows in the Training Data table.

    • To export the entire table, do not select individual rows.

  2. Open the Training Data table’s Actions menu.

  3. Under Export, select All training data as CSV or Selected rows as CSV.

  4. When the export is ready, use the download link in the Notifications menu.

An export option is unavailable if it would include documents that are still loading. Hover over an unavailable option to see why. Wait for loading to finish, or select only ready rows.

CSV contents

Training data CSV exports use status names, include layout names with layout-version identifiers, and contain the columns relevant to the model type.

Annotate training documents

Annotation establishes the ground truth that the specialized model learns from. Review every supported field against the source document, even when the system has pre-populated a value.

  1. In the Training Data tab, find a document that is ready to annotate.

  2. Select Annotate, or Edit annotations for a document that has already been annotated.

    • You can also open the annotation view by selecting the document’s Document ID link.

  3. Select a field and identify the text that represents its correct value.

  4. Review the field’s location and transcription. Correct the transcription to match the document exactly.

  5. Repeat for the remaining fields, and then click Save.

Text Segmentation

  • For faster and easier annotation, use the dotted lines around each value.

  • These boxes represent the text segmentation in the document. To learn more, see Text Segmentation.

  • Drag and drop your cursor to create a Bounding Box.

Text segmentation boxes shown as dotted lines around values in a document

While annotating:

  • Navigate between documents by selecting the left or right arrow (Left and right navigation arrows).

  • Documents appear in the order you’ve sorted them in the Training Data table.

To learn more about the annotation experience in TDM for ORCA VLMs, see the interactive walkthrough below.

Automatically transcribed fields

  • The application pre-populates some fields to reduce manual typing.

  • Treat these values as suggestions:

    • review each one and correct it before completing the annotation.

In v43.2 and later, moving or resizing a field’s bounding box updates its transcription with the text segments inside the box. Review the updated value and correct it if necessary.

For guidelines on entering values so that they provide consistent ground truth (exact transcription, reading order, and multi-page fields), see Training a Specialized Model.

Overview tab

Check training readiness

After annotating, check the Training summary on the Overview tab to see whether the model is ready to train.

Field or status

Description

Ready to train

At least one compatible base model is installed and has enough eligible annotated documents.

Reqs not met

No compatible base model is installed, or the training data requirement has not been met. Hover over the information icon to identify the unmet requirement.

Training in progress

Model training is running.

Training failed

The most recent training run failed.

Required Documents

Before the minimum is met, shows the number of eligible annotated documents available out of the 120 required for ORCA VLM training.

Total Documents

After the minimum is met, replaces Required Documents and shows the total number of eligible annotated documents.

Base Model

The base model used by the latest training (e.g., ORCA 1.0 or ORCA 2). Before a training run has selected a base model, the value is N/A.

Train a specialized model

Training uses annotated documents to create a specialized model on top of an installed ORCA base model. It does not modify the base model.

  1. Confirm that the Training summary shows Ready to train.

  2. Open the page-level Actions menu and select Train model.

  • If one eligible base model is installed, the system schedules training with that model.

  • If multiple base models are installed, select an eligible model in the Select model dialog, and then select Train model.

  • You cannot select a base model that does not have enough eligible documents.

Different base models provide different capabilities

The specialized models inherit the capabilities of the base model selected for the training run. In v43.2 and later, you can train on the ORCA 2 base model, which is available in addition to ORCA 1.0.

Evaluate and deploy the candidate

After training completes, a candidate model appears on the Overview tab, alongside the live model’s summary card if one exists. Review the candidate’s summary and its Projected Automation. Decide whether you want to deploy or reject it from the page-level Actions menu or the History tab.

For the complete guidance on evaluating a candidate, deploying it, and validating the deployed model with test documents, see Training a Specialized Model.

Action notifications

The application displays success notifications when model actions are completed. These actions include deploying or undeploying a model, rejecting a candidate, scheduling or canceling training, and exporting a model.

History tab

Manage model versions

The History tab lists the base models and specialized model versions associated with the model definition. Available actions depend on each model’s source, state, compatibility, and your permissions.

Column

Description

Name

The model name. Specialized-model names can contain up to 128 characters. Hover over a shortened name to see it in full.

UUID

The model version’s unique identifier. This column is hidden by default and displays the first eight characters when enabled.

State

Live, Candidate, Inactive, or Not installed. Archived models appear in the Archive view.

Compatibility

The model’s compatibility with Hyperscience versions. Orange indicates the current version, green one version ahead, and blue two versions ahead.

Layout version

The layout version used to train a specialized model. Base-model rows show n/a.

Source

Internal, Uploaded, or Base.

Base Model

The base model used to create the specialized model.

Proj auto

The projected automation and margin of error at the selected Test Target Accuracy. Rows without projection data show n/a.

Docs trained

The number of documents used to train the specialized model. Base-model rows show n/a.

Actions

The actions available for the model version. A row can show an inline Deploy or Install action and additional actions in its three-dot menu.

Show and copy a model UUID

  1. In the History tab, open the table-options menu next to Filter.

  2. Select Manage columns....

  3. Select UUID, and then select Save.

  4. Hover over a shortened UUID to see the full value, or select the UUID to copy it.

Resolve a missing base model

If a specialized model depends on a base model that is not installed, the specialized model is unavailable for deployment, and its row is grayed out. You can still download or archive the specialized model.

  1. In the History tab, find the required base-model row with the state Not installed.

  2. Select Install. Administration > Assets opens in a new tab.

  3. Install the base model, and then return to the model details page. To learn more, see Installing ORCA VLMs.

Archive and unarchive specialized models

Archive specialized model versions that you no longer actively use. Base models cannot be archived from the History tab. Undeploy a Live model or reject a Candidate model before archiving it.

  1. In the History tab, open the specialized model’s three-dot menu.

  2. Select Archive and confirm the action.

  3. To see archived models, open the table-options menu and select View Archive.

  4. To restore a model, open its three-dot menu and select Unarchive.

  5. Select View Active to return to active models.

Next steps

For the end-to-end specialization workflow — dataset requirements, annotation guidelines, and validating a deployed model with held-out test documents — see Training a Specialized Model.