--- title: "Labeling Anomaly Detection" slug: "labeling-anomaly-detection" updated: 2026-01-21T19:30:26Z published: 2026-01-21T19:30:26Z canonical: "help.hyperscience.ai/labeling-anomaly-detection" --- > ## Documentation Index > Fetch the complete documentation index at: https://help.hyperscience.ai/llms.txt > Use this file to discover all available pages before exploring further. # Labeling Anomaly Detection > [!WARNING] > **Accessing this feature** > > Your access to the feature described in this article depends on your license package and pricing plan. > > To learn which features are available to your organization and how to add more, contact your Hyperscience representative. A high-quality model requires consistent annotations. That's why identifying potential discrepancies in the training sets before model training is crucial. To help with this effort, we've included a tool called Labeling Anomaly Detection in Training Data Management (TDM). After completing the annotations, you can analyze the Training Data to find inconsistencies in your documents and the ones ineligible for training. To learn more about eligibility, see our [Document Eligibility Filtering](/v42/docs/document-eligibility-filtering) article. Labeling Anomaly Detection identifies and highlights potential anomalies in field and table annotations for review. Before using Labeling Anomaly Detection: 1. Upload the Required Number of Documents (100 minimum, 400 recommended). 2. Run Training Data Analysis. 3. Annotate your Training Set. 4. Reanalyze your data. Always re-run data analysis to get the most up-to-date information about your training set. Ineligibility details may change if documents have been added, removed, or modified since the last analysis. For more information, see Step 4 of our [Training a Semi-structured Model](/v42/docs/v42-training-a-semi-structured-model) article. ## Detecting anomalies If anomalies were detected during the training data analysis: - A count of documents with potential anomalies appears on the Training Data Health card. - In the Training Data table, each document containing potential anomalies is highlighted with a yellow bar on the left-hand side of its Doc ID. ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/indicator(1).gif) 1. Expand the **Filter** section, and select **Contains Anomalies** from the **Has Anomalies** drop-down list. 2. Click **Apply Filters**. 3. Click the **Edit Annotations** link for a document highlighted as having anomalies. 4. Review the annotations highlighted as being a potential anomaly. - Check how the field was annotated in other documents in the same group. That way, you'll ensure consistency throughout the training set. If the Annotation is not correct, adjust it accordingly and click **Save Changes**. - If the annotation is correct, click on it and then click **Ignore Anomaly.** A warning message appears in v42.1 and earlier: ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/28274732407309.png) - Click **Confirm** and then **Save Changes**. - In v42.1 and later, a green label indicates that a field is missing after Training Data Analysis. Click the **Mark field as missing** button (![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/image(212).png)), located at the bottom of the page preview, and continue with the review. You can also use the new keyboard shortcut (**Command** + **Option** + **M** for Mac, **Ctrl** + **Alt** + **M** for Windows) to speed up the process. ### Limitations of Labeling Anomaly Detection > [!NOTE] > - Anomalies can only be detected for fields with a single occurrence and a single Bounding Box. Therefore, you could still see incorrect annotations across fields with Multiple Occurrences (MOs) and Multiple Bounding Boxes (MBB). Make sure to double-check all fields before submitting the document. > - Table Anomaly Detection won't capture all errors in the annotations. If a column is missing, the other documents in the same group must have that column for the system to mark the missing column as an anomaly. > - You can run Labeling Anomaly Detection for up to 5,000 pages at a time. ### Anomaly indicators The indicators described in the table below appear in the document viewer if anomalies are detected in the document. | **Indicator** | **Description** | **Example** | | --- | --- | --- | | **Cell-anomaly indicator** - found in the right-hand sidebar | **Dotted line around a cell** - indicates a cell that needs to be reviewed | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/image-1750683163396.png) | | **Page Indicator** - found in the left-side sidebar | **Dotted line around a page** - indicates **cell-level** anomalies on the specific page | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542247367693.png) | | **Missing columns label** - found on the top of the document, next to the colored markers for each column | The indicator for a missing column is a dotted, transparent label, located next to the bookmark indicators for each column. It indicates all possible missing columns on a document level. | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542247371149.png) | | **Missing column tag** - found in the right-hand sidebar next to the specific column that is missing | This tag shows the specific missing column. | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542247373581.png) | | **Misplaced column** - found around the colored markers for columns | **Dotted line around the colored markers** - indicates misplaced columns | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542283409293.png) | | **Misplaced column tag** - found in the right-hand sidebar next to the specific column that is misplaced | This tag shows the specific misplaced column. | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542283437837.png) | | **Number of anomalies** - found in the right-hand sidebar above the list of columns in the document | This yellow indicator displays the current number of potential anomalies in the document. It is dynamic and changes after each interaction with an annotation labeled as an anomaly. If you have a nested table, the number of potential anomalies will also appear next to the name of the parent or child table. | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27542247411341.png) | | **Ignore anomaly** - action button, located in the colored label for a column | The bell button appears for single and multiple anomalies in a column. Click on it to ignore the detected anomaly **Single anomaly** - hover over it to see the specific anomaly **Multiple anomalies** - hover over to see the number of potential anomalies for this column | ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/27668711095309.png) | ### Re-analyzing data We recommend re-analyzing the data after reviewing all anomalies to ensure the training set is consistent and ready for model training. Click **Reanalyze data** to choose one of the two options listed below: - **Reanalyze with ignored anomalies** — If you decide to reanalyze the training set with the ignored anomalies, any anomalies that were previously ignored will not reappear. - **Reanalyze from scratch** — If you want to analyze the training set from scratch, the ignored anomalies will be included in the results. ![](https://cdn.us.document360.io/87894cef-4958-4f3f-be6f-b75a78c82548/Images/Documentation/image(193).png) A tool used to annotate, manage, import, and export training documents. It is also used to train models by working directly with the training data (“ground truth”) obtained from each document in the training set. The input used to teach machine learning models how to process documents accurately. Its structure depends on the model type: - For **Classification models**, training data consists of uploaded document pages grouped by layout. - For **Identification models**, training data includes manually annotated fields and tables to train the model to extract specific data points. The minimum amount of annotated data needed to train machine learning models in Hyperscience. - For **Classification models**, **at least 10 pages per layout** are required. - For **Identification models**, the system requires **a minimum of 100 documents**. Meeting these thresholds ensures the models can be successfully trained and begin learning layout- or document-type patterns. A tool in TDM that analyzes your training data to compute the importance of each training document and identify issues such as missing labels, overlapping fields or columns, or inconsistent annotations. This analysis helps you prioritize which documents to annotate and ensures clean, accurate data before you train a model. A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples. **Annotation** refers to a user-provided input that defines the correct prediction for a given machine learning task. Annotations are used to train supervised machine learning models. A rectangular subregion of a given page that specifies the location of text to be processed downstream or to be displayed to the user. Multiple Occurrences (MOs) are used to identify multiple different instances of a field. Used to annotate a single field's value when it spans across line or page breaks, ensuring the accurate capture of information across multiple locations or pages.