--- title: "Glossary" slug: "glossary" description: "Understand the key concepts of the Hyperscience platform with our comprehensive terminology guide to ensure clarity and consistency in your usage." tags: ["terminology"] updated: 2026-08-24T11:45:40Z published: 2026-08-24T11:45:40Z canonical: "help.hyperscience.ai/glossary" --- > ## Documentation Index > Fetch the complete documentation index at: https://help.hyperscience.ai/llms.txt > Use this file to discover all available pages before exploring further. # Glossary This article provides a list of terms as a reference to help you better understand key concepts related to the Hyperscience platform. It clarifies commonly used terminology, ensuring consistency across documentation and conversations. By using these definitions, you’ll gain a clearer understanding of how our platform works and how to make the most of its features. ## A | Term | Definition | Reference | | --- | --- | --- | | **Accuracy** | Accuracy measures the effectiveness of the models based on the proportion of correct predictions out of all predictions made. It helps you understand how often the system correctly predicts values compared to the actual values that reached consensus during QA. Accuracy can be influenced by factors like imbalanced datasets or inconsistent annotations. | [Accuracy](https://help.hyperscience.ai/latest/docs/accuracy) | | **Additional Layouts** | A layout type used to categorize pages where no data is extracted (e.g., a fax cover sheet). These layouts allow users to define custom categories for unmatched pages, helping to improve classification accuracy. Additional layouts can apply to Structured, Semi-structured, or Unstructured documents. | - [Creating Additional Layouts](https://help.hyperscience.ai/latest/docs/creating-additional-layouts) - [Understanding Document Types](https://help.hyperscience.ai/latest/docs/understanding-document-types) | | **Average Handling Time** | A metric that represents the average time it takes to process a submission. The value is averaged across multiple documents or submissions within a specific date range. | [Operational Value Reporting](https://help.hyperscience.ai/latest/docs/operational-value-reporting) | | **Annotation** | A user-provided input that defines the correct prediction for a given machine learning task. Annotations are used to train supervised machine learning models. | - [Text Segmentation](https://help.hyperscience.ai/latest/docs/text-segmentation) - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Anomaly** | A potential inconsistency or error in how a field or table is labeled in a document. Anomalies are flagged to help ensure consistent and accurate training data, which improves model performance. | [Labeling Anomaly Detection](https://help.hyperscience.ai/latest/docs/labeling-anomaly-detection) | | **API Blocks** | These blocks enable Hyperscience to interact with external systems through APIs, facilitating tasks like data retrieval, validation, or sending information to other applications. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Auto Thresholding** | An automated process that calculates a confidence threshold based on a target accuracy. Predictions below this threshold are sent to Supervision for a human review to ensure the system meets the desired accuracy. | - [Transcription Accuracy and Automation](https://help.hyperscience.ai/latest/docs/transcription-accuracy-and-automation) - [Accuracy](https://help.hyperscience.ai/latest/docs/accuracy) - [Automation](https://help.hyperscience.ai/latest/docs/automation) | | **Automation** | The processing of data without the need for human intervention. | [Automation](https://help.hyperscience.ai/latest/docs/automation) | | **Automation Rate** | The extent to which a machine can process data independently without requiring human supervision. It represents the proportion of extracted data with confidence scores exceeding a specified threshold. This threshold is determined by the level of accuracy you want your extracted data to have. | [Automation](https://help.hyperscience.ai/latest/docs/automation) | | **Auto-Splitting** | A feature in Hyperscience that automatically groups pages into documents using rules you define. It helps organize Semi-structured documents by deciding where one document ends and another begins based on page count, text patterns (like titles), or layout-specific logic. | [Auto-Splitting](https://help.hyperscience.ai/latest/docs/auto-splitting) | ## B | **Base Model** | A foundational model that provides core general-purpose capabilities and is not directly trained on customer-specific examples. Use-case specialization is achieved through additional training on top of the base model using customer-specific data. | [Model Definitions](https://help.hyperscience.ai/latest/docs/model-definitions) | | --- | --- | --- | | **Bounding Box** | A rectangular subregion of a given page that specifies the location of text to be processed downstream or to be displayed to the user. | [Text Segmentation](https://help.hyperscience.ai/latest/docs/text-segmentation) | | **Bundle** | A packaged file that contains everything needed to install or upgrade the Hyperscience platform. It includes the application and all required tools, helping to streamline setup and upgrade processes. | - [Installing Hyperscience](https://help.hyperscience.ai/deployment/docs/installing-hyperscience-hypercell) - [Upgrading Hyperscience](https://help.hyperscience.ai/deployment/docs/upgrading-hyperscience-hypercell) | | **Bypass Validation if Layout ID is Missing** | A flow-level setting that bypasses validation by layout identifier if the matched Structured layout variation doesn’t have an identifier specified. The bypass allows the system to continue classifying documents even without layout identifiers, ensuring that documents that are not tied to a specific layout variation are still processed. | [Structured Classification and Layout Identifiers](https://help.hyperscience.ai/latest/docs/structured-classification-and-layout-identifiers) | ## C | **Calibration** | A quality check performed after QA on Structured documents to evaluate model performance. It helps set target accuracy levels, define baseline automation thresholds, and assess how well different layouts, fields, or data types are processed before going live. | Contact your Hyperscience Representative for more information. | | --- | --- | --- | | **Candidate model** | A trained or imported model that is available for evaluation but is not currently used for document processing. | [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) [TDM for Identification Models](https://help.hyperscience.ai/latest/docs/tdm-for-identification-models) [TDM for ORCA VLMs](https://help.hyperscience.ai/latest/docs/tdm-for-orca-vlms) | | **Case** | A group of related documents, files, or pages that are processed together using a unique Case ID. | [Case Collation](https://help.hyperscience.ai/latest/docs/case-collation) | | **Cell** | A value in a table that holds a single piece of data, such as a name, number, date, or multiline entries like an address or description. In Hyperscience, cells are key to reading and extracting data from tables accurately. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | **Character** | Any single letter, number, or symbol found in a document. Hyperscience reads characters to understand and extract text. | - [Supported Characters and Data Types](https://help.hyperscience.ai/latest/docs/supported-characters-and-default-data-types) - [Default Data Types](https://help.hyperscience.ai/latest/docs/default-data-types) | | **Checkbox** | A non-text field used to capture two-option answers like “Yes/No” or “True/False.” | [Checkboxes and Signatures](https://help.hyperscience.ai/latest/docs/checkboxes-and-signatures) | | **Classification Model** | A machine learning model that automatically identifies a document’s type—Structured, Semi-structured, or Additional—and matches it to the correct layout. This classification helps Hyperscience process different document types accurately without manual intervention. | - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) - [Understanding Document Types](https://help.hyperscience.ai/latest/docs/understanding-document-types) - [Classification Models](https://help.hyperscience.ai/latest/docs/classification-models) | | **Classify Using Layout Identifier** | A flow-level setting that allows Structured documents to be matched using a layout identifier. See [Layout Identifier](/general-information/docs/glossary#l) for more context. | [Structured Classification and Layout Identifiers](https://help.hyperscience.ai/latest/docs/structured-classification-and-layout-identifiers) | | **Clustering** | The process of grouping similar documents or data points based on shared characteristics, often using machine learning algorithms. Clustering helps the platform to better organize and interpret large volumes of data by recognizing patterns and similarities. | [Document Drift Management (Layout Triage)](https://help.hyperscience.ai/latest/docs/document-drift-management-layout-triage) | | **Collation** | The process of grouping related files, documents, or pages into a single case using a unique identifier called a Case ID. For example, if you submit multiple documents for a loan application, collation ensures that all these documents are grouped together under one case for streamlined processing and review. | [Case Collation](https://help.hyperscience.ai/latest/docs/case-collation) | | **Column** | A list of values in a table that are of the same type of information, like names or prices, with one value per row. In simple tables, columns usually appear as vertical sections. However, in more complex tables, columns may not follow a vertical layout but still represent the same kind of data across rows. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | **Consensus** | A process used to confirm the correct value of a transcribed field. Consensus is reached when two matching transcriptions are provided for the same field, usually one from a human and one from the machine or two separate human-provided entries. This process ensures higher accuracy, especially when the system's confidence is low. | [Transcription Supervision Consensus](https://help.hyperscience.ai/latest/docs/transcription-supervision-consensus) | | **Continuous Field Locator model improvement** | When this setting is enabled, the system automatically retrains and updates Field Locator models using newly available QA data. This process allows the model to improve over time without manual intervention. It helps enhance accuracy for identifying field locations in Semi-structured documents. This setting should only be enabled if there’s enough training data in the environment to support it. | [Identification Settings](https://help.hyperscience.ai/latest/docs/identification-settings) | | **Copycat** | After you’ve annotated a single row from a table, you can use the copycat feature to copy the annotations to the remaining rows of the table. The copycat is not always accurate, so make sure to double-check the annotations before you submit. | - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) - [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | **Crop** | An image of the specific field you want to extract. | [Text Segmentation](https://help.hyperscience.ai/latest/docs/text-segmentation) | | **Custom Code Block** | A flexible component in Hyperscience flows that allows you to add custom Python logic to transform, validate, or enrich data before it's sent to downstream systems. It lets you apply your own business rules as part of document processing. | - [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) - [Modifying Custom Code Blocks](https://help.hyperscience.ai/latest/docs/modifying-custom-code-blocks) | | **Custom Data Type** | A user-defined format that tells Hyperscience what a specific type of data should look like, such as a Social Security Number or a policy ID. Custom data types help the system validate and extract field values more accurately based on expected patterns. | - [Creating Data Types with ML Configurations](https://help.hyperscience.ai/latest/docs/creating-data-types-with-ml-configurations) - [Creating Data Types with a List of Expected Values](https://help.hyperscience.ai/latest/docs/creating-data-types-with-a-list-of-expected-values) - [Creating Data Types with Custom Patterns](https://help.hyperscience.ai/latest/docs/creating-data-types-with-custom-patterns) | | **Custom Supervision** | A configurable task in Hyperscience that you can tailor to your business needs. It allows you to manually review, validate, or enrich data using flexible logic, custom fields, and decision types. | [Custom Supervision](https://help.hyperscience.ai/latest/docs/custom-supervision) | ## D | **Database Block** | A specific type of block that allows you to connect Hyperscience to external databases. These blocks allow the system to fetch or validate information during document processing. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | --- | --- | --- | | **Data Extraction** | The process of pulling specific information—like names, dates, or amounts—from a document. In Hyperscience, extraction happens after a page is matched to a layout and uses trained models to identify and capture the right data. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | **Dataset** | A group of documents used to help the system learn or improve. Datasets are used for training, testing, or evaluation of how well the system reads and extracts information. | [Preparing training data](https://help.hyperscience.ai/latest/docs/preparing-training-data) | | **Data Type** | A property that defines the format of the data expected in a field, like numbers, dates, or email addresses. For example, the data type **Date** accepts only valid dates (e.g., MM/DD/YYYY). Data types help Hyperscience understand what’s expected in a field and flag anything that doesn’t match. | [What is a Data Type?](https://help.hyperscience.ai/latest/docs/what-is-a-data-type) | | **Deployed Model / Live Model** | A trained machine learning model that has been activated within Hyperscience to process documents in real time. Once deployed, the model is live and is used to process documents—classifying them, locating fields, and extracting data based on what it has learned. | [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Document** | A group of one or more pages processed as a single unit in Hyperscience. Documents are categorized as Structured, Semi-structured, or Additional based on how consistent their layouts are and how fields can be extracted from their pages. | [Understanding Document Types](https://help.hyperscience.ai/latest/docs/understanding-document-types) | | **Document Classification Quality Assurance** | A task where you review and confirm whether Hyperscience correctly identified the type of each page. This process helps improve the system’s ability to match pages to the right layout, supports the training of the Classification model, and is used for document-classification reporting. | [Document Classification Task](https://help.hyperscience.ai/latest/docs/document-classification-task) | | **Document Classification Task** | The first step in Supervision. It is used to categorize and combine pages that were not classified by the machine. | [Document Classification Task](https://help.hyperscience.ai/latest/docs/document-classification-task) | | **Document Drift Management (Layout Triage)** | A post-processing feature that helps you manage documents that don't match layouts during Classification. When submissions don't meet the Structured Layout Match Threshold or are manually flagged as having incorrect or missing layouts, their pages are marked as unmatched. | [Document Drift Management (Layout Triage)](https://help.hyperscience.ai/latest/docs/document-drift-management-layout-triage) | | **Document Eligibility Filtering** | A feature in Training Data Management that indicates whether a document is eligible for training based on internal checks in the application and our machine learning logic. It provides additional information about documents that were excluded from the training set. | [Document Eligibility Filtering](https://help.hyperscience.ai/latest/docs/document-eligibility-filtering) | | **Document Renderer Block** | A step in a Hyperscience flow that turns processed documents into downloadable PDFs and generates links to access them. You can customize the page size and image quality of the PDFs to meet your needs. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Dropout** | A field-level setting that tells the system to ignore background text like pre-printed labels or symbols. When enabled, the system removes this background content and transcribes only new or handwritten text, helping to improve accuracy. This setting is enabled by default. | [Building a Structured Use Case](https://help.hyperscience.ai/latest/docs/building-a-structured-use-case) | ## E | **Eligible Number of Documents** | The number of documents that meet the requirements for training Identification models in Hyperscience. A minimum of 100 is needed to train, with 400 recommended for best results. This number applies to Field Identification and Table Identification models. | - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) - [Document Eligibility Filtering](https://help.hyperscience.ai/latest/docs/document-eligibility-filtering) | | --- | --- | --- | | **“.env” File** | A configuration file used to define environment-specific variables, such as API keys or database credentials. It allows Hyperscience to run securely and consistently across different instances. | [Editing the ".env" file and running the application](https://help.hyperscience.ai/deployment/docs/editing-the-env-file-and-running-the-application) | | **Excluded Documents** | Documents added to a Classification training set to show the system what should not be matched to a specific layout. They help improve model accuracy by teaching your model to ignore documents that look similar but don’t belong. | [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) | ## F | **False Negatives** | A type of error or outcome in machine learning evaluation. In Hyperscience, a false negative happens when the model fails to extract a field value that is clearly present and should have been captured. False negatives apply to Field Identification and Table Identification models. | [Scoring Field Identification Accuracy](https://help.hyperscience.ai/latest/docs/scoring-field-identification-accuracy) | | --- | --- | --- | | **False Positive** | A type of error or outcome in machine learning evaluation. In Hyperscience, a false positive happens when the model predicts a field value in the wrong location or extracts something that shouldn’t be considered a valid field at all. False positives apply to Field Identification and Table Identification models. | [Scoring Field Identification Accuracy](https://help.hyperscience.ai/latest/docs/scoring-field-identification-accuracy) | | **Field** | A labeled piece of information you want to capture from a document, like “Name,” “Date of Birth,” or “Total Amount.” In Hyperscience, fields allow you to specify the values that will be extracted from your documents. | [Field Identification](https://help.hyperscience.ai/latest/docs/field-identification) | | **Field Customization** | A feature that allows you override a field's default settings on a per-release basis. For example, you can set a layout's "Name" field to always be sent to Supervision for a given release, but not for other releases the field's layout appears in. | [Creating Field Customizations](https://help.hyperscience.ai/latest/docs/creating-field-customizations) | | **Field Dictionary** | A centralized place in Hyperscience where you define and manage field customizations for Structured layouts. It helps keep field names, data types, and output settings consistent across different documents | [Navigating the Field Dictionary](https://help.hyperscience.ai/latest/docs/navigating-the-field-dictionary) | | **Field Identification Quality Assurance** | A manual quality check for Semi-structured documents where you review and correct the system’s predicted field locations or humans’ input. Doing so helps improve how accurately your model finds and extracts information in future documents. The human’s input in QA is used to measure human performance for reporting purposes. | [Field Identification Quality Assurance](https://help.hyperscience.ai/latest/docs/field-identification-quality-assurance) | | **Field Identification Task** | A manual task in Hyperscience where you confirm or correct the location of fields in Semi-structured documents. You can adjust or draw bounding boxes around field values to help your model learn where to look for the data you want to extract. You may need to perform a Field ID task when the machine is not confident enough in its prediction for a field, based on the target accuracy. | [Field Identification](https://help.hyperscience.ai/latest/docs/field-identification) | | **Field Identification Model / Field Locator Model** | A machine learning model in Hyperscience that learns where fields are located in Semi-structured documents. It uses examples from training to predict the position of each field on a page so the system can extract the right data. | [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Finetuning** | A process that improves the accuracy by using your data to adjust system thresholds automatically. It helps ensure the system makes more accurate predictions and flags uncertain results for review, reducing errors and improving overall performance. | [Transcription Models Overview](https://help.hyperscience.ai/latest/docs/transcription-models-overview) | | **Field-Level Accuracy Targets (FLAT)** | A configuration in Hyperscience that allows you to set different accuracy levels for specific fields or table columns. For example, if you need higher accuracy for fields like addresses or account numbers, you can set a higher target for them while keeping other fields at a lower target accuracy. Doing so helps improve the precision of critical fields without adding extra tasks. | [Identification Settings](https://help.hyperscience.ai/latest/docs/identification-settings) | | **Flexible Extraction** | A task in Hyperscience that involves human intervention to validate or correct data extraction for Structured documents. This task is used when automatic extraction isn’t fully reliable, allowing you to transcribe or adjust specific fields to ensure accuracy. | [Flexible Extraction](https://help.hyperscience.ai/latest/docs/flexible-extraction) | | **Flexible Extraction Block** | A component in Hyperscience that allows you to define when documents or fields should undergo Flexible Extraction. It enables the validation of transcriptions or adding data to documents that were manually categorized or skipped regular Transcription Supervision. To use it, you need a Custom Code Block to set the specific rules. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Flow (Workflow)** | A customizable workflow in Hyperscience that automates processes, including steps like classification, data extraction, validation, and output. Flows streamline operations by handling tasks sequentially with minimal manual effort. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | **Flow Block** | A modular step in a flow that performs a specific processing, integration, or control function. A flow block can receive inputs, apply configured logic or settings, and pass its output to subsequent blocks. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Flow Identifiers** | A unique name or identifier for a specific flow within the system. It helps distinguish flows clearly, especially during testing and debugging. This identifier may differ from labels displayed in the UI. For more information, contact your Hyperscience representative. | [Testing and Debugging Flows](https://help.hyperscience.ai/latest/docs/testing-and-debugging-flows) | | **Flow Run** | The complete execution cycle of a specific flow within Hyperscience. Each flow run encompasses all steps from initiation to completion for a given submission, allowing you to monitor, troubleshoot, and manage document processing. | [Flow Runs Page](https://help.hyperscience.ai/latest/docs/flow-runs-page) | | **Flow Studio** | Flow Studio enables you to view and configure flows. In Flow Studio, you can inspect a flow’s structure, update block settings, and use Build Mode to modify the blocks included in a flow. You can also add or remove blocks, or configure their settings directly in the application, using the Flow Builder. | [Editing flows in Flows Studio](https://help.hyperscience.ai/latest/docs/editing-flows-in-flow-studio) | | **Full Page Transcription (FPT)** | A process in Hyperscience that captures all visible text on a page, not just specific fields. It’s useful for documents with Unstructured layouts, enabling broader data extraction and analysis. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | ## G | **Graphics Processing Unit (GPU)** | A processing unit initially designed to increase the speed of graphics calculations. It is efficient in completing matrix computations, making it faster than a central processing unit (CPU) in some calculations related to machine learning. | - [Enabling Trainers with GPUs in On-Premise Podman Deployments](https://help.hyperscience.ai/deployment/docs/enabling-trainers-with-gpus-in-on-premise-podman-deployments) - [Enabling Trainers with GPUs in On-Premise Docker Deployments](https://help.hyperscience.ai/deployment/docs/enabling-trainers-with-gpus-in-on-premise-docker-deployments) - [Enabling Trainers with GPUs in On-Premise Kubernetes Deployments](https://help.hyperscience.ai/deployment/docs/enabling-trainers-with-gpus-in-on-premise-kubernetes-deployments) | | --- | --- | --- | | **Ground Truth** | Manually annotated data used to train our machine learning models. We use a subset of this data to assess the performance of your models | [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Groups** | A collection of documents in the Training Data Curator used to organize training data for machine learning models. Groups help you manage, annotate, and track documents based on specific use cases, such as invoice processing. | [Training Data Curator](https://help.hyperscience.ai/latest/docs/training-data-curator) | ## H | **Human-in-the-Loop (HITL)** | A process where people review and correct data that the Hyperscience platform has low confidence in. It helps improve accuracy and ensures high-quality results, especially when the system is unsure. Also known as Supervision. | [What is Supervision?](https://help.hyperscience.ai/latest/docs/what-is-supervision) | | --- | --- | --- | | **Hypercell** | The core software platform developed by Hyperscience for building and operating document-processing solutions. It brings together the application, models, flows, human-in-the-loop task interfaces, and integrations as a unified system. | | ## I | **Intelligent Document Processing (IDP)** | The customized automation of data extraction from paper-based documents or document images to integrate with specific digital business processes. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | --- | --- | --- | | **IDP Flow** | The end-to-end process in Hyperscience where documents are uploaded, classified, and automatically processed to extract key data. It includes steps like document ingestion, data extraction, validation (with Human-in-the-Loop if needed), and structured output delivery to downstream systems. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | **Image (Page Image)** | A file (e.g., a scanned document or photo) that the Hyperscience platform processes to extract text and data from. | [How a File Becomes a Submission](https://help.hyperscience.ai/latest/docs/how-a-file-becomes-a-submission) | | **Image Correction** | A setting in the Machine Classification Block that identifies and corrects the orientation of page images by automatically rotating them. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Incremental Training** | A process that enables you to efficiently update your existing model by incorporating new data or minor annotation changes without losing previously learned information. | [Improving Model Performance](https://help.hyperscience.ai/latest/docs/improving-model-performance) | | **Input Blocks** | Blocks that enable you to integrate and process documents from your organization's data sources (such as inboxes, message queues, or network folders) within the system. | - [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) - [Input Blocks](https://help.hyperscience.ai/latest/docs/input-blocks) | ## J | **Job** | A logical unit of work to be accomplished within the system. | [Jobs Page](https://help.hyperscience.ai/latest/docs/jobs-page) | | --- | --- | --- | ## K | **Knowledge Store** | A database-like feature that allows you to store business data (e.g., vendor names, addresses) that can be displayed as decision choices during Custom Supervision. Configuring Custom Supervision tasks to retrieve validated sets of choices from the Knowledge Store prevents errors and reduces time spent on Supervision. | [Knowledge Store](https://help.hyperscience.ai/latest/docs/knowledge-store) | | --- | --- | --- | ## L | **Layout** | The way a page is arranged in a document. In Hyperscience, layouts guide how data is extracted from different types of documents, like forms, invoices, and purchase orders. There are three types of layouts: Structured, Semi-structured, and Additional, depending on how consistent the content and its arrangement are across pages. | - [Understanding Document Types](https://help.hyperscience.ai/latest/docs/understanding-document-types) - [What is a Layout?](https://help.hyperscience.ai/latest/docs/what-is-a-layout) - [Determining Layout Type](https://help.hyperscience.ai/latest/docs/determining-layout-type) | | --- | --- | --- | | **Layout Identifier** | A string of characters that appears on a specific location on a page in a Structured layout. It allows the machine to distinguish among similar layouts that a submission matches to. It can contain letters, numbers, or a combination of both. Layout identifiers help the machine achieve the best match for submitted pages. | [Layout Identifiers](https://help.hyperscience.ai/latest/docs/layout-identifiers) | | **Layout Variation** | Layout variations occur when documents of the same type (such as HCFA or W-8 forms in the US) have the same key information but differ in how that information is arranged on the page. | [Adding a Variation to a Layout](https://help.hyperscience.ai/latest/docs/adding-a-variation-to-a-layout) | | **Layout Version** | An iteration of a layout that reflects the state of a layout at a given point in time. You can create new versions and restore older ones based on your needs. | [What is a Layout Version?](https://help.hyperscience.ai/latest/docs/what-is-a-layout-version) | | **Large Language Models (LLMs)** | Advanced models that understand and generate human-like text. They are used to enhance document-processing tasks, such as data extraction, summarization, and classification, by interpreting complex content. As a result, LLMs can improve automation and accuracy. | [Using the General Prompting Block](https://help.hyperscience.ai/latest/docs/using-the-general-prompting-block) | | **Large Language Models Install Block** | A component that checks for and installs a Large Language Model (LLM) in the system if it’s not already present. This block ensures the necessary LLM is available for processing tasks that require advanced language understanding. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | ## M | **Machine Classification Block** | A flow component that automatically detects and categorizes document types using AI, helping route them through the correct processing steps. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | --- | --- | --- | | **Machine Collation Block** | A flow component that automatically groups related documents or data points, streamlining the organization and processing of information. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Machine Identification Block** | A flow component that automatically finds and labels key fields on a page so the system knows what data to extract. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Machine Transcription Block** | A flow component that reads and converts text from a document into structured, machine-readable data. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Manual Classification Block** | A flow component that enables human reviewers to assign the correct document type when the model is not sure. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Manual Identification Block** | A flow component that allows a human to annotate fields on a page to teach the model where information is located. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Manual Transcription Block** | A flow component that allows a human reviewer to transcribe text when the system is not confident that its transcription is accurate. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Margin of Error (MoE)** | The range of uncertainty in the system's estimate of accuracy. It shows you how much the estimate may differ from the true value. The smaller the margin of error is, the more confident the system is in its estimate. | [Accuracy](https://help.hyperscience.ai/latest/docs/accuracy) | | **Model** | A construct in our system that performs specific tasks, like classifying documents, identifying fields, or extracting text. The model improves over time as it learns from more examples. | - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) - [Building a Structured Use Case](https://help.hyperscience.ai/latest/docs/building-a-structured-use-case) - [Transcription Models Overview](https://help.hyperscience.ai/latest/docs/transcription-models-overview) - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) | | **Model Definition** | A configuration object used to manage Vision Language Models (VLMs). Model definitions define a model's scope, task, compatibility, deployment status, and associated model versions. | [Model Definitions](https://help.hyperscience.ai/latest/docs/model-definitions) | | **Multiple Bounding Boxes (MBB)** | Used to annotate a single field's value when it spans across line or page breaks, ensuring the accurate capture of information across multiple locations or pages. | [Field Identification](https://help.hyperscience.ai/latest/docs/field-identification) | | **Multiple Occurrences (MO)** | Used to identify multiple distinct instances of a field. | [Field Identification](https://help.hyperscience.ai/latest/docs/field-identification) | | **Model Validation Tasks (MVTs)** | Tasks used to review and correct data predictions made by Hyperscience models. They help evaluate model accuracy and provide feedback that can be used to improve future model performance. | [Model Validation Tasks](https://help.hyperscience.ai/latest/docs/model-validation-tasks) | | **Multiline** | A setting in the Layout Editor for fields or columns that may span more than one line of text (e.g., Address, Description). When this setting is not enabled and a field's value spans more than one line, it will not be extracted in its entirety. | - [Creating Structured Layouts](https://help.hyperscience.ai/latest/docs/creating-structured-layouts) - [Creating Semi-structured Layouts](https://help.hyperscience.ai/latest/docs/creating-semi-structured-layouts) | ## N | **Named Entity Recognition Block (NER)** | A component in Hyperscience that automatically identifies and extracts specific types of information—like names, addresses, and organizations—from unstructured text. This block is used in tandem with full-page transcription to process documents containing freeform text. | [Entity Recognition Block](https://help.hyperscience.ai/latest/docs/entity-recognition-block) | | --- | --- | --- | | **Nested Table** | Data structure that represents a table within another table. It is embedded as a row into another table, creating a hierarchical or a “nested” structure. Nested tables allow you to extract data from tables with complicated structures where child-row data points inherit data points from parent rows. | [What is a Nested Table?](https://help.hyperscience.ai/latest/docs/what-is-a-nested-table) | | **Non-Structured Layout Classifier (NLC)** | Finds the correct Semi-structured or Additional layout for a given set of submission pages based on the words in the submitted documents. Note that NLC works on a page level. | [Semi-structured Document Classification](https://help.hyperscience.ai/latest/docs/semi-structured-document-classification) | | **Normalization** | The process of converting extracted data into a consistent format. In Hyperscience, normalization helps standardize values like dates, amounts, or addresses so they’re easier to use in downstream systems. | - [Default Data Types](https://help.hyperscience.ai/latest/docs/default-data-types) - [Creating Data Types with Custom Patterns](https://help.hyperscience.ai/latest/docs/creating-data-types-with-custom-patterns) | ## O | **Optical Intelligent Character Recognition (OICR)** | A technology that automatically identifies and converts printed or handwritten text within digital images into machine-encoded text. In Hyperscience, OICR is a key step that enables the platform to extract and work with text during document processing. | [Transcription Models Overview](https://help.hyperscience.ai/latest/docs/transcription-models-overview) | | --- | --- | --- | | **ORCA (Optical Reasoning and Cognition Agent)** | Hyperscience's Vision Language Model (VLM) framework for extracting information from semi-structured documents. ORCA supports multiple base models, such as ORCA 1.0 and ORCA 2. You can use these base models as-is or train specialized models on top of them to improve extraction performance for your specific use case. | [ORCA (Optical Reasoning and Cognition Agent) VLMs](https://help.hyperscience.ai/latest/docs/orca-optical-reasoning-and-cognition-agent-vlms) | | **Out-of-Memory (OOM) Error** | An out-of-memory error happens when a process in Hyperscience tries to use more memory than is available. It can interrupt processing and usually indicates that the job or document is too large or complex for the allocated resources. | [Memory Management](https://help.hyperscience.ai/deployment/docs/memory-management) | | **Out-of-the-Box (OOTB) Model** | A pre-trained machine learning model that comes ready to use in Hyperscience without needing additional training. It works well for common document types and use cases, helping teams get started quickly. | [ORCA (Optical Reasoning and Cognition Agent) VLMs](https://help.hyperscience.ai/latest/docs/orca-optical-reasoning-and-cognition-agent-vlms) | | **Output Blocks** | After processing, these blocks send extracted data to designated destinations like databases, applications, or other systems, ensuring seamless integration with existing workflows. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | ## P | **Page** | A logical entity in Hyperscience. A submission consists of one or more pages, and each page belongs to at most one document. Each page is analyzed individually for pre-processing and classification, and as part of a document for subsequent tasks (where applicable). | [How a File Becomes a Submission](https://help.hyperscience.ai/latest/docs/how-a-file-becomes-a-submission) | | --- | --- | --- | | **PDF Extraction** | A tool in the system that assists in creating layouts from PDFs by automatically suggesting field locations and names. PDF Extraction can only be used when creating the first variation of a layout. It cannot be used when creating subsequent layout variations. Only available for Structured layouts and is disabled by default. | [Application Settings Overview](https://help.hyperscience.ai/latest/docs/application-settings-overview) | | **Permission Group** | A group of users who share the same set of permissions. | [Permission Groups](https://help.hyperscience.ai/latest/docs/permission-groups) | | **Personally Identifiable Information (PII)** | Any data that can be used to identify an individual, such as a name, address, or Social Security number. | [PII Data Deletion](https://help.hyperscience.ai/latest/docs/pii-data-deletion) | | **PII Data Deletion** | A setting that allows you to remove PII from submissions to enhance security or comply with organizational policies. In Hyperscience, enabling this feature deletes all document image data, including extracted data and the original and processed images. | [PII Data Deletion](https://help.hyperscience.ai/latest/docs/pii-data-deletion) | | **PII Deletion Policy** | A configurable setting in Hyperscience that determines when PII is deleted based on either the submission-completion date or the date submitted. You can specify the number of days after the chosen date and set the exact time for deletion. | [PII Data Deletion](https://help.hyperscience.ai/latest/docs/pii-data-deletion) | | **Production Environment** | The live Hyperscience system, where real documents and data are processed as part of day-to-day business operations. This environment is used by end users and must meet high standards for performance, stability, and data security. It is distinct from development or testing environments. | [Deployment Information](https://help.hyperscience.ai/deployment/docs) | | **Projected Automation** | The predicted automation based on the desired target accuracy. The projection is derived from the model’s training data. The system automatically ensures that the same data is not used for both projections and training. | [Automation](https://help.hyperscience.ai/latest/docs/automation) | ## Q | **Quality Assurance (QA)** | Process that ensures the accuracy and reliability of system outputs. In Hyperscience, QA tasks allow users to review and correct errors in classification, identification, and transcription. Documents may be randomly sampled for QA from all processed data. | [What is Quality Assurance?](https://help.hyperscience.ai/latest/docs/what-is-quality-assurance) | | --- | --- | --- | | **Quality Assurance (QA) Records** | Results of **Quality Assurance (QA)** tasks in Hyperscience. These records are used to evaluate accuracy and help the system calculate the QA sample rate, which determines how many fields are selected for review to ensure consistent data quality. | [Transcription Settings](https://help.hyperscience.ai/latest/docs/transcription-settings) | ## R | **Recommended Number of Documents** | The suggested minimum and optimal amount of labeled data required to train machine learning models in Hyperscience. - For **Classification models**, at least **10 pages per layout** are needed, with **120 pages per layout** recommended. - For **Identification models**, the minimum is **100 documents**, and the recommended amount is **400 documents**. Providing diverse, high-quality examples ensures better model accuracy and reliability. | - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) - [TDM for Identification Models](https://help.hyperscience.ai/latest/docs/tdm-for-identification-models) | | --- | --- | --- | | **Release** | A release is a package of one or more committed layout variation versions. This collection of layouts or layout variations should reflect the various document types that the machine should expect. To use the layouts that you have created to process documents, you will need to deploy a flow that contains a release. | [What is a Release?](https://help.hyperscience.ai/latest/docs/what-is-a-release) | | **Reprocessing Block** | When initiated from other Supervision tasks, the Document Classification task allows users to manually classify machine-misclassified documents. As a result, the users can reclassify all pages of the submission and submit them for reprocessing. | [Document Classification Task](https://help.hyperscience.ai/latest/docs/document-classification-task) | | **Required Number of Documents** | The minimum amount of annotated data needed to train machine learning models in Hyperscience. - For **Classification models**, at least **10 pages per layout** are required. - For **Identification models**, the system requires a **minimum of 100 documents**. Meeting these thresholds ensures the models can be successfully trained and begin learning layout- or document-type patterns. | - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) - [TDM for Identification Models](https://help.hyperscience.ai/latest/docs/tdm-for-identification-models) | | **Resubmitting a Submission** | In Hyperscience, the process of reprocessing a previously submitted document using the same flow and configuration. This action is typically used to address issues or errors encountered during the initial processing. "Resubmitting" and "retrying" are used interchangeably and retain the original submission ID. | [Testing and Debugging Flows](https://help.hyperscience.ai/latest/docs/testing-and-debugging-flows) | | **Retrying a Submission** | In Hyperscience, the process of reprocessing a previously submitted document using the same flow and configuration. This action is typically used to address issues or errors encountered during the initial processing. "Resubmitting" and "retrying" are used interchangeably and retain the original submission ID. | [Testing and Debugging Flows](https://help.hyperscience.ai/latest/docs/testing-and-debugging-flows) | | **Routing Blocks** | Blocks that direct documents or data through specific paths in the flow based on predefined conditions, ensuring each item follows the appropriate processing route. ​ | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Row** | A single horizontal grouping of field values extracted by the Table Identification model within a tabular region of a document. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | ## S | **Segment** | A distinct region on a document image that contains text, as identified by the Segmentation model. Segments are the building blocks used for downstream tasks like Classification and Transcription. Each segment includes positional data (in the form of a bounding box) and text content, helping the system understand the document’s layout and structure. | [Text Segmentation](https://help.hyperscience.ai/latest/docs/text-segmentation) | | --- | --- | --- | | **Segmentation** | The process of partitioning an image into regions containing text. It is the first step of downstream processing tasks such as Classification and Transcription. | [Text Segmentation](https://help.hyperscience.ai/latest/docs/text-segmentation) | | **Semi-structured Layout** | A configuration within Hyperscience designed to process documents where fields and table cells are present, but their positions can vary among documents. Unlike Structured layouts, Semi-structured layouts do not rely on fixed field locations. Instead, a model is trained to find the fields and cells based on provided training examples. | [Creating Semi-structured layouts](https://help.hyperscience.ai/latest/docs/creating-semi-structured-layouts) | | **Signature Field** | A non-text field data type used to detect and extract handwritten signatures from documents. In Hyperscience, the Signature data type allows the system to identify areas where a person has signed, facilitating the extraction of these signatures for verification or record-keeping purposes. Compatible with both Structured and Semi-structured layouts. | [Checkboxes and Signatures](https://help.hyperscience.ai/latest/docs/checkboxes-and-signatures) | | **Software-Defined Management (SDM)** | An approach to managing infrastructure and operations through software-based controls rather than manual or hardware-specific methods. In Hyperscience, this approach enables centralized, flexible configuration and orchestration of workflows, environments, and resources, supporting automation, scalability, and easier maintenance. | | | **Software Development Kit (SDK)** | A collection of tools, libraries, and documentation that developers use to build software applications for a specific platform or system. In Hyperscience, the SDK allows teams to create custom integrations, automate tasks, or extend platform functionality—such as building Code Blocks or interacting with the API—within a developer-friendly framework. | [Custom Supervision](https://help.hyperscience.ai/latest/docs/custom-supervision) | | **Specialized models** | A model created by training on customer-specific, annotated documents on top of a base model. It is tailored to a particular layout and use case to improve extraction accuracy and automation. | [Training a specialized model](https://help.hyperscience.ai/v43/docs/training-a-specialized-model) | | **Straight-Through Processing (STP)** | Any data that passes through the system without human intervention. When used as a metric, STP allows you to measure the number of documents that pass through our system without human review. However, STP is not recommended, as it may bypass critical quality checks. Human-in-the-loop validation ensures accuracy, especially for high-impact or sensitive data, and helps maintain trust and compliance in real-world operations. HITL also ensures optimal average handling time ([AHT](/general-information/docs/glossary#a)) by flagging values for review *only* when needed. This can be further optimized by setting appropriate [Field-Level Accuracy Targets (FLAT)](/general-information/docs/glossary#f) as needed. | [Automation](https://help.hyperscience.ai/latest/docs/automation) | | **Structured Layout** | In Hyperscience, a Structured layout is used to process documents where key fields appear in consistent positions. This layout type helps the system quickly and accurately find and extract data from standardized documents like W-8 or HCFA forms in the US. | [Creating Structured Layouts](https://help.hyperscience.ai/latest/docs/creating-structured-layouts) | | **Subflow** | A smaller flow that's part of a larger flow group. It's used to break down complex flows into reusable pieces. When you deploy a flow group, all its subflows are deployed together, making it easier to manage and reuse common processing steps across different workflows. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | **Submission** | A set of files uploaded together for processing. The system interprets each file as a page, matches it to existing layouts, and groups it into documents based on these matches. | - [What is a Submission?](https://help.hyperscience.ai/latest/docs/what-is-a-submission) - [How a File Becomes a Submission?](https://help.hyperscience.ai/latest/docs/how-a-file-becomes-a-submission) | | **Submission Bootstrap** | In Hyperscience, the Submission Bootstrap refers to the Submission Initialization Block within a document-processing flow. This block manages the initial setup and configuration for incoming submissions, including data-ingestion parameters. By configuring the Submission Bootstrap, you can control how submissions are initialized. | [Document Processing Subflow Settings](https://help.hyperscience.ai/latest/docs/document-processing-subflow-settings) | | **Submission Initialization Block** | The first step in a document-processing flow is to set up the submission by determining how documents are grouped and where they are ingested from. | [Flow Blocks](https://help.hyperscience.ai/latest/docs/flow-blocks) | | **Supervision** | A manual task that is created when the system’s confidence in a prediction is below the confidence threshold. Supervision allows a human to review and correct the output, ensuring data accuracy through human-in-the-loop input. | [What is Supervision?](https://help.hyperscience.ai/latest/docs/what-is-supervision) | ## T | **Table** | A logical structure used to organize and present information in rows and columns. It is used to present values in a readable format. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | --- | --- | --- | | **Table Identification Quality Assurance** | The process of reviewing and validating the system's identification of tables within Semi-structured documents. This quality-assurance task allows you to measure the system’s accuracy in table identification. | [Table Identification Quality Assurance](https://help.hyperscience.ai/latest/docs/table-identification-quality-assurance) | | **Table Identification Task** | A manual task in Hyperscience where you confirm or correct the location of tables in Semi-structured documents. You may adjust or draw bounding boxes around the table cells to help your model learn where to look for the data you want to extract. Rows are represented by horizontal separators. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | **Table Locator** | A machine learning model in Hyperscience that learns where cells and rows are located in Semi-structured documents. It uses examples from training to predict the position of each cell and row on a page so the system can extract the correct data. | [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Target Accuracy** | A setting specified by the user. It indicates the desired overall system accuracy, including tasks performed by humans. It allows you to evaluate how well the system is expected to perform. See also [Field-Level Accuracy Targets (FLAT)](/general-information/docs/glossary#f). | - [Accuracy](https://help.hyperscience.ai/latest/docs/accuracy) - [Automation](https://help.hyperscience.ai/latest/docs/automation) - [Identification Settings](https://help.hyperscience.ai/latest/docs/identification-settings) | | **Task** | A step in processes where a human reviews or confirms part of the system’s output. Tasks are generated based on the system settings when the system is uncertain about a result or needs to check accuracy. There are two main types: - **Supervision Tasks** - **Quality Assurance (QA) Tasks** | - [What are Tasks?](https://help.hyperscience.ai/latest/docs/what-are-tasks) - [What is Supervision?](https://help.hyperscience.ai/latest/docs/what-is-supervision) - [What is Quality Assurance?](https://help.hyperscience.ai/latest/docs/what-is-quality-assurance) | | **Task Queue** | A list of tasks waiting for human review. It organizes and stores all active Supervision and Quality Assurance tasks that need to be completed. | [Task Queue Tab](https://help.hyperscience.ai/latest/docs/task-queue-tab) | | **Task Restriction** | An access-control rule that limits which users can perform specific Supervision or Quality Assurance tasks. Task restrictions use permission groups and can be applied at the submission, layout, or flow-block level. | [Task restrictions and permission groups](https://help.hyperscience.ai/latest/docs/task-restrictions-and-permission-groups) | | **Template Row** | The lead row in your table. It doesn’t need to be the first one, but it should be representative of the rows in your table. Hyperscience uses the Copycat tool to populate the annotation for the rest of the rows in your table. The Copycat is not always accurate, so make sure to double-check the annotations. | [Table Identification](https://help.hyperscience.ai/latest/docs/table-identification) | | **Text Classification** | A machine learning model that reads unstructured text—like comments, emails, or notes, and assigns them to predefined categories. This categorization helps automate decisions and organize freeform text based on business rules. | [Text Classification](https://help.hyperscience.ai/latest/docs/text-classification) | | **Threshold** | The confidence limit used to decide if a machine prediction should be sent for human review to ensure accuracy. | - [Accuracy](https://help.hyperscience.ai/latest/docs/accuracy) - [Automation](https://help.hyperscience.ai/latest/docs/automation) - [Transcription Settings](https://help.hyperscience.ai/latest/docs/transcription-settings) - [Identification Settings](https://help.hyperscience.ai/latest/docs/identification-settings) | | **Top-Level Flow** | The main flow that manages the end-to-end document processing, coordinating with subflows to handle specific components of the process. | [Flows Overview](https://help.hyperscience.ai/latest/docs/flows-overview) | | **Trainer** | A separate machine dedicated to handling resource-heavy tasks like training Identification models. It operates independently and connects to the main application through the API. | [Trainer](https://help.hyperscience.ai/latest/docs/trainer) | | **Training Data** | The input used to teach machine learning models how to process documents accurately. Its structure depends on the model type: - For Classification models, training data consists of uploaded document pages grouped by layout. - For Identification models, training data includes manually annotated fields and tables to train the model to extract specific data points. | - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) - [TDM for Identification Models](https://help.hyperscience.ai/latest/docs/tdm-for-identification-models) | | **Training Data Analysis** | A tool in TDM that analyzes your training data to compute the importance of each training document and identify issues such as missing labels, overlapping fields or columns, or inconsistent annotations. This analysis helps you prioritize which documents to annotate and ensures clean, accurate data before you train a model. | - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) - [TDM for Identification Models](https://help.hyperscience.ai/latest/docs/tdm-for-identification-models) | | **Training Data Curator** | A tool in Hyperscience that helps you select the most valuable documents for training your model. It highlights documents that are likely to improve accuracy so you can focus your annotation efforts where they matter most. | [Training Data Curator](https://help.hyperscience.ai/latest/docs/training-data-curator) | | **Training Data Management (TDM)** | A tool used to annotate, manage, import, and export training documents. It is also used to train models by working directly with the training data (“ground truth”) obtained from each document in the training set. | [Training Data Management](https://help.hyperscience.ai/latest/docs/training-data-management) | | **Training Set** | A dataset used to teach the system how to recognize and extract information. It includes documents with labeled fields so the system can learn from real examples. | [Training Data Management](https://help.hyperscience.ai/latest/docs/training-data-management) | | **Transcription Task** | A Supervision task that allows you to review or enter text the system couldn’t confidently read from a document. This task enables you to ensure accurate final data when the system’s confidence is low. | [Transcription](https://help.hyperscience.ai/latest/docs/transcription) | | **Transcription Model** | A machine learning model that automatically extracts text from scanned document images. It supports both printed and handwritten text. When the model’s confidence in the extracted text is low, the system generates a Transcription task for human review to ensure accuracy. | [Transcription Models Overview](https://help.hyperscience.ai/latest/docs/transcription-models-overview) | | **Transcription Quality Assurance Task** | A type of QA task where a set percentage of transcribed fields or table cells are reviewed by a human. It helps measure and improve the accuracy of both machine and human transcriptions. | [Transcription Quality Assurance](https://help.hyperscience.ai/latest/docs/transcription-quality-assurance) | | **True Positive** | A result where the model correctly predicts something as positive, for example, it says a field is present, and it is present. | [Scoring Field Identification Accuracy](https://help.hyperscience.ai/latest/docs/scoring-field-identification-accuracy) | | **Technical Validation Event (TVE)** | A trial phase where a potential customer tests the platform with clear success criteria. It’s designed to show that the product works well for their needs. This phase ensures that both sides are aligned before moving forward. | [TVE (POC) Installation Instructions](https://help.hyperscience.ai/deployment/docs/tve-poc-installation-instructions) | ## U | **Unmatched Document** | A set of pages that the system couldn’t match to any known layout during Classification. | - [TDM for Classification Models](https://help.hyperscience.ai/latest/docs/tdm-for-classification-models) - [Document Drift Management (Layout Triage)](https://help.hyperscience.ai/latest/docs/document-drift-management-layout-triage) - [Semi-structured Document Classification](https://help.hyperscience.ai/latest/docs/semi-structured-document-classification) | | --- | --- | --- | | **Unstructured Extraction** | Processing documents with little to no layout consistency. Key information appears anywhere, often embedded in long paragraphs or freeform text. Examples include contracts, title deeds, and annual reports. | - [Understanding Document Types](https://help.hyperscience.ai/latest/docs/understanding-document-types) - [Long-form Extraction](https://help.hyperscience.ai/latest/docs/long-form-extraction) - [Field Identification](https://help.hyperscience.ai/latest/docs/field-identification) - [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | ## V | **Vision Language Models (VLM)** | Models that understand both the text and images in a document. They combine what’s written with where it appears on the page to help the system read and extract information more accurately. | [ORCA (Optical Reasoning and Cognition Agent) VLMs](https://help.hyperscience.ai/latest/docs/orca-optical-reasoning-and-cognition-agent-vlms) | | --- | --- | --- | | **Vendor** | A third-party entity that does business with your company and sends you documents. The meaning of the data can vary depending on the vendor. | [Training an Identification Model](https://help.hyperscience.ai/latest/docs/training-an-identification-model) | | **Visual Page Classifier (VPC)** | An automated component that matches submission pages to the correct layouts from the Layout Library. It ensures accurate and efficient processing of Structured documents. | [Structured Document Classification](https://help.hyperscience.ai/latest/docs/structured-document-classification) |