Document Understanding modern projects user guide

适用平台：

上次更新日期 2026年4月6日

基本功能

要自动化文档处理，需要四项基本功能：数字化、分类、提取和验证。

Figure 1. Fundamental capabilities 描述 Document Understanding 基本功能的示意图。系统首先要将文档数字化，然后将其分类，最后执行提取。分类和提取都有一个额外的验证步骤。

数字化

数字化将物理文档转换为机器可读文本，然后可以对文本进行数字化处理。尽管光学字符识别 (OCR) 是数字化的重要组成部分，但数字化流程更加复杂，涉及各个步骤，包括 OCR。

例如，在处理 PDF 文档时，数字化算法可以区分扫描 PDF 和原生 PDF，或者包含扫描图像和原生文本的混合 PDF。大多数文本可以直接从原生 PDF 文档中提取，但在某些情况下，可能需要使用 OCR 读取一些徽标。数字化流程可以处理所有这些情况，以确保文本检测具有最高的准确性，同时快速高效地运行。

You can change the OCR used in your project from Project settings. For more information, check the Configure project settings page. You can check the available OCR engines and the supported languages from the Supported languages section of the user guide.

You can check the Known limitations page for more information on the supported files, image size limits, and more specifications.

分类

分类的目的是扫描文档并确定其所属的文档类型。了解文档的类型非常重要，因为不同的文档类型需要不同的处理技术。例如，发票需要由发票提取模型处理，以确保提取所有相关字段。

Figure 2. Document classifier 该图像描述未知文档类型的文档如何通过文档分类器。之后，该文档将被分类为发票。

提取

Data extraction is the process of selecting and retrieving only the relevant information from a document. Extracting specific data from a lengthy document using string manipulation can be challenging. However, Document Understanding^TM provides various extraction methodologies for different document types and formats. For example, we only want to extract the Vendor Name, Billing Name, Due Date, and Total fields from an invoice.

Figure 3. Data extraction 描述如何从发票中提取数据的映像。提取的字段包括“供应商名称”、“账单名称”、“到期日期”和“总计”。

验证

在分类和提取中，软件机器人使用置信度概念，该概念用于衡量良好执行特定任务的确定性级别。此任务可能是识别文档类型、识别字段或读取其中的数据。在这些情况下，Document Understanding 框架允许您让人类用户来审核和验证机器人的输出。在最佳情况下，系统会使用人工输入，通过机器学习来训练机器人的准确性。

在此页面上

数字化
分类
提取
验证

此页面有帮助吗？

前一个文档类型

下一个关键概念

Document Understanding modern projects user guide

数字化​

分类​

提取​

验证​

此页面有帮助吗？

数字化

分类

提取

验证