- 概要
- Document Processing Contracts
- リリース ノート
- Document Processing Contracts について
- Box クラス
- IPersistedActivity インターフェイス
- PrettyBoxConverter クラス
- IClassifierActivity インターフェイス
- IClassifierCapabilitiesProvider インターフェイス
- ClassifierDocumentType クラス
- ClassifierResult クラス
- ClassifierCodeActivity クラス
- ClassifierNativeActivity クラス
- ClassifierAsyncCodeActivity クラス
- ClassifierDocumentTypeCapability クラス
- ContentValidationData クラス
- EvaluatedBusinessRulesForFieldValue クラス
- EvaluatedBusinessRuleDetails クラス
- ExtractorAsyncCodeActivity クラス
- ExtractorCodeActivity クラス
- ExtractorDocumentType クラス
- ExtractorDocumentTypeCapabilities クラス
- ExtractorFieldCapability クラス
- ExtractorNativeActivity クラス
- ExtractorResult クラス
- FieldValue クラス
- FieldValueResult クラス
- ICapabilitiesProvider インターフェイス
- IExtractorActivity インターフェイス
- ExtractorPayload クラス
- DocumentActionPriority 列挙型
- DocumentActionData クラス
- DocumentActionStatus 列挙型
- DocumentActionType 列挙型
- DocumentClassificationActionData クラス
- DocumentValidationActionData クラス
- UserData クラス
- Document クラス
- DocumentSplittingResult クラス
- DomExtensions クラス
- Page クラス
- PageSection クラス
- Polygon クラス
- PolygonConverter クラス
- Metadata クラス
- WordGroup クラス
- Word クラス
- ProcessingSource 列挙型
- ResultsTableCell クラス
- ResultsTableValue クラス
- ResultsTableColumnInfo クラス
- ResultsTable クラス
- Rotation 列挙型
- ルール クラス
- RuleResult クラス
- RuleSet クラス
- RuleSetResult クラス
- SectionType 列挙型
- WordGroupType 列挙型
- IDocumentTextProjection インターフェイス
- ClassificationResult クラス
- ExtractionResult クラス
- ResultsDocument クラス
- ResultsDocumentBounds クラス
- ResultsDataPoint クラス
- ResultsValue クラス
- ResultsContentReference クラス
- ResultsValueTokens クラス
- ResultsDerivedField クラス
- ResultsDataSource 列挙型
- ResultConstants クラス
- SimpleFieldValue クラス
- TableFieldValue クラス
- DocumentGroup クラス
- DocumentTaxonomy クラス
- DocumentType クラス
- Field クラス
- FieldType 列挙型
- FieldValueDetails クラス
- LanguageInfo クラス
- MetadataEntry クラス
- TextType 列挙型
- TypeField クラス
- ITrackingActivity インターフェイス
- ITrainableActivity インターフェイス
- ITrainableClassifierActivity インターフェイス
- ITrainableExtractorActivity インターフェイス
- TrainableClassifierAsyncCodeActivity クラス
- TrainableClassifierCodeActivity クラス
- TrainableClassifierNativeActivity クラス
- TrainableExtractorAsyncCodeActivity クラス
- TrainableExtractorCodeActivity クラス
- TrainableExtractorNativeActivity クラス
- BasicDataPoint クラス
- BasicValue クラス
- ComponentCollectionFacade クラス
- DataPointFacade基本クラス
- ExtractionResultHandler クラス
- FieldGroupDataPoint クラス
- FieldGroupValue クラス
- FieldLookup 基本クラス
- FieldRedactionSettings クラス
- RedactionOptions クラス (プレビュー)
- RedactionType 列挙型
- ResultsValueFacadeBase クラス
- TableDataPoint クラス
- TableRow クラス
- TableValue クラス
- WildcardDataPoint クラス
- WildcardDataPointCollection クラス
- Document Understanding ML
- Document Understanding OCR ローカル サーバー
- Document Understanding
- リリース ノート
- Document Understanding アクティビティ パッケージについて
- プロジェクトの対応 OS
- 外部接続を設定する
- Document Understanding のコード化されたオートメーション API (プレビュー)
- IntelligentOCR
- リリース ノート
- IntelligentOCR アクティビティ パッケージについて
- プロジェクトの対応 OS
- タクソノミーを読み込み
- ドキュメントをデジタル化
- ドキュメント分類スコープ
- キーワード ベースの分類器
- Document Understanding プロジェクト分類器
- インテリジェント キーワード分類器
- ドキュメント分類アクションを作成
- ドキュメント検証成果物を作成
- ドキュメント検証成果物を取得
- ドキュメント分類アクション完了まで待機し再開
- 分類器トレーニング スコープ
- キーワード ベースの分類器トレーナー
- インテリジェント キーワード分類器トレーナー
- データ抽出スコープ
- Document Understanding プロジェクト抽出器
- Document Understanding プロジェクト抽出器トレーナー
- 正規表現ベースの抽出器
- フォーム抽出器
- インテリジェント フォーム抽出器
- ドキュメントを墨消し
- ドキュメント検証アクションを作成
- ドキュメント検証アクション完了まで待機し再開
- 抽出器トレーニング スコープ
- 抽出結果をエクスポート
- マシン ラーニング抽出器
- マシン ラーニング抽出器トレーナー
- マシン ラーニング分類器
- マシン ラーニング分類器トレーナー
- 生成 AI 分類器
- 生成 AI 抽出器
- 認証を構成する
- ML サービス
- OCR
- OCR Contracts
- リリース ノート
- OCR コントラクトについて
- プロジェクトの対応 OS
- IOCRActivity インターフェイス
- OCRAsyncCodeActivity クラス
- OCRCodeActivity クラス
- OCRNativeActivity クラス
- Character クラス
- OCRResult クラス
- Word クラス
- FontStyles 列挙型
- OCRRotation 列挙型
- OCRCapabilities クラス
- OCRScrapeBase クラス
- OCRScrapeFactory クラス
- ScrapeControlBase クラス
- ScrapeEngineUsages 列挙型
- ScrapeEngineBase
- ScrapeEngineFactory クラス
- ScrapeEngineProvider クラス
- OmniPage
- PDF
- [リストから削除済] ABBYY
- [リストから削除済] ABBYY Embedded
Document Understanding アクティビティ パッケージ用のコード化されたオートメーション API です。ドキュメントの分類、データ抽出、検証成果物に対応します。
UiPath.DocumentUnderstanding.Activities
ドキュメントの分類、構造化データの抽出、ドキュメント検証成果物の作成と取得のためのコード化されたワークフロー API です。これらの API は、コード化されたオートメーションを設計する際に使用できます。コード化されたオートメーションの概要と、API を使用してそれらのオートメーションを設計する方法については、「 コード化されたオートメーション」をご覧ください。
- サービス アクセサ:
du(IDocumentUnderstandingService型) - 必要なパッケージ: 依存関係
project.json"UiPath.DocumentUnderstanding.Activities": "*"。
自動インポートされる名前空間
Document Understanding パッケージをインストールすると、コード化されたワークフローで以下の名前空間が自動的に利用可能になります。
UiPath.DocumentUnderstanding.Activities.ApiUiPath.Platform.ResourceHandlingUiPath.DocumentProcessing.Contracts.ActionsUiPath.IntelligentOCR.StudioWeb.Activities.DataExtractionUiPath.IntelligentOCR.StudioWeb.Activities.DocumentClassification
サービスの概要
du サービスでは、Document Understanding の各操作が直接メソッド呼び出しとして公開されます。開くコネクションまたはスコープがありません。サービス アクセサーのメソッドを直接呼び出します。
var extracted = du.ExtractDocumentData(@"C:\invoices\invoice.pdf", "Invoices", "Production", "invoice");
var extracted = du.ExtractDocumentData(@"C:\invoices\invoice.pdf", "Invoices", "Production", "invoice");
このサービスは、Studio Web の [ドキュメントを分類]、[ ドキュメント データを抽出]、[ ドキュメント検証成果物を作成]、および [ ドキュメント検証成果物を取得 ] の各アクティビティをミラーリングします。
プロジェクトのバージョンまたはタグ
projectVersionOrTag パラメーターは、Studio の単一の [バージョン] ドロップダウンにマップされます。バージョン名 (例: v3) またはタグ (例: Production、 Staging、 live) を渡します。ランタイムは、プロジェクトに一致する種類を解決します。[Predefined] プロジェクトでは、 を使用します Production。
共通パラメーター
timeoutMs(Int) - 分類と抽出のタイムアウト (ミリ秒単位)既定値は3600000(1 時間) です。docType(String) - ドキュメントの種類の ID です。複数のドキュメントの種類を含むプロジェクトにはドキュメントの種類名を渡します。バージョンごとに抽出器が 1 つである IXP スタイルのプロジェクトには空の文字列を渡します。
分類
DocumentData ClassifyDocument(string documentPath, string projectName, string projectVersionOrTag, int timeoutMs = 3600000)
ローカル ファイル パスのドキュメントを Document Understanding プロジェクトに分類します。予測されたドキュメントの種類を含む DocumentData オブジェクトを返します。
DocumentData ClassifyDocument(IResource file, string projectName, string projectVersionOrTag, int timeoutMs = 3600000)
指定されたドキュメント リソースを分類します。[パスの存在を確認] や [ローカル ファイル/フォルダーを取得] アクティビティの出力など、任意のIResourceを受け入れます。
データ抽出
IDocumentData<DictionaryData> ExtractDocumentData(string documentPath, string projectName, string projectVersionOrTag, string docType, int timeoutMs = 3600000)
ローカル ファイル パスにあるドキュメントから構造化データを抽出します。抽出されたフィールドは IDocumentData<DictionaryData>として返されます。
IDocumentData<DictionaryData> ExtractDocumentData(IResource file, string projectName, string projectVersionOrTag, string docType, int timeoutMs = 3600000)
指定されたドキュメント リソースから構造化データを抽出します。
IDocumentData<DictionaryData> ExtractDocumentData(DocumentData classifiedDocument, string projectName, string projectVersionOrTag, string docType = null, int timeoutMs = 3600000)
すでに ClassifyDocumentで分類されているドキュメントから構造化データを抽出します。これにより、ファイルが再デジタル化されないようになります。docType がnullの場合、分類されたドキュメントのドキュメントの種類から派生します。
検証の成果物
ContentValidationData CreateDocumentValidationArtifacts(IDocumentData<DictionaryData> automaticExtractionResults, string orchestratorFolderName, string orchestratorBucketName = null)
抽出結果を Orchestrator のストレージにアップロードし、 ContentValidationData ハンドルを返します。ハンドルを使用して検証アクションを作成するか、後で を使用して結果を取得する RetrieveDocumentValidationArtifactsを使用します。orchestratorBucketName が nullの場合、既定のバケットが使用されます。
IDocumentData<DictionaryData> RetrieveDocumentValidationArtifacts(ContentValidationData contentValidationData, object completedAppAction = null, bool removeDataFromStorage = false, bool returnAutomaticExtractionResults = false)
検証済みの抽出結果をストレージから取得します。removeDataFromStorage がtrueの場合、ストレージ成果物は取得後に削除されます。returnAutomaticExtractionResults がtrueの場合、検証済みの結果ではなく元の自動抽出結果が返されます。
アクティビティとの関係
コード化された API アクティビティと Document Understanding XAML アクティビティは同じランタイムで実行されるため、コード化されたワークフローは、[ ドキュメント データを抽出 ] アクティビティおよび [ ドキュメントを分類 ] アクティビティと同じ分類、抽出、検証成果物の動作になります。違いは呼び出しの形状です。このサービスはプレーンなファイル パスまたはIResource入力を受け取り、Document Data オブジェクトを直接返します。また、アクティビティの [エラー発生時に実行を継続] オプションの代わりに通常のtry/catchを使用します。
一般的なパターン
分類してから抽出
このパターンでは、ドキュメントを分類し、分類された結果を再利用してデータを抽出するため、ファイルを 2 度目にデジタル化するのを回避できます。
[Workflow]
public void Execute()
{
var classified = du.ClassifyDocument(@"C:\docs\file.pdf", "MyProject", "Production");
Log($"Document type: {classified.DocumentType.DisplayName}");
// Reuses the classified document, so the file is not digitized again.
var extracted = du.ExtractDocumentData(classified, "MyProject", "Production");
}
[Workflow]
public void Execute()
{
var classified = du.ClassifyDocument(@"C:\docs\file.pdf", "MyProject", "Production");
Log($"Document type: {classified.DocumentType.DisplayName}");
// Reuses the classified document, so the file is not digitized again.
var extracted = du.ExtractDocumentData(classified, "MyProject", "Production");
}
検証コントロールを含むアプリ タスクのデータを抽出して準備する
このパターンは、データを抽出し、アプリ タスクで確認するために検証成果物としてストレージにアップロードし、アクションの完了後に検証済みの結果を取得します。
[Workflow]
public void Execute()
{
var extracted = du.ExtractDocumentData(@"C:\docs\invoice.pdf", "Invoices", "Production", "invoice");
// Upload the results so they can be reviewed in Action Center.
var artifacts = du.CreateDocumentValidationArtifacts(extracted, "Shared");
// After the Action Center validation action completes, retrieve the validated results.
var validated = du.RetrieveDocumentValidationArtifacts(artifacts, removeDataFromStorage: true);
}
[Workflow]
public void Execute()
{
var extracted = du.ExtractDocumentData(@"C:\docs\invoice.pdf", "Invoices", "Production", "invoice");
// Upload the results so they can be reviewed in Action Center.
var artifacts = du.CreateDocumentValidationArtifacts(extracted, "Shared");
// After the Action Center validation action completes, retrieve the validated results.
var validated = du.RetrieveDocumentValidationArtifacts(artifacts, removeDataFromStorage: true);
}
- 自動インポートされる名前空間
- サービスの概要
- プロジェクトのバージョンまたはタグ
- 共通パラメーター
- 分類
DocumentData ClassifyDocument(string documentPath, string projectName, string projectVersionOrTag, int timeoutMs = 3600000)DocumentData ClassifyDocument(IResource file, string projectName, string projectVersionOrTag, int timeoutMs = 3600000)- データ抽出
IDocumentData<DictionaryData> ExtractDocumentData(string documentPath, string projectName, string projectVersionOrTag, string docType, int timeoutMs = 3600000)IDocumentData<DictionaryData> ExtractDocumentData(IResource file, string projectName, string projectVersionOrTag, string docType, int timeoutMs = 3600000)IDocumentData<DictionaryData> ExtractDocumentData(DocumentData classifiedDocument, string projectName, string projectVersionOrTag, string docType = null, int timeoutMs = 3600000)- 検証の成果物
ContentValidationData CreateDocumentValidationArtifacts(IDocumentData<DictionaryData> automaticExtractionResults, string orchestratorFolderName, string orchestratorBucketName = null)IDocumentData<DictionaryData> RetrieveDocumentValidationArtifacts(ContentValidationData contentValidationData, object completedAppAction = null, bool removeDataFromStorage = false, bool returnAutomaticExtractionResults = false)- アクティビティとの関係
- 一般的なパターン
- 分類してから抽出
- 検証コントロールを含むアプリ タスクのデータを抽出して準備する