- 概要
- Document Processing Contracts
- リリース ノート
- Document Processing Contracts について
- Box クラス
- IPersistedActivity インターフェイス
- PrettyBoxConverter クラス
- IClassifierActivity インターフェイス
- IClassifierCapabilitiesProvider インターフェイス
- ClassifierDocumentType クラス
- ClassifierResult クラス
- ClassifierCodeActivity クラス
- ClassifierNativeActivity クラス
- ClassifierAsyncCodeActivity クラス
- ClassifierDocumentTypeCapability クラス
- ContentValidationData クラス
- EvaluatedBusinessRulesForFieldValue クラス
- EvaluatedBusinessRuleDetails クラス
- ExtractorAsyncCodeActivity クラス
- ExtractorCodeActivity クラス
- ExtractorDocumentType クラス
- ExtractorDocumentTypeCapabilities クラス
- ExtractorFieldCapability クラス
- ExtractorNativeActivity クラス
- ExtractorResult クラス
- FieldValue クラス
- FieldValueResult クラス
- ICapabilitiesProvider インターフェイス
- IExtractorActivity インターフェイス
- ExtractorPayload クラス
- DocumentActionPriority 列挙型
- DocumentActionData クラス
- DocumentActionStatus 列挙型
- DocumentActionType 列挙型
- DocumentClassificationActionData クラス
- DocumentValidationActionData クラス
- UserData クラス
- Document クラス
- DocumentSplittingResult クラス
- DomExtensions クラス
- Page クラス
- PageSection クラス
- Polygon クラス
- PolygonConverter クラス
- Metadata クラス
- WordGroup クラス
- Word クラス
- ProcessingSource 列挙型
- ResultsTableCell クラス
- ResultsTableValue クラス
- ResultsTableColumnInfo クラス
- ResultsTable クラス
- Rotation 列挙型
- ルール クラス
- RuleResult クラス
- RuleSet クラス
- RuleSetResult クラス
- SectionType 列挙型
- WordGroupType 列挙型
- IDocumentTextProjection インターフェイス
- ClassificationResult クラス
- ExtractionResult クラス
- ResultsDocument クラス
- ResultsDocumentBounds クラス
- ResultsDataPoint クラス
- ResultsValue クラス
- ResultsContentReference クラス
- ResultsValueTokens クラス
- ResultsDerivedField クラス
- ResultsDataSource 列挙型
- ResultConstants クラス
- SimpleFieldValue クラス
- TableFieldValue クラス
- DocumentGroup クラス
- DocumentTaxonomy クラス
- DocumentType クラス
- Field クラス
- FieldType 列挙型
- FieldValueDetails クラス
- LanguageInfo クラス
- MetadataEntry クラス
- TextType 列挙型
- TypeField クラス
- ITrackingActivity インターフェイス
- ITrainableActivity インターフェイス
- ITrainableClassifierActivity インターフェイス
- ITrainableExtractorActivity インターフェイス
- TrainableClassifierAsyncCodeActivity クラス
- TrainableClassifierCodeActivity クラス
- TrainableClassifierNativeActivity クラス
- TrainableExtractorAsyncCodeActivity クラス
- TrainableExtractorCodeActivity クラス
- TrainableExtractorNativeActivity クラス
- BasicDataPoint クラス
- BasicValue クラス
- ComponentCollectionFacade クラス
- DataPointFacade基本クラス
- ExtractionResultHandler クラス
- FieldGroupDataPoint クラス
- FieldGroupValue クラス
- FieldLookup 基本クラス
- FieldRedactionSettings クラス
- RedactionOptions クラス (プレビュー)
- RedactionType 列挙型
- ResultsValueFacadeBase クラス
- TableDataPoint クラス
- TableRow クラス
- TableValue クラス
- WildcardDataPoint クラス
- WildcardDataPointCollection クラス
- Document Understanding ML
- Document Understanding OCR ローカル サーバー
- Document Understanding
- IntelligentOCR
- リリース ノート
- IntelligentOCR アクティビティ パッケージについて
- プロジェクトの対応 OS
- タクソノミーを読み込み
- ドキュメントをデジタル化
- ドキュメント分類スコープ
- キーワード ベースの分類器
- Document Understanding プロジェクト分類器
- インテリジェント キーワード分類器
- ドキュメント分類アクションを作成
- ドキュメント検証成果物を作成
- ドキュメント検証成果物を取得
- ドキュメント分類アクション完了まで待機し再開
- 分類器トレーニング スコープ
- キーワード ベースの分類器トレーナー
- インテリジェント キーワード分類器トレーナー
- データ抽出スコープ
- Document Understanding プロジェクト抽出器
- Document Understanding プロジェクト抽出器トレーナー
- 正規表現ベースの抽出器
- フォーム抽出器
- インテリジェント フォーム抽出器
- ドキュメントを墨消し
- ドキュメント検証アクションを作成
- ドキュメント検証アクション完了まで待機し再開
- 抽出器トレーニング スコープ
- 抽出結果をエクスポート
- マシン ラーニング抽出器
- マシン ラーニング抽出器トレーナー
- マシン ラーニング分類器
- マシン ラーニング分類器トレーナー
- 生成 AI 分類器
- 生成 AI 抽出器
- 認証を構成する
- ML サービス
- OCR
- OCR Contracts
- リリース ノート
- OCR コントラクトについて
- プロジェクトの対応 OS
- IOCRActivity インターフェイス
- OCRAsyncCodeActivity クラス
- OCRCodeActivity クラス
- OCRNativeActivity クラス
- Character クラス
- OCRResult クラス
- Word クラス
- FontStyles 列挙型
- OCRRotation 列挙型
- OCRCapabilities クラス
- OCRScrapeBase クラス
- OCRScrapeFactory クラス
- ScrapeControlBase クラス
- ScrapeEngineUsages 列挙型
- ScrapeEngineBase
- ScrapeEngineFactory クラス
- ScrapeEngineProvider クラス
- OmniPage
- PDF
- リリース ノート
- PDF アクティビティ パッケージについて
- プロジェクトの対応 OS
- PDF のコード化されたオートメーション API (プレビュー)
- [リストから削除済] ABBYY
- [リストから削除済] ABBYY Embedded
PDF アクティビティ パッケージのコード化されたオートメーション API です。PDF および XPS ファイルの読み取り、結合、変換、操作に対応しています。
UiPath.PDF.Activities
PDF および XPS ファイルの読み取り、結合、変換、操作を行うためのコード化されたワークフロー API です。これらの API は、コード化されたオートメーションを設計する際に使用できます。コード化されたオートメーションの概要と、API を使用してそれらのオートメーションを設計する方法については、「 コード化されたオートメーション」をご覧ください。
- サービス アクセサ:
pdf(IPdfService型) - 必要なパッケージ: 依存関係
project.json"UiPath.PDF.Activities": "*"。
自動インポートされる名前空間
PDF パッケージをインストールすると、コード化されたワークフローで以下の名前空間が自動的に利用可能になります。
UiPath.PDF.Activities.ApiUiPath.PDFUiPath.PDF.Activities.PDF.EnumsUiPath.Platform.ResourceHandlingSystem.Net.Mail
サービスの概要
pdf サービスでは、各操作が直接メソッド呼び出しとして公開されます。開くコネクションまたはスコープがありません。ほとんどの操作には 2 つの形式があります。1 つはファイル パスを stringとして受け取る形式と、 IResource 形式 (たとえば、[ パスの存在を確認] アクティビティや [ ローカル ファイル/フォルダーを取得 ] アクティビティの出力など) です。サービス アクセサーのメソッドを直接呼び出します。
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
ファイルを作成または変換するメソッドは、出力を指す ILocalResource を返します。outputFileName が nullの場合、サービスによって名前が自動的に生成されます。パスワードで保護されたファイルは、 password パラメーターを渡すことで開かれます。range パラメーターには、"All"範囲または明示的な範囲 (例: "1-3,5") を指定できます。
テキストの読み上げ
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)
PDF ファイルからテキストを読み取ります。テキスト レイアウトを維持するには、 preserveFormatting を [ true ] に設定します。IResourceを受け入れるオーバーロードも使用できます。
string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
付属の OCR エンジンを使用して、スキャンした PDF からテキストを読み取ります。OCR サービスから OCR エンジンを取得します (例: IOcrService.GetUiPathDocumentOcr)。degreeOfParallelism 一度に処理するページ数を設定します。Windows 専用です。
string ReadXpsText(string fileName, string range = "All")
XPS ファイルからテキストを読み取ります。IResourceを受け入れるオーバーロードも使用できます。Windows 専用です。
string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)
提供された OCR エンジンを使用して XPS ファイルからテキストを読み取ります。Windows 専用です。
ページ操作とドキュメント操作
int GetPdfPageCount(string fileName, string password = null)
PDF ファイルのページ数を返します。IResourceを受け入れるオーバーロードも使用できます。
ILocalResource JoinPdf(string[] fileList, string outputFileName = null)
複数の PDF ファイルを指定の順序で 1 つの PDF に結合します。IResource[]を受け入れるオーバーロードも使用できます。
ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)
PDF から PDF から新しい PDF にページの範囲 (例: "1-3,5"など) を抽出します。IResourceを受け入れるオーバーロードも使用できます。
ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
1 つのページ (1 から開始 pageNumber) を画像としてエクスポートします。画像形式は、 outputFileName 拡張子から推論されます。IResourceを受け入れるオーバーロードも使用できます。
作成と変換
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)
画像ファイルのリストから PDF を作成します (1 ページに 1 つの画像)。サポートされている入力形式は、PNG、JPG、JPEG、TIF、TIFF、BMP です。IResource[]を受け入れるオーバーロードも使用できます。
ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
HTML コンテンツを PDF に変換します。背景色とグラフィックを保持するには、[ keepBackground ] を [ true ] に設定します。
ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
メール メッセージを PDF に変換します。
ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)
プレーン テキストを PDF に変換します。フォント サイズとテキストの配置を設定できます。
画像と添付ファイルを抽出する
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)
PDF からすべての画像を抽出し、選択した形式で保存します。IResourceを受け入れるオーバーロードも使用できます。
IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)
PDF に埋め込まれたファイルを抽出します。filterを特定の拡張機能 (例: new[] { ".docx", ".xlsx" }) に制限するように設定します。IResourceを受け入れるオーバーロードも使用できます。
パスワードを管理する
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)
PDF のユーザーと所有者のパスワードを設定または削除します。削除するには、新しいパスワードとして null または空の文字列を渡します。IResourceを受け入れるオーバーロードも使用できます。
オプション
ConvertToPdfOptions
変換方法の共有レイアウトとコンテンツ オプション。
HeaderHtmlString- ページ ヘッダーとして使用される HTML スニペットです。既定値はnullです (ヘッダーなし)。FooterHtmlString- ページ フッターとして使用される HTML スニペットです。既定値はnullです (フッターなし)。PaperSizePaperSize- 出力 PDF の用紙サイズです。既定値はA4です。MarginInt- ページ余白 (ポイント)。既定値は0です。ScaleDouble- ページのレンダリング規模です (0.1から2まで)。既定値は1.0です。
列挙型のリファレンス
ImageExtension
ExtractImagesFromPdfが使用 します。
PNG- PNG 画像形式です。JPEG- JPEG 画像形式です。TIFF- TIFF 画像形式です。BMP- BMP 画像形式です。
ImageDpi
ReadPdfWithOcr と ExportPdfPageAsImageで使用します。
Low- 96 DPIMedium- 150 DPI既定。High- 270 DPI。
TextAlignment
ConvertTextToPdfが使用 します。
Justify- 両端揃え。既定。Left- 左揃え。Right- 右揃え。Center- 中央揃え。
PaperSize
ConvertToPdfOptionsが使用 します。
A4(既定)、A3、A5、A2、A6- ISO A シリーズのサイズです。Letter、Legal、Tabloid、Ledger、Executive、Statement- 北米サイズ。B4、B5- ISO Bシリーズサイズ。Number10Envelope、DLEnvelope、C5Envelope、C4Envelope- エンベロープのサイズです。
アクティビティとの関係
コード化された API アクティビティと PDF XAML アクティビティは同じランタイムで実行されるため、コード化されたワークフローは、 PDF アクティビティ パッケージのアクティビティと同じ読み取り、変換、操作の動作になります。違いは呼び出しの形状です。サービスはプレーンなファイル パスまたはIResource入力を受け取り、string、int、またはILocalResource結果を直接返します。また、アクティビティの [エラー発生時に実行を継続] オプションの代わりに、通常の try/catch を使用します。
一般的なパターン
PDF のテキストを読み込む
このパターンは、PDF ファイルからページ範囲のテキストを読み取ります。ReadPdfText は、抽出されたテキストを stringとして返します。
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
PDF を結合してページ範囲を抽出する
このパターンは、2 つの PDF ファイルを 1 つに結合し、結合結果からページ範囲を抽出します。各ステップは、出力ファイルを指す ILocalResource を返します。
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
HTML を余白のあるレターサイズの PDF に変換する
このパターンでは、 ConvertToPdfOptions を使用して Letter 用紙サイズとページ マージンを設定し、HTML 文字列を PDF に変換します。ConvertHtmlToPdf は、作成した PDF を指す ILocalResource を返します。
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
- 自動インポートされる名前空間
- サービスの概要
- テキストの読み上げ
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)string ReadXpsText(string fileName, string range = "All")string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)- ページ操作とドキュメント操作
int GetPdfPageCount(string fileName, string password = null)ILocalResource JoinPdf(string[] fileList, string outputFileName = null)ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)- 作成と変換
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)- 画像と添付ファイルを抽出する
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)- パスワードを管理する
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)- オプション
ConvertToPdfOptions- 列挙型のリファレンス
ImageExtensionImageDpiTextAlignmentPaperSize- アクティビティとの関係
- 一般的なパターン
- PDF のテキストを読み込む
- PDF を結合してページ範囲を抽出する
- HTML を余白のあるレターサイズの PDF に変換する