- Vue d'ensemble (Overview)
- Document Processing Contracts
- Notes de publication
- À propos des contrats de traitement de documents
- Classe Zone
- Interface ActivitéIPersisted
- Classe PrettyBoxConverter
- Interface ActivitéIClassifier
- Interface FournisseurIClassifieurCapacités
- Classe TypeDocumentClassifieur
- Classe RésultatClassifieur
- Classe ActivitéCodeClassifieur
- Classe ActivitéClassifieurNatif
- Classe ActivitéClassifieurCodeAsync
- Classe CapacitéClassifieurTypeDocument
- ContentValidationData Class
- EvaluatedBusinessRulesForFieldValue Class
- EvaluatedBusinessRuleDetails Class
- Classe ActivitéExtracteurCodeAsync
- Classe ActivitéExtracteurCode
- Classe ExtracteurTypeDocument
- Classe ExtracteurDocumentTypeCapacités
- Classe ExtracteurChampCapacités
- Classe ActivitéExtracteurNatif
- Classe ExtracteurRésultat
- FieldValue Class
- FieldValueResult Class
- Interface FournisseurICapabilities
- Interface ActivitéIExtractor
- Classe ChargeUtileExtracteur
- Énumération PrioritéActionDocument
- Classe DocumentActionData
- Énumération StatutActionDocument
- Énumération TypeActionDocument
- Classe DocumentClassificationActionData
- Classe DocumentValidationActionData
- Classe DonnéesUtilisateur
- Classe Documents
- Classe RésultatDivisionDocument
- Classe ExtensionDom
- Classe Page
- Classe SectionPage
- Classe Polygone
- Classe ConvertisseurPolygones
- Classe de métadonnées
- Classe GroupeMot
- Classe Mot
- Énumération SourceTraitement
- Classe CelluleRésultatsTable
- Classe ValeurTableRésultats
- Classe InformationsColonnesTableRésultats
- Classe TableRésultats
- Énumération Rotation
- Rule Class
- RuleResult Class
- RuleSet Class
- RuleSetResult Class
- Énumération TypeSection
- Énumération TypeGroupeMot
- ProjectionTexteIDocument Interface
- Classe RésultatClassification
- Classe RésultatExtraction
- Classe ResultatsDocument
- Classe ResultatsLimitesDocument
- Classe ResultatsDonnéesPoint
- Classe RésultatsValeur
- Classe ResultatsContenuRéference
- Classe ResultatsValeurJetons
- Classe ResultatsChampDérivé
- Énumération ResultatsSourceDonnées
- Classe ResultatsConstantes
- Classe ChampValeurSimple
- Classe ValeurChampTable
- Classe GroupeDocument
- Classe DocumentTaxonomie
- Classe TypeDocument
- Classe Champ
- Énumération TypeChamp
- FieldValueDetails Class
- Classe InfoLangage
- Classe SaisieMétadonnées
- Énumération TypeTexte
- Classe TypeFieldTypeField Class
- Interface ActivitéISuivi
- ITrainableActivity Interface
- Interface ActivitéClassifieurITrainable
- Interface ActivitéExtracteurITrainable
- Classe ActivitéFormationClassifieurCodeAsync
- Classe ActivitéFormationClassifieurCode
- Classe ActivitéFormationClassifieurNatif
- Classe ActivitéFormationExtracteurCodeAsync
- Classe ActivitéFormationExtracteurCode
- Classe ActivitéFormationExtracteurNative
- BasicDataPoint Class
- BasicValue Class
- ComponentCollectionFacade Class
- DataPointFacadeBase Class
- ExtractionResultHandler Class
- FieldGroupDataPoint Class
- FieldGroupValue Class
- FieldLookupBase Class
- FieldRedactionSettings Class
- RedactionOptions Class (Preview)
- RedactionType Enum
- ResultsValueFacadeBase Class
- TableDataPoint Class
- TableRow Class
- TableValue Class
- WildcardDataPoint Class
- WildcardDataPointCollection Class
- Document Understanding ML
- Serveur local OCR Document Understanding
- Document Understanding
- Notes de publication
- À propos du package d’activités Document Understanding
- Compatibilité du projet
- Configuration de la connexion externe
- Document Understanding coded automation APIs (preview)
- Définir le mot de passe du PDF
- Merge PDFs
- Get PDF Page Count
- Extraire le texte PDF (Extract PDF Text)
- Extract PDF Images
- Extract PDF Page Range
- Extraire les données du document
- Create Validation Task and Wait
- Attendre la tâche de validation et reprendre
- Create Validation Task
- Créer une action de validation de document (Create Document Validation Action)
- Retrieve Document Validation Artifacts
- Classer un document (Classify Document)
- Créer une tâche de validation de classification (Create Classification Validation Task)
- Créer une tâche de validation de classification et attendre (Create Classification Validation Task and Wait)
- Attendre la tâche de validation de la classification et reprendre
- IntelligentOCR
- Notes de publication
- À propos du package d'activités IntelligentOCR
- Compatibilité du projet
- Load Taxonomy
- Digitize Document
- Classify Document Scope
- Keyword Based Classifier
- Classifieur de projet Document Understanding (Document Understanding Project Classifier)
- Intelligent Keyword Classifier
- Create Document Classification Action
- Créer une action de validation de document (Create Document Validation Action)
- Retrieve Document Validation Artifacts
- Attendre l'action de classification du document et reprendre
- Tester l'étendue des classifieurs
- Outil d'entraînement de classifieur basé sur des mots-clés
- Intelligent Keyword Classifier Trainer
- Data Extraction Scope
- Extracteur de projet Document Understanding (Document Understanding Project Extractor)
- Entraîneur d’extracteur de projet Document Understanding
- Regex Based Extractor
- Form Extractor
- Extracteur de formulaires intelligents
- Caviarder le document
- Create Document Validation Action
- Wait For Document Validation Action And Resume
- Tester l'étendue des extracteurs
- Export Extraction Results
- Extracteur d'apprentissage automatique
- Machine Learning Extractor Trainer
- Machine Learning Classifier
- Machine Learning Classifier Trainer
- Classifieur génératif
- Extracteur génératif
- Configuration de l'authentification
- Valider des documents avec des actions App
- Valider manuellement des documents numérisés
- Extraction de données basée sur des ancres à l'aide de l'Extracteur de formulaires intelligent
- Station de validation
- Activités génératives - Bonnes pratiques
- Extracteur génératif - Bonnes pratiques
- Classifieur génératif - Bonnes pratiques
- Services ML
- OCR
- Contrats OCR
- Notes de publication
- À propos des contrats OCR
- Compatibilité du projet
- Interface ActivitéIOCR
- Classe OCRCodeAsync
- Classe ActivitéCodeOCR
- Classe ActivitéOCRNatif
- Classe Caractère
- Classe RésultatOCR
- Classe Mot
- Énumération StylesPolice
- Énumération RotationOCR
- Classe OCRCapabilities
- Classe BaseCaptureOCR
- Classe UsineCaptureOCR
- Classe BaseContrôleCapture
- Énumération UtilisationCaptureMoteur
- ScrapeEngineBase
- Classe ScrapeEngineFactory
- Classe ScrapeEngineProvider
- OmniPage
- PDF
- Notes de publication
- À propos du package d'activités PDF
- Compatibilité du projet
- PDF coded automation APIs (preview)
- Convert Email to PDF (preview)
- Convert HTML to PDF (preview)
- Convert Text to PDF (preview)
- Export PDF Page As Image
- Extraire les pièces jointes d’un fichier PDF
- Extract Images From PDF
- Extract PDF Page Range
- Get PDF Page Count
- Join PDF Files
- Manage PDF Password
- Lire le texte PDF (Read PDF Text)
- Lire le PDF avec OCR (Read PDF With OCR)
- Lire le texte XPS (Read XPS Text)
- Lire le XPS avec OCR (Read XPS With OCR)
- [Non listé] Abbyy
- Notes de publication
- À propos du package d'activités Abbyy
- Compatibilité du projet
- Reconnaissance optique des caractères ABBYY (ABBYY OCR)
- Reconnaissance optique des caractères ABBYY Cloud (ABBYY Cloud OCR)
- FlexiCapture Classifier
- FlexiCapture Extractor
- FlexiCapture Scope
- Classer un document (Classify Document)
- Traiter le document (Process Document)
- Valider le document (Validate Document)
- Exporter le document (Export Document)
- Obtenir le champ (Get Field)
- Obtenir la table (Get Table)
- Prepare Validation Station Data
- [Non listé] Abbyy intégré
Coded automation APIs for the PDF activities package, covering reading, merging, converting, and manipulating PDF and XPS files.
UiPath.PDF.Activities
Coded workflow API for reading, merging, converting, and manipulating PDF and XPS files. These APIs are available when designing coded automations. For an introduction to coded automations and how to design them using APIs, see Coded Automations.
- Accéder au service:
pdf(typeIPdfService) - Package requis:
"UiPath.PDF.Activities": "*"dans les dépendancesproject.json.
Espaces de noms importés automatiquement
The following namespaces are automatically available in coded workflows when the PDF package is installed:
UiPath.PDF.Activities.ApiUiPath.PDFUiPath.PDF.Activities.PDF.EnumsUiPath.Platform.ResourceHandlingSystem.Net.Mail
Vue d’ensemble du service
The pdf service exposes each operation as a direct method call. There is no connection or scope to open. Most operations provide two forms: one that takes a file path as a string, and one that takes an IResource (for example, the output of the Path Exists or Get Local File or Folder activity). Call methods on the service accessor directly:
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
Methods that create or transform a file return an ILocalResource pointing to the output. When outputFileName is null, the service generates a name automatically. Password-protected files are opened by passing the password parameter. The range parameter accepts "All" or explicit ranges such as "1-3,5".
Reading text
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)
Reads the text from a PDF file. Set preserveFormatting to true to keep the text layout. An overload accepting an IResource is also available.
string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
Reads the text from a scanned PDF using the supplied OCR engine. Obtain an OCR engine from the OCR service (for example, IOcrService.GetUiPathDocumentOcr). degreeOfParallelism sets how many pages are processed at once. Windows only.
string ReadXpsText(string fileName, string range = "All")
Reads the text from an XPS file. An overload accepting an IResource is also available. Windows only.
string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)
Reads the text from an XPS file using the supplied OCR engine. Windows only.
Page and document operations
int GetPdfPageCount(string fileName, string password = null)
Returns the number of pages in a PDF file. An overload accepting an IResource is also available.
ILocalResource JoinPdf(string[] fileList, string outputFileName = null)
Merges multiple PDF files into a single PDF, in the order given. An overload accepting an IResource[] is also available.
ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)
Extracts a range of pages (for example, "1-3,5") from a PDF into a new PDF. An overload accepting an IResource is also available.
ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
Exports a single page (1-based pageNumber) as an image. The image format is inferred from the outputFileName extension. An overload accepting an IResource is also available.
Creating and converting
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)
Creates a PDF from a list of image files, one image per page. Supported input formats: PNG, JPG, JPEG, TIF, TIFF, BMP. An overload accepting an IResource[] is also available.
ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
Converts HTML content to a PDF. Set keepBackground to true to preserve background colors and graphics.
ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
Converts an email message to a PDF.
ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)
Converts plain text to a PDF, with a configurable font size and text alignment.
Extracting images and attachments
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)
Extracts all images from a PDF and saves them in the chosen format. An overload accepting an IResource is also available.
IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)
Extracts the files embedded in a PDF. Set filter to limit to specific extensions (for example, new[] { ".docx", ".xlsx" }). An overload accepting an IResource is also available.
Managing passwords
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)
Sets or removes the user and owner passwords on a PDF. Pass null or an empty string as a new password to remove it. An overload accepting an IResource is also available.
Options
ConvertToPdfOptions
Shared layout and content options for the conversion methods.
HeaderHtmlString- HTML snippet used as the page header. Defaults tonull(no header).FooterHtmlString- HTML snippet used as the page footer. Defaults tonull(no footer).PaperSizePaperSize- Paper size for the output PDF. Defaults toA4.MarginInt- Page margin in points. Defaults to0.ScaleDouble- Page rendering scale, between0.1and2inclusive. Defaults to1.0.
Enum référence
ImageExtension
Utilisé par ExtractImagesFromPdf.
PNG- PNG image format.JPEG- JPEG image format.TIFF- TIFF image format.BMP- BMP image format.
ImageDpi
Utilisé par ReadPdfWithOcr et ExportPdfPageAsImage.
Low- 96 DPI.Medium- 150 DPI. Default.High- 270 DPI.
TextAlignment
Utilisé par ConvertTextToPdf.
Justify- Justified alignment. Default.Left- Left alignment.Right- Right alignment.Center- Center alignment.
PaperSize
Utilisé par ConvertToPdfOptions.
A4(default),A3,A5,A2,A6- ISO A-series sizes.Letter,Legal,Tabloid,Ledger,Executive,Statement- North American sizes.B4,B5- ISO B-series sizes.Number10Envelope,DLEnvelope,C5Envelope,C4Envelope- Envelope sizes.
Relation avec les activités
The coded API and the PDF XAML activities run on the same runtime, so a coded workflow reaches the same reading, conversion, and manipulation behavior as the activities in the PDF activity package. The difference is the call shape: the service takes plain file paths or IResource inputs and returns string, int, or ILocalResource results directly, and you use ordinary try/catch instead of the activity Continue On Error option.
Modèles communs
Read the text of a PDF
This pattern reads the text of a page range from a PDF file. ReadPdfText returns the extracted text as a string.
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
Merge PDFs and extract a page range
This pattern merges two PDF files into one and then extracts a page range from the merged result. Each step returns an ILocalResource pointing to the output file.
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
Convert HTML to a Letter-size PDF with a margin
This pattern converts an HTML string to a PDF, using ConvertToPdfOptions to set a Letter paper size and a page margin. ConvertHtmlToPdf returns an ILocalResource pointing to the created PDF.
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
- Espaces de noms importés automatiquement
- Vue d’ensemble du service
- Reading text
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)string ReadXpsText(string fileName, string range = "All")string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)- Page and document operations
int GetPdfPageCount(string fileName, string password = null)ILocalResource JoinPdf(string[] fileList, string outputFileName = null)ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)- Creating and converting
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)- Extracting images and attachments
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)- Managing passwords
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)- Options
ConvertToPdfOptions- Enum référence
ImageExtensionImageDpiTextAlignmentPaperSize- Relation avec les activités
- Modèles communs
- Read the text of a PDF
- Merge PDFs and extract a page range
- Convert HTML to a Letter-size PDF with a margin