- Información general
- Contratos de procesamiento de documentos
- Notas relacionadas
- Acerca de los contratos de procesamiento de documento
- Clase Cuadro
- Interfaz IPersistedActivity
- Clase PrettyBoxConverter
- Interfaz IClassifierActivity
- Interfaz IClasificadorProveedorDeCapacidades
- Clase ClassifierDocumentType
- Clase ClassifierResult
- ClassifierCodeActivity Class
- ClassifierNativeActivity Class
- ClassifierAsyncCodeActivity Class
- Clase ClasificadorCapacidadDeTipoDeDocumento
- ContentValidationData Class
- EvaluatedBusinessRulesForFieldValue Class
- EvaluatedBusinessRuleDetails Class
- Clase
- Clase
- Clase ExtractorDocumentType
- Clase ExtractorDocumentTypeCapabilities
- Clase ExtractorFieldCapability
- Clase
- Clase ExtractorResult
- FieldValue Class
- FieldValueResult Class
- Interfaz ICapabilitiesProvider
- Interfaz IExtractorActivity
- Clase ExtractorPayload
- Enumeración DocumentActionPriority
- Clase DocumentActionData
- Enumeración DocumentActionStatus
- DocumentActionType Enum
- Clase DocumentClassificationActionData
- Clase DocumentValidationActionData
- Clase UserData
- Clase Documento
- Clase DocumentoDividirResultado
- Clase DomExtensions
- Clase Página
- Clase SecciónDePágina
- Clase de polígono
- Clase PolygonConverter
- Clase de metadatos
- Clase GrupoDeWord
- Clase Word
- Enum FuenteDeProcesamiento
- Clase ResultadosTablaCelda
- Clase ResultadosTablaValor
- Clase ResultadosTablaColumnaInfo
- Clase TablaDeResultados
- Enum Rotación
- Rule Class
- RuleResult Class
- RuleSet Class
- RuleSetResult Class
- Enum TipoDeSección
- Enum TipoDeGrupoDeWord
- Interfaz IDocumentTextProjection
- Clase ResultadoDeClasificación
- Clase ResultadoDeExtracción
- Clase ResultadosDeDocumento
- Clase ResultadosDeLímitesDeDocumento
- Clase ResultadosDePuntoDeDatos
- Clase ResultadosDeValor
- Clase ResultadosDeContenidoDeReferencia
- Clase ResultadosDeValorDeTokens
- Clase ResultadosDeCampoDerivado
- Enum ResultadosDeFuenteDeDatos
- Clase ResultadoDeConstantes
- Clase ValorDeCampoSimple
- Clase ValorDeCampoDeTabla
- Clase GrupoDeDocumento
- Clase TaxonomíaDeDocumento
- Clase TipoDeDocumento
- Clase Campo
- Enum TipoDeCampo
- FieldValueDetails Class
- Clase InformaciónDeLenguaje
- Clase MetadataEntry
- Enumeración de tipo de texto
- Clase TipoDeCampo
- Interfaz de actividad de ITracking
- Interfaz de ITrainableActivity
- Interfaz ITrainableClassifierActivity
- Interfaz ITrainableExtractorActivity
- Clase TrainableClassifierAsyncCodeActivity
- Clase TrainableClassifierCodeActivity
- Clase TrainableClassifierNativeActivity
- Clase TrainableExtractorAsyncCodeActivity
- Clase TrainableExtractorCodeActivity
- Clase TrainableExtractorNativeActivity
- BasicDataPoint Class
- BasicValue Class
- ComponentCollectionFacade Class
- DataPointFacadeBase Class
- ExtractionResultHandler Class
- FieldGroupDataPoint Class
- FieldGroupValue Class
- FieldLookupBase Class
- FieldRedactionSettings Class
- RedactionOptions Class (Preview)
- RedactionType Enum
- ResultsValueFacadeBase Class
- TableDataPoint Class
- TableRow Class
- TableValue Class
- WildcardDataPoint Class
- WildcardDataPointCollection Class
- Document Understanding ML
- Servidor local de OCR de Document Understanding
- Document Understanding
- Notas relacionadas
- Acerca del paquete de actividades Document Understanding
- Compatibilidad de proyectos
- Configurar la conexión externa
- Document Understanding coded automation APIs (preview)
- Establecer contraseña de PDF
- Fusionar PDF
- Obtener el recuento de páginas del PDF
- Extraer texto en PDF
- Extraer imágenes en PDF
- Extraer rango de página en PDF
- Extraer datos del documento
- Cree una tarea de validación y espere
- Esperar la tarea de validación y continuar
- Crear tarea de validación
- Crear artefactos de validación de documentos
- Recuperar artefactos de validación de documentos
- Clasificar documento
- Crear tarea de validación de clasificación
- Crear tarea de validación de clasificación y esperar
- Esperar la tarea de validación de clasificación y reanudar
- OCRInteligente
- Notas relacionadas
- Acerca del paquete de actividades IntelligentOCR
- Compatibilidad de proyectos
- Cargar taxonomía
- Digitalizar documento
- Clasificar ámbito de documento
- Clasificador basado en palabras clave
- Clasificador de proyectos de Document Understanding
- Clasificador inteligente de palabra clave
- Crear acción de clasificación de documentos
- Crear artefactos de validación de documentos
- Recuperar artefactos de validación de documentos
- Esperar la acción de clasificación de documentos y reanudar
- Entrenar el alcance de los clasificadores
- Entrenador del clasificador basado en palabras clave
- Entrenador del clasificador inteligente de palabra clave
- Alcance de la extracción de información
- Extractor de proyectos de Document Understanding
- Entrenador del extractor de proyectos de Document Understanding
- Extractor basado en regex
- Extractor de forma
- Extractor inteligente de formularios
- Redactar documento
- Crear acción de validación de documentos
- Esperar la acción de validación de documentos y reanudar
- Entrenar el alcance de los Extractores
- Exportar resultados de extracción
- Extractor con aprendizaje automático
- Entrenador de extractor con aprendizaje automático
- Clasificador de aprendizaje automático
- Entrenador del clasificador de aprendizaje automático
- Clasificador generativo
- Extractor generativo
- Configurar autenticación
- Validar documentos con acciones de la aplicación
- Validación manual para digitalizar documentos
- Extracción de datos basada en anclajes utilizando el extractor inteligente de formularios
- Estación de validación
- Actividades generativas: buenas prácticas
- Extractor generativo: buenas prácticas
- Clasificador generativo: buenas prácticas
- Servicios ML
- OCR
- Contratos OCR
- Notas relacionadas
- Acerca de los contratos OCR
- Compatibilidad de proyectos
- IOCRActivity Interface
- OCRAsyncCodeActivity Class
- OCRCodeActivity Class
- OCRNativeActivity Class
- Clase Carácter
- Clase OCRResult
- Clase Word
- FontStyles Enum
- OCRRotation Enum
- Clase OCRCapabilities
- OCRScrapeBase Class
- OCRScrapeFactory Class
- ScrapeControlBase Class
- Enum ScrapeEngineUsages
- ExtraerBaseDelEctor
- Clase ScrapeEngineFactory
- Clase ExtraerEngineProvider
- OmniPage
- PDF
- Notas relacionadas
- Acerca del paquete de actividades de PDF
- Compatibilidad de proyectos
- PDF coded automation APIs (preview)
- Convert Email to PDF (preview)
- Convert HTML to PDF (preview)
- Convert Text to PDF (preview)
- Exportar página en PDF como imagen
- Extraer archivos adjuntos de PDF
- Extraer imágenes de PDF
- Extraer rango de página en PDF
- Obtener el recuento de páginas del PDF
- Unir archivos PDF
- Administrar contraseña de PDF
- Leer texto en PDF
- Leer PDF con OCR
- Leer texto en XPS
- Leer XPS con OCR
- [No en la lista] Abbyy
- [No en la lista] Abbyy incrustado
Coded automation APIs for the PDF activities package, covering reading, merging, converting, and manipulating PDF and XPS files.
UiPath.PDF.Activities
Coded workflow API for reading, merging, converting, and manipulating PDF and XPS files. These APIs are available when designing coded automations. For an introduction to coded automations and how to design them using APIs, see Coded Automations.
- Descriptor de acceso de servicio:
pdf(tipoIPdfService) - Paquete requerido:
"UiPath.PDF.Activities": "*"en las dependenciasproject.json.
Espacios de nombres importados automáticamente
The following namespaces are automatically available in coded workflows when the PDF package is installed:
UiPath.PDF.Activities.ApiUiPath.PDFUiPath.PDF.Activities.PDF.EnumsUiPath.Platform.ResourceHandlingSystem.Net.Mail
Descripción general del servicio
The pdf service exposes each operation as a direct method call. There is no connection or scope to open. Most operations provide two forms: one that takes a file path as a string, and one that takes an IResource (for example, the output of the Path Exists or Get Local File or Folder activity). Call methods on the service accessor directly:
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
string text = pdf.ReadPdfText(@"C:\docs\report.pdf");
Methods that create or transform a file return an ILocalResource pointing to the output. When outputFileName is null, the service generates a name automatically. Password-protected files are opened by passing the password parameter. The range parameter accepts "All" or explicit ranges such as "1-3,5".
Reading text
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)
Reads the text from a PDF file. Set preserveFormatting to true to keep the text layout. An overload accepting an IResource is also available.
string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
Reads the text from a scanned PDF using the supplied OCR engine. Obtain an OCR engine from the OCR service (for example, IOcrService.GetUiPathDocumentOcr). degreeOfParallelism sets how many pages are processed at once. Windows only.
string ReadXpsText(string fileName, string range = "All")
Reads the text from an XPS file. An overload accepting an IResource is also available. Windows only.
string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)
Reads the text from an XPS file using the supplied OCR engine. Windows only.
Page and document operations
int GetPdfPageCount(string fileName, string password = null)
Returns the number of pages in a PDF file. An overload accepting an IResource is also available.
ILocalResource JoinPdf(string[] fileList, string outputFileName = null)
Merges multiple PDF files into a single PDF, in the order given. An overload accepting an IResource[] is also available.
ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)
Extracts a range of pages (for example, "1-3,5") from a PDF into a new PDF. An overload accepting an IResource is also available.
ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)
Exports a single page (1-based pageNumber) as an image. The image format is inferred from the outputFileName extension. An overload accepting an IResource is also available.
Creating and converting
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)
Creates a PDF from a list of image files, one image per page. Supported input formats: PNG, JPG, JPEG, TIF, TIFF, BMP. An overload accepting an IResource[] is also available.
ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
Converts HTML content to a PDF. Set keepBackground to true to preserve background colors and graphics.
ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)
Converts an email message to a PDF.
ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)
Converts plain text to a PDF, with a configurable font size and text alignment.
Extracting images and attachments
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)
Extracts all images from a PDF and saves them in the chosen format. An overload accepting an IResource is also available.
IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)
Extracts the files embedded in a PDF. Set filter to limit to specific extensions (for example, new[] { ".docx", ".xlsx" }). An overload accepting an IResource is also available.
Managing passwords
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)
Sets or removes the user and owner passwords on a PDF. Pass null or an empty string as a new password to remove it. An overload accepting an IResource is also available.
Opciones
ConvertToPdfOptions
Shared layout and content options for the conversion methods.
HeaderHtmlString- HTML snippet used as the page header. Defaults tonull(no header).FooterHtmlString- HTML snippet used as the page footer. Defaults tonull(no footer).PaperSizePaperSize- Paper size for the output PDF. Defaults toA4.MarginInt- Page margin in points. Defaults to0.ScaleDouble- Page rendering scale, between0.1and2inclusive. Defaults to1.0.
Referencia de enumeración
ImageExtension
Utilizado por ExtractImagesFromPdf.
PNG- PNG image format.JPEG- JPEG image format.TIFF- TIFF image format.BMP- BMP image format.
ImageDpi
Utilizado por ReadPdfWithOcr y ExportPdfPageAsImage.
Low- 96 DPI.Medium- 150 DPI. Default.High- 270 DPI.
TextAlignment
Utilizado por ConvertTextToPdf.
Justify- Justified alignment. Default.Left- Left alignment.Right- Right alignment.Center- Center alignment.
PaperSize
Utilizado por ConvertToPdfOptions.
A4(default),A3,A5,A2,A6- ISO A-series sizes.Letter,Legal,Tabloid,Ledger,Executive,Statement- North American sizes.B4,B5- ISO B-series sizes.Number10Envelope,DLEnvelope,C5Envelope,C4Envelope- Envelope sizes.
Relación con las actividades
The coded API and the PDF XAML activities run on the same runtime, so a coded workflow reaches the same reading, conversion, and manipulation behavior as the activities in the PDF activity package. The difference is the call shape: the service takes plain file paths or IResource inputs and returns string, int, or ILocalResource results directly, and you use ordinary try/catch instead of the activity Continue On Error option.
Patrones comunes
Read the text of a PDF
This pattern reads the text of a page range from a PDF file. ReadPdfText returns the extracted text as a string.
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
[Workflow]
public void Execute()
{
string text = pdf.ReadPdfText(@"C:\docs\report.pdf", range: "1-3");
Log(text);
}
Merge PDFs and extract a page range
This pattern merges two PDF files into one and then extracts a page range from the merged result. Each step returns an ILocalResource pointing to the output file.
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
[Workflow]
public void Execute()
{
var merged = pdf.JoinPdf(new[] { @"C:\docs\a.pdf", @"C:\docs\b.pdf" }, @"C:\docs\merged.pdf");
var firstTwoPages = pdf.ExtractPdfPageRange(merged.LocalPath, "1-2", outputFileName: @"C:\docs\excerpt.pdf");
}
Convert HTML to a Letter-size PDF with a margin
This pattern converts an HTML string to a PDF, using ConvertToPdfOptions to set a Letter paper size and a page margin. ConvertHtmlToPdf returns an ILocalResource pointing to the created PDF.
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
[Workflow]
public void Execute()
{
var options = new ConvertToPdfOptions { PaperSize = PaperSize.Letter, Margin = 20 };
var pdfFile = pdf.ConvertHtmlToPdf("<h1>Report</h1><p>Generated by a coded workflow.</p>", @"C:\docs\report.pdf", options: options);
}
- Espacios de nombres importados automáticamente
- Descripción general del servicio
- Reading text
string ReadPdfText(string fileName, string password = null, string range = "All", bool preserveFormatting = false)string ReadPdfWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null, ImageDpi imageDpi = ImageDpi.Medium)string ReadXpsText(string fileName, string range = "All")string ReadXpsWithOcr(string fileName, IOCRActivity ocrEngine, string range = "All", int degreeOfParallelism = 1, string password = null)- Page and document operations
int GetPdfPageCount(string fileName, string password = null)ILocalResource JoinPdf(string[] fileList, string outputFileName = null)ILocalResource ExtractPdfPageRange(string fileName, string range, string password = null, string outputFileName = null)ILocalResource ExportPdfPageAsImage(string fileName, int pageNumber, string outputFileName, string password = null, ImageDpi imageDpi = ImageDpi.Medium)- Creating and converting
ILocalResource CreatePdfFromImages(string[] fileList, string outputFileName = null)ILocalResource ConvertHtmlToPdf(string html, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertEmailToPdf(MailMessage email, string outputFileName = null, bool keepBackground = true, ConvertToPdfOptions options = null)ILocalResource ConvertTextToPdf(string text, string outputFileName = null, int fontSize = 12, TextAlignment textAlignment = TextAlignment.Justify, ConvertToPdfOptions options = null)- Extracting images and attachments
IEnumerable<ILocalResource> ExtractImagesFromPdf(string fileName, string password = null, ImageExtension imageExtension = ImageExtension.PNG, string outputFolderName = null)IEnumerable<ILocalResource> ExtractAttachmentsFromPdf(string fileName, string password = null, string[] filter = null, string outputFolderName = null)- Managing passwords
ILocalResource ManagePdfPassword(string fileName, string oldUserPassword, string newUserPassword, string oldOwnerPassword, string newOwnerPassword, string outputFileName = null)- Opciones
ConvertToPdfOptions- Referencia de enumeración
ImageExtensionImageDpiTextAlignmentPaperSize- Relación con las actividades
- Patrones comunes
- Read the text of a PDF
- Merge PDFs and extract a page range
- Convert HTML to a Letter-size PDF with a margin