- Información general
- Crear modelos
- Consumir modelos
- Paquetes ML
- 1040: paquete ML
- 1040 Anexo C - Paquete ML
- 1040 Anexo D - Paquete ML
- 1040 Anexo E - Paquete ML
- 1040x: paquete ML
- 3949a: paquete ML
- 4506T: paquete ML
- 709: paquete ML
- 9465: paquete ML
- ACORD125: paquete ML
- ACORD126 - Paquete ML
- ACORD131 - Paquete ML
- ACORD140 - Paquete ML
- ACORD25 - Paquete ML
- Extractos bancarios: paquete ML
- Conocimientos de embarque: paquete ML
- Certificado de incorporación: paquete ML
- Certificado de origen: paquete ML
- Cheques: paquete ML
- Certificado de producto secundario: paquete ML
- CMS1500 - Paquete ML
- Declaración de conformidad de la UE: Paquete ML
- Estados financieros: paquete ML
- FM1003: paquete ML
- I9 - Paquete ML
- Documentos de identidad: paquete ML
- Facturas: paquete ML
- FacturasAustralia: paquete ML
- FacturasChina - Paquete ML
- Facturas en hebreo: paquete ML
- FacturasIndia - Paquete ML
- FacturasJapón - Paquete ML
- Envío de facturas: paquete ML
- Listas de embalaje: paquete ML
- Nóminas - - Paquete ML
- Pasaportes: paquete ML
- Órdenes de compra: paquete ML
- Recibos: paquete ML
- ConsejosDeRemesas: paquete ML
- UB04 - Paquete ML
- Facturas de servicios públicos: paquete ML
- Títulos de vehículos: paquete ML
- W2 - Paquete ML
- W9 - Paquete ML
- Puntos finales públicos
- Idiomas admitidos
- Datos y seguridad
- Lógica de licencias y tarificación
- Tutorial
Guía del usuario de Document Understanding
UiPath® DocPath
The DocPath large language model (LLM) is our latest data extraction model technology, designed to replace current generation models used within UiPath® Document UnderstandingTM. While DocPath operates similarly to previous models, it was trained using a wide variety of documents. This enables it to process common document types with little to no training needed. What sets DocPath LLM apart is its generative architecture, which significantly improves accuracy and simplifies extraction. Additionally, you can also fine-tune the model with your unique datasets.
To gain further insights into the DocPath architecture and the techniques used for training, check the DocPath page from our AI blog.
Currently, UiPath DocPath is only available for US-based tenants. Support for other regions is planned to roll out in early 2025.
DocPath LLM offers numerous enhancements over previous models. It improves accuracy, especially with tables, adapts to various document layouts to reduce annotation efforts, and boosts automation rates.
- Improved accuracy: DocPath LLM delivers a higher accuracy rate and superior F1 score for semi-structured documents such as invoices, receipts, and purchase orders. This ensures precise and consistent data extraction.
- Effortless annotation: The model reduces manual work by only requiring one annotation per document, eliminating the need to annotate each field instance on every page.
- Enhanced automation: With a greater correlation between confidence level and accuracy, DocPath LLM enhances automation rates while reducing the number of documents sent to Action Center for the same accuracy level.
From our internal tests, DocPath outperformed its predecessor in performance. It reduced the false positive rate by around 15%, and the false negative rate dropped by nearly 17%.
The DocPath LLM is available exclusively for Document Understanding modern projects. Despite the introduction of DocPath, all existing project versions will still use current model versions. This ensures a seamless transition without any disruption to ongoing production workflows.
To start training an exisiting document type on DocPath, unconfirm and confirm all fields in a few documents.
The field names you choose can greatly impact the performance of the model. To ensure optimal results, use natural language and proper grammar for field names. You should only use widely recognized acronyms such as Number (No), Account (Acct), Address (Addr), and Apartment (Apt). Currently, only West European languages are supported, so make sure that the chosen field names align with these languages. Refrain from using non-descriptive names, such as "Column 3", unless the document specifically uses that terminology.
- The extracted fields must match exactly with the text in the documents. This process does not include summarization or other types of text analysis.
- Custom training is not applicable for the following document types. If you attempt to use DocPath for these, it will result in an error:
- Facturas China
- Facturas en hebreo
- Facturas Japón