# About ML Packages

> Using a Document Understanding ML Package involves these steps:

Using a Document Understanding ML Package involves these steps:

* Collect document samples and the requirements of the data points that need to be extracted.
* Label documents using [Data Manager](data-manager.md).

  Data Manager itself connects to an [OCR Service](ocr-services1.md).
* Download or export labeled documents as a Training dataset and upload that exported folder to AI Center Storage.
* Download or export labeled documents as an Evaluation dataset and upload that exported folder to AI Center Storage.
* Run a Training Pipeline on AI Center.
* Evaluate the model performance with an Evaluation Pipeline on AI Center.
* Deploy the trained model as an ML Skill in AI Center.
* Query the ML Skill from an RPA Workflow using the [UiPath.DocumentUnderstanding.ML](https://docs.uipath.com/activities/docs/about-the-documentunderstanding-ml-activities-pack) activity package.

  :::note
  Remember that using Document Understanding ML Packages requires that the machine on which AI Center is installed can access `https://du-metering.uipath.com`.
  :::

  :::important
  When creating a UiPath.DocumentUnderstanding.ML.Activities Package in AI Center, the package name should not be any python reserved keyword, such as `class` , `break`, `from`, `finally`, `global`, `None`, etc. Note that this list is not exhaustive since the package name is used for `class <pkg-name>` and `import <pkg-name>`.
  :::

These are out-of-the-box Machine Learning Models to classify and extract any commonly occurring data points from semi-structured or unstructured documents, including regular fields, table columns, and classification fields, in a template-less approach.

![docs image](https://dev-assets.cms.uipath.com/assets/images/document-understanding/document-understanding-116175-5a7c46d2.webp)

:::note
Out-of-the-box Machine Learning Packages that are delivered by UiPath have version 0 and are already available on your tenant, meaning that there is no need to download them.

Download is available only for versions 1 or higher, that were already trained by you.
:::

Document Understanding contains multiple ML Packages split into six main categories:

* [UiPathDocumentOCR](about-ml-packages.md#uipathdocumentocr)
* [UiPathDocumentOCR_CPU Preview](about-ml-packages.md)
* [DocumentUnderstanding](about-ml-packages.md#documentunderstanding)
* [DocumentClassifier](about-ml-packages.md#documentclassifier)
* [Out-of-the-box Pre-trained ML Packages](about-ml-packages.md#out-of-the-box-pre-trained-ml-packages)
* [Other Out-of-the-box ML Packages](about-ml-packages.md#other-out-of-the-box-ml-packages)

## UiPathDocumentOCR

This is a non-retrainable model which can be used with the [UiPath Document OCR](https://docs.uipath.com/activities/docs/ui-path-document-ocr) engine activity as part of the [Digitize Document](https://docs.uipath.com/activities/docs/digitize-document) activity. To be used, the ML Skill must first be made public so that a URL can be copy-pasted into the UiPath Document OCR engine activity.

UiPathDocumentOCR requires access to the Document Understanding metering server at https://du.uipath.com/metering if the ML skill is running on an AI Center on-premises regular deployment. No internet access is needed on AI Center on-premises air-gapped deployments.

The UiPathDocumentOCR ML Package in AI Center is optimized for running on GPU, so we strongly recommend using it on GPU. If no GPU is available, we recommend using the standalone docker container for versions before 2021.10. Starting with 2021.10, the ML package can also be run in AI Center on-premises, but we advise having at least a 4-core CPU or ideally an 8-core CPU.

## UiPathDocumentOCR_CPU Preview

This ML Package can be deployed exactly the same way as the **UiPathDocumentOCR** ML Package, with the following differences:

* it is optimized to run on CPU, so you should see a 3-4x speedup when running in workflow, and 5-10x speedup when using it to import documents into **Document Manager**
* accuracy is slightly lower than the **UiPathDocumentOCR** ML Package, and it is similar to the [UiPath.DocumentUnderstanding.OCR.LocalServer](https://activities.uipath.com/docs/about-the-documentunderstanding-ocr-local-server-pack) Studio package
* due to being faster, the CPU is also recommended when documents are large (over 20 pages per doc) in the absence of a GPU, which is ideal.

## DocumentUnderstanding

This is a generic, retrainable model for extracting any commonly occurring data points from any type of structured or semi-structured documents, building a model from scratch. This ML Package must be trained. If deployed without training first, deployment fails with an error stating that the model is not trained.

## DocumentClassifier

This is a generic, retrainable model for classifying any type of structured or semi-structured documents, building a model from scratch. This ML Package must be trained. If deployed without training first, deployment fails with an error stating that the model is not trained.

## Out-of-the-box Pre-trained ML Packages

These are retrainable ML Packages that hold the knowledge of different Machine Learning Models.

They can be customized to extract additional fields or support additional languages using Pipeline runs. Using state-of-the-art transfer learning capabilities, this model can be retrained on additional labeled documents and tailored to specific usecases or expanded for additional Latin, Cyrillic or Greek language support.

The dataset used may have the same fields, a subset of the fields, or have additional fields. To benefit from the intelligence already contained in the pre-trained model, you need to use fields with the same names as in the out-of-the-box model itself.

These ML Packages are:

* **Invoices**: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/invoices/info/model).
* **InvoicesAustralia**: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/invoices_au/info/model).
* **InvoicesIndia**: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/invoices_india/info/model).
* **InvoicesJapan** `Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/invoices_japan/info/model).

  Retraining using data from **Validation Station** is currently not supported.
* **InvoicesChina** `Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/invoices_china/info/model).

  Retraining using data from **Validation Station** is currently not supported.
* **Receipts**: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/receipts/info/model).
* **Purchase Orders**: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/purchase_orders/info/model).
* **Utility Bills**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/utility_bills/info/model).
* **ID Cards**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/id_cards/info/model).
* **Passports**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/passports/info/model).
* **RemittanceAdvices**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/remittance_advices/info/model).
* **DeliveryNotes**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/delivery_notes/info/model).
* **W2**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/w2/info/model).
* **W9**`Preview`: The fields extracted out-of-the-box can be found [here](https://du.uipath.com/ie/w9/info/model).

These models are deep learning architectures built by UiPath. A GPU can be used both at serving time and training time but is not mandatory. A GPU delivers&gt;10x improvement in speed for Training in particular.

## Other Out-of-the-box ML Packages

These are non-retrainable Packages that are required for non-ML components of the Document Understanding suite.

These ML Packages are:

* **FormExtractor**: Deploy as Public Skill and paste the URL into the [Form Extractor](data-extraction-form-extractor.md) activity.
* **IntelligentFormExtractor**: Deploy as Public Skill and paste the URL into the [Intelligent Form Extractor](data-extraction-intelligent-form-extractor.md) activity. Make sure to first deploy the **HandwritingRecognition** ML Skill and configure that as OCR for the this package.
* **IntelligentKeywordClassifier**: Deploy as Public Skill and paste the URL into the [Intelligent Keyword Classifier](document-classification-intelligent-keyword-classifier.md) activity.
* **HandwritingRecognition**: Deploy as Public Skill and use as OCR when creating the **IntelligentFormExtractor** package.
