Document extraction
Document Extraction (/ocr) reads a document on the server. It returns either the
document’s text or the structure of the source code in it. To open it, go to
Fetch and select Document Extraction beside ACTIVE
CONNECTIONS.
### Choose a mode
Select **Text** to extract prose, or **Code** to extract source-code structure.
Changing the mode clears the previous result.
### Point to the document
Enter the **Document path**, for example `/path/to/file.pdf`. The server reads the
file at this path, so it must be a document the FACE server can open. Choose a
**Format**: AUTO, PDF, DOCX or PPTX.
### Extract
Select **Extract**. The button stays disabled until a path is entered.
Results
- Text mode shows Extracted text, with a confidence pill. The pill shows the confidence when the server measured it, and Confidence not measured otherwise.
- Code mode shows Language:, the number of functions and classes, and the code.
Before the first extraction the page says Enter a document path and extract. No text extracted. means the document was read and held no text. A failed request shows its error instead.
Behind this page: OCRService.ExtractText and ExtractCode. See the
API reference.