openapi: 3.0.0 info: title: Amazon Textract Async Operations Text Detection API description: Amazon Textract is a machine learning service that automatically extracts text, handwriting, and data from scanned documents, going beyond simple OCR to identify and extract data from forms and tables. version: '2018-06-27' contact: name: Kin Lane email: kin@apievangelist.com url: https://aws.amazon.com/textract/ license: name: Apache 2.0 url: https://www.apache.org/licenses/LICENSE-2.0 servers: - url: https://textract.amazonaws.com description: Amazon Textract API endpoint tags: - name: Text Detection description: Operations for detecting text in documents. paths: /: post: operationId: DetectDocumentText summary: Amazon Textract Detect Document Text description: Detects text in the input document, returning lines and words of detected text along with confidence levels and bounding box information. headers: X-Amz-Target: description: The target API operation. schema: type: string enum: - Textract.DetectDocumentText requestBody: required: true content: application/x-amz-json-1.1: schema: type: object properties: Document: type: object description: The input document as base64-encoded bytes or an S3 object. properties: Bytes: type: string format: byte S3Object: type: object properties: Bucket: type: string Name: type: string Version: type: string required: - Document responses: '200': description: Successful text detection response. content: application/json: schema: $ref: '#/components/schemas/DocumentAnalysis' tags: - Text Detection x-microcks-operation: delay: 0 dispatcher: FALLBACK components: schemas: DocumentAnalysis: type: object description: Response object containing the results of document text detection or analysis, including detected blocks of text, tables, and forms. properties: DocumentMetadata: type: object description: Metadata about the document. properties: Pages: type: integer description: The number of pages detected in the document. Blocks: type: array description: The items detected in the document. items: type: object properties: BlockType: type: string description: The type of text item detected. enum: - KEY_VALUE_SET - PAGE - LINE - WORD - TABLE - CELL - SELECTION_ELEMENT - MERGED_CELL - TITLE - QUERY - QUERY_RESULT - SIGNATURE - TABLE_TITLE - TABLE_FOOTER - LAYOUT_TEXT - LAYOUT_TITLE - LAYOUT_HEADER - LAYOUT_FOOTER - LAYOUT_SECTION_HEADER - LAYOUT_PAGE_NUMBER - LAYOUT_LIST - LAYOUT_FIGURE - LAYOUT_TABLE - LAYOUT_KEY_VALUE Confidence: type: number format: float description: The confidence score for the detected block. Text: type: string description: The word or line of text recognized. Geometry: type: object properties: BoundingBox: type: object properties: Width: type: number Height: type: number Left: type: number Top: type: number Polygon: type: array items: type: object properties: X: type: number Y: type: number Id: type: string description: The identifier for the recognized text block. Relationships: type: array items: type: object properties: Type: type: string enum: - VALUE - CHILD - COMPLEX_FEATURES - MERGED_CELL - TITLE - ANSWER - TABLE - TABLE_TITLE - TABLE_FOOTER Ids: type: array items: type: string Page: type: integer description: The page on which the block was detected. AnalyzeDocumentModelVersion: type: string description: The version of the model used to analyze the document.