Amazon Web Services

Amazon Textract

Amazon Textract is a machine learning (ML) service that automatically extracts printed text, handwriting, layout elements and data from scanned documents, going beyond simple optical character recognition (OCR) to identify, understand and extract specific data from documents such as PDFs, images, tables and forms. It offers pretrained and customizable features to automate document processing for use cases such as loan and mortgage processing, invoices and receipts, and healthcare intake forms, extracting data in minutes instead of hours or days.

Official product page
5published use cases
3industries
1countries on record
Document Intelligence
top AI capability

Evidence mix: High 5 · Medium 0 · Low 0 — bands are computed from each record's evidence signals.

Industry
Country

5 use cases

Financial ServicesGenerative AIConversational AIDocument Intelligence

EXL Transforms Insurance Underwriting with Generative AI Assistant Built on Amazon Bedrock

EXL, a global data analytics and digital solutions company, built LDS Underwriting Assist, a retrieval-augmented generation chatbot integrated into its Life Digital Suite (LDS) platform to help insurers streamline the underwriting assessment and review stages that previously required underwriters to manually review hundreds of pages of documents. The service uses Anthropic Claude 3 Sonnet on Amazon Bedrock with Amazon Kendra as the chatbot interface, Amazon Textract to extract data from scanned documents, and Amazon Comprehend to analyze text and redact PII and protected health information for compliance with regulations such as India's PII guidelines. EXL tested multiple foundation models (Mistral AI, Claude 2, Sonnet 3, Amazon Titan) to minimize hallucinations before launching the service in August 2024, just 60 days after starting development. EXL reports the solution reduces underwriting processing time from several days to a few hours and has the potential to cut underwriting costs by up to 80%.

EXLAmazon Bedrock · Anthropic Claude 3 Sonnet · Amazon Kendra +2
HealthcareNatural Language ProcessingMachine LearningDocument Intelligence

CDPHP modernizes infrastructure and improves medical data extraction with AWS AI/ML

CDPHP, a not-for-profit health plan serving 400,000 members in Upstate New York, used AWS services including Amazon Comprehend Medical, Amazon Textract, and Amazon SageMaker to automate its data processing pipeline for unstructured medical records and health data. The organization processed over seven million records during initial migration and now processes 3,000 electronic health records weekly. CDPHP achieved a 60% improvement in overall efficiency and reduced HEDIS report generation from 4-5 days (three data scientists) to two reports produced daily.

CDPHP (Capital District Physicians' Health Plan Inc.)· United StatesAmazon Comprehend Medical · Amazon Textract · Amazon SageMaker
Government & Public SectorGenerative AINatural Language Processing

Contra Costa County District Attorney's Office Makes Unbiased Charging Decisions with ScaleCapacity Generative AI Solution on AWS

The Contra Costa County District Attorney's Office worked with AWS Partner ScaleCapacity to build a Race-Blind Charging solution to comply with California's AB 2778 mandate. The solution uses Amazon Bedrock and Amazon Textract to automatically redact race, ethnicity and other identifying details from police reports and case documents before charging decisions are made, with Amazon S3, DynamoDB, Cognito and SES supporting document storage, metadata and user access. The office processes around 17,000 cases annually, achieved compliance in six months, and can test and deploy new redaction rule changes in under a week.

Contra Costa County District Attorney's Office· United StatesAmazon Bedrock · Amazon Textract · Amazon S3 +3
HealthcareGenerative AILarge Language ModelsDocument Intelligence

Myriad Genetics speeds document processing with AWS GenAI Intelligent Document Processing Accelerator

Myriad Genetics partnered with the AWS Generative AI Innovation Center to replace an Amazon Textract/Comprehend pipeline with Amazon Bedrock foundation models (Nova Pro for classification, Nova Premier for extraction) using the open-source GenAI IDP Accelerator. Document classification accuracy rose from 94% to 98%, classification cost per page fell 77% (3.1 cents to 0.7 cents), and classification time fell 80% (8.5 minutes to 1.5 minutes per document). Automated key information extraction reached 90% accuracy matching the manual baseline, with a projected $132K in annual savings and 300 hours saved monthly across 9,000 prior authorizations in the Women's Health unit alone.

Myriad GeneticsAmazon Bedrock · Amazon Nova Pro · Amazon Nova Premier +2