Home / Resources / Blogs

Automating Customer Onboarding with Lumiq’s Document AI and Amazon Textract

Technical JAN 23, 2024 LUMIQ Team, Lalit Khatter Data in FSI

The fourth industrial revolution is upon us, and the domino effect can be seen in the banking and insurance sector in India. The Indian BFSI (Banking, Financial Services, and Insurance) industry is moving towards advanced automation to enhance productivity, achieve cost optimizations on manual efforts, and deliver a seamless customer experience.

The Indian BFSI market specifically is peculiar compared to other regions. Each leg of the insurance process has a long cycle, is riddled with complexity and evolving fraud risk, and has an inordinate amount of customer back-and-forth.

Innovations in artificial intelligence (AI) technologies are yielding novel approaches to automation that fit the Indian market. Lumiq is spearheading such efforts and leveraging deep domain knowledge of the industry and machine learning (ML) to help create solutions for the future, today.

Lumiq is an AWS Advanced Consulting Partner) and a full-stack data science company that has expertise in data engineering, data science, ML Engineering, and MLOps, helps enterprises make sense of their data and monetize it in the most optimal manner.

In this post, Lumiq showcases how Drishti Document AI, an intelligent document parsing solution built on Amazon Textract helps optimize and accelerate customer onboarding journeys.

Benefits for BFSI Organizations with AI/ML

Back in 2017, in an Intelligent Document Processing Automation article, McKinsey, reported that organizations experimenting with AI were obtaining impressive results, including:

  • Automating 50–70 percent of tasks, which had translated into 20–35 percent annual run-rate cost efficiencies.
  • Reduction of straight-through process time of 50–60 percent with return on investment most often in triple-digit percentages.

In India, and more broadly in APJ, results are meeting or exceeding estimations. For example:

  • Achieved 95 percent manual effort reduction in Auto PII (Personal Identifiable Information) Redaction for one of India’s leading life insurance players by automating AADHAAR numbers mandatory masking as per the Supreme Court of India mandate. The first eight digits of all Aadhaar numbers, irrespective of location, are redacted out.
  • 67 percent reduction in data entry quality check process post introduction of AI guided data entry process. Information from KYC documents, handwritten proposal forms, and other relevant documents is extracted and auto-filled into a database, reducing manual labor.
  • 90 percent automation of information extraction including tabular data for a large private bank in India

Automating Information Extraction from KYC Documents

*Prerequisite Knowledge: Basic knowledge of AWS Technologies (Amazon Textract, Amazon S3) & conceptual knowledge of ML algorithms.

*

The BFSI sector is one of the most regulated industries in India, and given the KYC (Know Your Customer) mandate, the customer onboarding process is one of the most document-heavy steps in the industry. Automating instantaneous capture of data from unstructured and structured documents reduced turnaround times and the manual effort required which, in turn, enhanced customer experience.

KYC documents include the application forms and supporting documents such as government IDs, to validate the information provided in the application form.

Amazon Textract is pre-trained on millions of documents and dozens of document classes. It performs well on both printed and handwritten forms and other documents.

Image Pre-Processing

Pre-processing of documents can have a marked improvement on the quality of Textract output. Based on the kind of document scans and images we received, we implemented the following preprocessing steps before sending the documents to Amazon Textract. This step increases the overall accuracy of the text extraction:

  • Document boundary detection using the UNet model.
  • Orientation correction using the RotNet architecture.
  • Document quality improvements including but not limited to blur reduction, and contrast correction.

Text Post-Processing

The solution performs post-processing on the results provided by Amazon Textract, such as post-OCR correction with domain knowledge. For example, based on our expertise in the insurance domain, if Amazon Textract gives us “police,” we can be very sure that it’s in fact “policy”, and then correct for the same.

Information Extraction

The Drishti Document AI Engine uses a novel deep learning architecture that takes as input the text extracted (from Amazon Textract) and the interactions between text and the layout information, which is beneficial for a great number of real-world document image understanding tasks such as information extraction from scanned documents.

The Engine uses image features such as font size, font-weight, font-family, and underlining to incorporate as much visual context as possible. This allows information to be extracted as easy-to-use key-value pairs directly from the document.

The extracted information can be used for multiple use cases, including automated data entry, auto data quality check, and auto personal identifiable information redaction.

The solution has an overall accuracy of more than 95 percent.

  • Unlike most OCR tools which are template-based, Amazon Textract is template-agnostic and allows Drishti Document AI Engine to be truly template-agnostic, as well.
  • Amazon Textract can capture checkboxes and radio button selections.
  • The underlying Drishti document AI model keeps learning as more data is processed and progressively keeps improving the accuracy.

Key Takeaways

  • With the proprietary AI/ML latest techniques, you can achieve 95 percent accuracy and unlock a field-level confidence score divided into green (>95%), yellow (60% to 95%), and red zones (<60%). This allows any DEQC (data entry and quality check) operator to focus on just the poor quality (yellow and red) extractions for quick corrections, in case required.
  • Following the template-agnostic philosophy, once the AI engine has seen any single variant of a document class, it can extract information from any variant of that class with little to no drop in accuracy levels.
  • The Drishti Document AI Engine has a broad and varied catalog of document classes on which it’s trained on. To add new document classes to the catalog, the engine needs to be trained on as low as 50 documents, with a model training time of less than 24 hours!
  • Accuracy levels for other document classes will not decrease in the slightest, and for the new document class, you can see an initial accuracy greater than 95 percent! The accuracy will improve as the engine sees more documents of the new class.

Conclusion

Customer onboarding processes across industries involve large amounts of document processing and in most scenarios, the automation is limited to text extraction or OCR. Lumiq’s Drishti Document AI enables accurate text extraction and provides a high level of AI-driven business process automation leading to high savings in terms of human effort, reduced turn-around-time, and most importantly elevated customer experience.

Shoaib Mohammad, CEO & Founder, Lumiq.ai

Rapid, relevant, and correct customer service is the cornerstone of all successful businesses.

The Drishti Document AI Engine is modular and portable: it can be deployed in-situ, that is, in the field agent devices (generally cheap android tablets) itself.

  • If there is any discrepancy in the documents (and information) provided by the customer, the agent can request clearer documents on the spot. This reduces to-and-fro between the DEQC operator and customer. This reduces the onboarding time and prevents any embitterment on part of the customer.
  • Furthermore, for any straight-through-processing applications, which are generally small-risk tickets, all rules and validations can be processed onsite and such applications can be cleared in minutes!
  • If the device has access to historical data, deduplication of existing customers can be done at the time of onboarding itself.
  • Multiple other use cases can be addressed, limited only by the imagination of the business.

We invite you to learn more in regards to Lumiq Intelligent Document Processing capabilities and Drishti Document AI.

Ready to turn your data into decisions?

Tell us where your data is slowing you down. We will show you what production-grade looks like in your own AWS cloud.

Book a briefing