OCR Field Extraction

Automated extraction of PAN, Folio numbers from scanned documents using DocTR and custom post-processing.

Problem

Financial services require extracting specific fields (PAN, Folio, Account numbers) from scanned documents. Traditional OCR gives raw text - identifying and extracting specific fields requires manual effort.

Solution

Built a specialized OCR pipeline using DocTR for text detection and custom regex/AI post-processing to identify and extract specific financial identifiers with high accuracy.

Impact

Key features

Tech stack

What K Laxman learned

← All 27 projects by K Laxman

Explore more

GitHub · LinkedIn · Email