Financial PDF OCR Data Extraction
Employer not named by the sourceRemote
Frontier is not the employer and does not collect applications.
About this role
Python, Excel, Software Architecture, OCR, Visual Basic for Apps, OpenCV, Data Extraction, Data Analysis · I have a backlog of invoices, receipts and bank statements, all supplied as searchable and non-searchable PDFs. From each document I only need two categories of information pulled out:
• the dates and amounts that appear on every page • the full itemised lines (description, quantity, unit price, line total)
Customer names or addresses are not required this time, so the workflow can stay tightly focused on these data points.
Ideally you will set up an OCR pipeline—Tesseract, ABBYY FlexiCapture, Amazon Textract, or a custom Python script with OpenCV—anything you are comfortable with that gets reliable accuracy. The final output should land in a neatly structured CSV or Excel workbook that I can import straight into my accounting software.
Acceptance criteria • ≥ 98 % field-level accuracy on a random 50-document sample • Consistent column order: Document ID, Date, Amount, Line Item, Qty, Unit Price, Line Total • Clear, commented code or repeatable tool configuration so I can rerun the process on new PDFs
Let me know approximately how long you’ll need for an initial batch of 200 documents and which stack you prefer so we can get started right away.