PDF Text Extraction & Categorization

Employer not named by the sourceRemote

Full Stack

Apply on the company’s site

Frontier is not the employer and does not collect applications.

About this role

Python, Data Processing, Excel, LaTeX, Data Extraction, Data Analysis, Automation, Data Management · I have a collection of PDFs and I need every piece of text pulled out and neatly organized in Excel. The files aren’t consistent—some pages follow clear headings, others read more like free-form notes—so I’m looking for someone comfortable interpreting whatever structure shows up and still producing a clean, well-labeled spreadsheet.

Here’s what I expect: • Accurate extraction of all text only (no images or tables). • Smart categorization into separate Excel columns that reflect the key information in each document rather than simply mirroring the PDF layout. • A single .xlsx file returned for each batch, ready for filtering and analysis.

I’m fine with whichever approach you prefer—Python with libraries such as PyPDF2, Tabula, or PDFMiner, Adobe Acrobat scripting, or another proven method—as long as the final sheet is complete and typo-free. Let me know how you plan to tackle the varying formats and an estimated turnaround time per 100 pages, and I’ll share a sample file so you can confirm the process works before we move ahead with the full set.