Text Topic Modeling Project
Employer not named by the sourceRemote
Frontier is not the employer and does not collect applications.
About this role
Python, Statistics, Statistical Analysis, SPSS Statistics, Data Science, Data Visualization, Data Analysis, Natural Language Processing · I have a corpus of plain text that needs to be explored through topic modeling. The job is strictly data analysis: no data entry or cleaning tasks are required beyond the usual NLP-oriented preprocessing steps.
What I expect you to do • Pre-process the text (tokenisation, stop-word removal, lemmatisation, n-grams as needed). • Build and tune one or more unsupervised models—LDA, NMF or any well-justified alternative—for coherent topic extraction. • Interpret each topic with clear keyword lists and, where helpful, short example snippets. • Provide reproducible Python code (Jupyter notebook or .py script) plus a brief write-up of methodology, parameter choices and results.
Acceptance criteria • Code runs end-to-end on my machine without missing dependencies. • Topics reach an acceptable coherence score (please propose your preferred metric). • Deliverables are received within the agreed timeline and can be iterated once if needed.
Python (gensim, scikit-learn, spaCy), R, or another modern NLP stack is fine as long as the environment is documented. If you have previous work demonstrating strong results in topic modeling on text data, feel free to mention it when you