Advanced Web Scraping: Historical Hackers Directory
Employer not named by the sourceRemote
Frontier is not the employer and does not collect applications.
About this role
Python, Web Scraping, Data Mining, Data Extraction, Automation, Data Collection · As part of a university project, I require the public directory of historical notable hackers from the database at https://soldierx.com/ cybersecurity-forum, downloaded and saved. This is publicly available webdata, no login or registration is required. But the website has some security measures in place that make scraping trickier. Some high risk IP ranges are blocked, (eventually Proxy needed) and simple automation tools like HTTrack are detected, and the IP even gets banned sometimes. It requires a tool like PlayWright or similar to work. Formats are flexible as long as they are editable and not screenshots, allowing integration of text and images into the digital Project-Landcape. I need the entire information. This covers both the data of the index-pages and their specific subsections containing the detailed profile information for each individual.
All information on the following pages is required:
https://www.soldierx.com/hdb https://www.soldierx.com/hdb?page=1 https://www.soldierx.com/hdb?page=2 https://www.soldierx.com/hdb?page=3 https://www.soldierx.com/hdb?page=4 https://www.soldierx.com/hdb?page=5 https://www.soldierx.com/hdb?page=6 https://www.soldierx.com/hdb?page=7