Python Developer Needed: Advanced Web Scraping (Directory + External Email Discovery)

Remote Full-time
We are looking for an experienced Python developer (or Web Automation expert) to build a scraper for a public speaker directory. The Goal: We need to extract approximately 10,000 profiles into a clean CSV/Google Sheets database. The Challenge (2-Step Logic): The email addresses are NOT listed on the directory itself. The script must perform a "Deep Scrape": Step 1 (Directory): Scrape the profile on the main platform (our website) to get: Name, Topics, Location, Profile-URL, and the Link to their Personal Website. Step 2 (External Enrichment): The script must visit the personal website of each speaker. Step 3 (Email Extraction): On the personal website, the script must crawl for the email address. Note: The websites are in German. The script needs to look for keywords like "Impressum" (Legal Notice), "Kontakt" (Contact), or "Datenschutzerklärung" to find the page where the email is listed. It needs to handle simple regex extraction and common obfuscations (e.g., info [at] domain). Deliverables: The Dataset: A CSV/Google Sheet containing: - Name - Topics - City/Country - Personal Website URL - Extracted Email Address (if found) The Source Code: Well-documented Python script (e.g., Scrapy, Selenium, Playwright) so we can run it again in the future. Requirements: - Proven experience with Python (Scrapy/BeautifulSoup) or Headless Browsers (Selenium/Playwright). - Experience in scraping data from multiple different domain structures (since every personal website looks different). - Ability to handle potential anti-bot measures (IP rotation/delays) to scrape respectfully and avoid blocking. - Bonus: Experience with German websites (understanding the structure of "Impressum" pages). Apply tot his job
Apply Now
← Back to Home