About the Role
You'll be an embedded engineer on a fast-moving AI startup team, owning the reliability and quality of a web extraction infrastructure that processes millions of pages daily across hundreds of scripts. Reporting to a Forward Deployed Engineer, this role is critical to keeping a high-volume data pipeline running smoothly and accurately.
What You'll Do
-
Monitor and triage a queue of broken scrapers and data quality alerts to keep operations running without interruption.
-
Build and deploy new web scraping scripts for sites requiring data extraction.
-
Validate scraped data for accuracy and investigate discrepancies when they arise.
-
Create dashboards to visualize and monitor scraped data in real time.
-
Collaborate with AI agents to fill capability gaps, resolve issues, and improve existing scripts.
What We're Looking For
-
1+ years of professional experience building or maintaining web scraping solutions.
-
1+ years working with TypeScript, Node.js, and SQL.
-
Hands-on experience with Puppeteer or similar browser automation libraries.
-
Experience debugging and fixing broken scraping scripts in production environments.
-
Familiarity with message queues (e.g. RabbitMQ) and in-memory caching systems (e.g. Redis).
-
Experience validating data quality and accuracy in extracted datasets.
-
Comfort with GCP or equivalent cloud infrastructure, bash scripting, and git.
-
Exposure to data warehousing tools (e.g. BigQuery) and data visualization or dashboard tooling is a plus.
-
Strong async communicator who proactively flags blockers and sends updates without prompting.
-
Thrives in ambiguous, fast-paced environments with a bias toward action and efficiency.
Location
Fully remote, open to any time zone. Candidates must be available 8+ hours per day with meaningful overlap during US business hours. Visa sponsorship is not available.