About the Role
You'll own the reliability and quality of a web extraction infrastructure that processes millions of pages daily through hundreds of scripts, embedded within a fast-moving AI startup team. Reporting to the Forward Deployed Engineer, this role is critical to keeping a high-throughput data pipeline running smoothly and accurately.
What You'll Do
-
Monitor and triage a queue of broken scrapers and data quality alerts to maintain smooth operations.
-
Build and deploy new web scraping scripts for sites that require data extraction.
-
Validate scraped data for accuracy and investigate discrepancies as they arise.
-
Create dashboards to visualize and monitor scraped data in real time.
-
Collaborate with AI agents to fill capability gaps, fix issues, and improve existing scripts.
What We're Looking For
-
1+ years of hands-on professional experience building or maintaining web scraping solutions.
-
Proficiency with TypeScript and Node.js for building and debugging scraping scripts.
-
Proficiency with SQL for data querying and validation.
-
Experience with Puppeteer or similar browser automation libraries.
-
Experience debugging and fixing broken scrapers in production environments.
-
Experience with message queues such as RabbitMQ or similar asynchronous job processing systems.
-
Experience with Redis or similar in-memory caching systems.
-
Familiarity with Google Cloud Platform or equivalent cloud infrastructure (BigQuery experience is a plus).
-
Experience with bash scripting, git, and gRPC.
-
Experience building dashboards or data visualization tools is a plus.
-
Strong async communicator who proactively flags blockers and sends updates without prompting.
-
Comfortable with ambiguity and energized by a fast-paced, early-stage environment.
Location
Fully remote, open to any time zone. Candidates must have 8+ hours of daily availability with overlap during US business hours. Visa sponsorship is not available.