Data Scraping Genius
Moovsoon · Egypt
Apply & track with Apply EdgeFull-Time | Remote | Growing StartupWe are looking for someone who is exceptionally good at large-scale web data collection.Not someone who has built a few Scrapy spiders. Not someone whose answer is “we can use Bright Data.” Not someone who plans to point an AI agent at websites and hope it works.We need an engineer who has already collected property listing data across the United States at meaningful scale and understands what happens when a scraper has to operate every day, across thousands of markets, against constantly changing websites and anti bot systems.Previous experience scraping U.S. real estate or property listing sites is required. Please do not apply without it.You Must Submit A Quick Note With Your Resume That Higlights This Previous Experience.What We Are BuildingWe need near real time coverage of active U.S. property listings.The objective is straightforward:Capture the newest property listing activity across the entire United States, every day, reliably and cost-effectively.This includes discovering new listings quickly, detecting changes to existing listings, maintaining broad geographic coverage, normalizing records across sources, and operating the collection infrastructure continuously.We already understand this problem and have existing scraping infrastructure. You will be speaking with technical people who have done this before.We are hiring someone who can take the system significantly further.What You Will OwnYou will eventually own the company's data acquisition and scraping infrastructure.That includes:Nationwide property listing collectionNew listing discoveryIncremental and high-frequency crawlingListing update detectionProperty detail extractionImage and media collection where appropriateSource prioritizationGeographic coverage monitoringData normalizationDuplicate detection and entity resolutionScraper health monitoringFailure detection and automatic recoveryProxy infrastructureBrowser automation infrastructureAnti-bot resilienceCost optimizationQueueing and distributed crawlingData freshness measurementSource-specific maintenanceInfrastructure for adding new data sources quicklyThe role will expand beyond property listings as our company adds additional proprietary datasets.You Should Already UnderstandWe expect candidates to be comfortable discussing, in detail:Proxy InfrastructureResidential vs ISP vs datacenter proxiesSticky sessionsIP rotation strategiesGeo-targeted proxiesASN diversityProxy reputationSuccess rate by provider and sourceBandwidth economicsCost per successful requestWhen a premium provider such as Bright Data, Oxylabs, Decodo, NetNut, etc. makes senseWhen building additional infrastructure around those services is necessarySimply saying “use Bright Data Web Unlocker” is not an architecture.Managed unblockers can be part of the stack, but you should understand when they become prohibitively expensive, where they fail, and how to build a more customized collection system around them.Fingerprinting and Anti-Bot SystemsYou should understand modern detection systems including concepts such as:TLS fingerprintsHTTP fingerprintsBrowser fingerprintsHeader consistencyUser agent consistencyCookies and session stateJavaScript challengesBrowser behaviorRequest velocityIP reputationASN detectionGeographic inconsistenciesHeadless browser detectionCAPTCHA systemsRate limitingBehavioral detectionSession managementYou should know why simply rotating IP addresses is usually not enough.Browser and Request ArchitectureYou should know when to use:Direct HTTP requestsScrapyPlaywrightPuppeteerChromiumBrowser poolsRemote browsersManaged unblockersSource APIsMobile or alternate endpoints where legitimately availableHybrid request/browser architecturesThe best solution is not necessarily the most complicated one. We care about reliability, freshness, and cost per usable record.Scale MattersThe biggest requirement for this position is understanding the difference between:“I scraped a real estate website.” and “I operated nationwide real estate collection infrastructure every day.”They are completely different problems.At scale you need to think about:Crawl schedulingIncremental crawlingPrioritizing high velocity marketsAvoiding unnecessary recrawlsIdentifying listings without repeatedly crawling entire inventoriesDetecting stale recordsDistributed workersQueue architectureBackpressureRetriesCircuit breakersRequest budgetsProxy budgetsSource specific success ratesSchema changesDOM changesSilent extraction failuresData validationMonitoringAlertingHistorical stateDeduplicationAddress normalizationListing identity across multiple sourcesFreshness SLAsWe care just as much about how intelligently you crawl as how successfully you get a page to load.Cost Is Part of the Engineering ProblemA solution that technically works but costs an enormous amount per million pages is not a successful solution.You should be able to reason about:Cost per request → success rate → records extracted → unique usable listings → cost per usable listing.The InterviewThis will be a technical interview.We will ask you to walk us through how you would design a system to collect the latest property listing activity across the entire United States every day.Be prepared to get specific.We are going to push beyond surface-level answers.Ideal BackgroundYou are likely a fit if you have:4+ years of backend, data acquisition, crawling, scraping, or infrastructure experienceDirect experience collecting U.S. property listing dataExperience running production crawlers continuouslyPython expertiseStrong knowledge of HTTP and browser behaviorExperience with distributed systemsExperience with proxy networksExperience with Playwright, Puppeteer, Scrapy, or comparable toolingExperience with queues such as Kafka, RabbitMQ, SQS, Redis, or similar systemsExperience with large-scale data pipelinesStrong SQL skillsExperience monitoring scraping reliability and data qualityA demonstrated ability to reduce infrastructure and proxy costsExperience owning hundreds of millions of requests, millions of records, or similarly large collection workloads is a major plus.What This Role Can Become:This is not a maintenance scraping position.You would start with one of the most important datasets in our business and have the opportunity to become the technical owner of our broader data infrastructure and acquisition platform.As we expand, the mandate expands with it.We are a growing startup with a strong engineering culture. We move quickly; we already know enough about this space to recognize hand-waving, and we are looking for someone who knows considerably more about large-scale data collection than we do.