1

Python Web Crawler Jobs (NOW HIRING)

$250/hr

Proficiency in at least one backend language (we use Python or Rust). * Are fluent in distributed ... Have experience building a web crawler. * Have extensive experience understanding and scaling ...

... the modern web with a sophisticated crawler at its core. This job is at the heart of crawler ... coding with Python and Javascript * Ability to write quality code * Willing to firefight

... the modern web with a sophisticated crawler at its core. This job is at the heart of crawler ... coding with Python and Javascript * Ability to write quality code * Willing to firefight

... the modern web with a sophisticated crawler at its core. This job is at the heart of crawler ... coding with Python and Javascript * Ability to write quality code * Willing to firefight

... the modern web with a sophisticated crawler at its core. This job is at the heart of crawler ... coding with Python and Javascript * Ability to write quality code * Willing to firefight

Research Crawling Engineer

London, KY · On-site

$172K/yr

Build a distributed crawler for a continuously updated, high‑quality web project * Design a ... Go, Rust, Python, Java, or C++ * Experience working for reputable companies * Experience building ...

... Python-based backend services • Develop and operate resilient web crawling and automation ... • Writing Crawler experience (such as: Scrapy, Apache Nutch, Heritrix) • Familiarity with ...

Monitor crawler performance, identify failures, and resolve data extraction issues. * Collaborate ... Build intelligent crawling workflows using Python and modern web automation frameworks. * Implement ...

Building new features on top of our Tracker Radar crawler to detect emerging privacy threats ... Go, Node.js, Python, Perl. * Excellent communication skills. You can validate your decisions and ...

NY · On-site

... Web Vitals, JS rendering, hreflang, and multilingual SEO. * Experience leading pre- and post ... Experience building technical SEO automations -- Python, API integrations, or workflow tools; you ...

Python Web Crawler information

See salary details

$44

$57

$64

How much do python web crawler jobs pay per hour?

As of Sep 9, 2026, the average hourly pay for python web crawler in the United States is $57.88, according to ZipRecruiter salary data. Most workers in this role earn between $52.88 and $62.50 per hour, depending on experience, location, and employer.

What is a python web crawler?

A Python Web Crawler is a program written in Python that automatically navigates the web, retrieves information from websites, and often stores or processes this data for further use. Web crawlers, also known as spiders or bots, are commonly used for indexing website content, data mining, or collecting data from multiple web pages. Python is a popular language for creating web crawlers due to its robust libraries such as BeautifulSoup, Scrapy, and Requests, which make it easier to handle HTTP requests, parse HTML, and manage large-scale crawling. Properly designed crawlers respect website rules by following robots.txt files and rate limits to avoid overloading servers. They are widely used in search engines, price comparison sites, and research projects.

What are the key skills and qualifications needed to thrive as a python web crawler?

To thrive as a Python Web Crawler, you need strong programming skills in Python, a solid understanding of web protocols (HTTP/HTTPS), and experience with data extraction techniques, often supported by a relevant computer science degree. Familiarity with frameworks and libraries like Scrapy, BeautifulSoup, Requests, and sometimes Selenium, as well as knowledge of version control systems like Git, is typically required. Attention to detail, problem-solving abilities, and persistence are vital soft skills for overcoming challenges such as website structure changes and data inconsistency. These skills ensure efficient, reliable data collection and adaptability to evolving web environments, which are crucial for success in this role.

What are some common challenges faced by python web crawlers and how can they be addressed?

Python Web Crawlers often encounter challenges such as dealing with dynamically loaded content, handling rate limits or CAPTCHAs, and managing large volumes of data efficiently. To address these, it's important to use libraries like Selenium or Playwright for dynamic pages, implement respectful crawling practices to avoid IP bans, and utilize robust data storage solutions. Collaborating closely with data engineers and following best practices for error handling and code modularity also helps ensure smooth and scalable web crawling operations.

What is the difference between Python Web Crawler vs Data Analyst?

AspectPython Web CrawlerData Analyst
Required SkillsPython programming, web scraping, data extractionData analysis, statistics, Excel, SQL
Work EnvironmentTech companies, data-driven projects, remote or officeBusiness environments, reporting, data visualization
Industry UsageWeb data collection, SEO, market researchBusiness insights, decision-making, reporting

While both roles involve working with data, a Python Web Crawler focuses on developing scripts to extract data from websites using Python, whereas a Data Analyst interprets and visualizes data to support business decisions. The roles often overlap in data handling but differ in technical focus and end goals.

What are popular job titles related to Python Web Crawler jobs?

For Python Web Crawler jobs, the most frequently searched job titles are:

Infographic showing various Python Web Crawler job openings in the United States as of September 2026, with employment types broken down into 100% Full Time. Highlights an 67% In-person, and 33% Hybrid job distribution, with an average salary of $120,397 per year, or $57.9 per hour.

Software Engineer, Data Infra

San Francisco, CA • On-site

Mosaic.tech
Software Development • 51 - 200 employees

$350K - $475K/yr

Other

Medical, Dental, Vision, PTO

Posted 6 days ago


Job description

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We’re looking for an engineer to join us and contribute to data infrastructure. You'll join a small, high-impact team responsible for architecting and scaling the core infrastructure behind distributed training pipelines, multimodal data catalogs, and intelligent processing systems that operate over petabytes of data.

Infrastructure is critical to us: it's the bedrock that enables every breakthrough. You'll work directly with researchers to accelerate experiments, develop new datasets, improve infrastructure efficiency, and enable key insights across our data assets.

If you're excited by distributed systems, large-scale data mining, open-source tools like Spark, Kafka, Beam, Ray, and Delta Lake, and enjoy building from the ground up, we'd love to hear from you.

What You’ll Do
  • Design, build, and operate scalable, fault-tolerant infrastructure for LLM Research: distributed compute, data orchestration, and storage across modalities.

  • Develop high-throughput systems for data ingestion, processing, and transformation — including training data catalogs, deduplication, quality checks, and search.

  • Build systems for traceability, reproducibility, and robust quality control at every stage of the data lifecycle.

  • Implement and maintain monitoring and alerting to support platform reliability and performance.

  • Collaborate with research teams to unlock new features, improve data quality, and accelerate training cycles.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.

  • Proficiency in at least one backend language (we use Python or Rust).

  • Are fluent in distributed compute frameworks such as Apache Spark or Ray.

  • Are deeply familiar with cloud infrastructure, data lake architectures, and batch and streaming pipelines.

  • Comfort operating across the stack and owning projects end-to-end.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications — we encourage you to apply if you meet some but not all of these:

  • Have hands-on experience with Kafka, dbt, Terraform, and Airflow.

  • Have experience building a web crawler.

  • Have extensive experience understanding and scaling deduplication, data mining, and search.

  • Have strong knowledge of file formats and storage systems (e.g., Parquet, Delta Lake, etc.) and how they impact performance and scalability.

  • Are proactive about documentation, testing, and empowering your teammates with good tooling.

Logistics
  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can’t guarantee success for every candidate or role, if you’re the right fit, we’re committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

#J-18808-Ljbffr