1

Data Scraping Jobs in California (NOW HIRING)

Integration Engineer

Oakland, CA · On-site +1

$119K - $160K/yr

... of data collection. You will use a custom Python-based macro system and write HTML, text, and OCR parsers to interface with it. OCR and web-scraping experience required . This role can be remote.

Integration Engineer

Oakland, CA · On-site

$119K - $160K/yr

... of data collection. You will use a custom Python-based macro system and write HTML, text, and OCR parsers to interface with it. OCR and web-scraping experience required . This role can be remote.

Staff Engineer

San Francisco, CA · On-site

$250 - $300/hr

Design, build, and operate core infrastructure powering a people-data and AI agent platform -- including APIs, LLM-powered workflows, web scraping systems, and large-scale knowledge graph ...

Founding Engineer

San Francisco, CA · On-site

$150K - $220K/yr

About the Role A seed-funded B2B SaaS startup in the people-data and AI agents space is looking for ... Design, build, and operate core infrastructure including APIs, LLM-powered workflows, web scraping ...

Founding Engineer

San Francisco, CA · On-site

$150K - $220K/yr

About the Role A seed-funded B2B SaaS startup in the people-data and AI agents space is looking for ... Design, build, and operate core infrastructure including APIs, LLM-powered workflows, web scraping ...

Data cleansing, scraping unstructured data and converting into structured data Implement data processing using map reduce, pig, hive scripts, Hadoop streaming, kafka, scala, spark and scheduling jobs

Experience in some of the following: full‑stack web application development, developing and designing large software systems, web scraping and other automation scripting, ETL and data pipelines ...

Showing results 41-60

Data Scraping information

What is a data scraping?

A Data Scraping job involves extracting structured data from websites or online sources using automated tools or scripts. This data is then used for various purposes, such as market research, competitive analysis, or business intelligence. Professionals in this field typically use programming languages like Python or tools like Scrapy and BeautifulSoup to collect and organize data efficiently. However, ethical considerations and legal guidelines must be followed to ensure compliance with website policies and data protection laws.

What are the key skills and qualifications needed to thrive in data scraping, and why are they important?

To excel in Data Scraping, proficiency in programming languages such as Python or JavaScript, familiarity with web protocols, and understanding of data extraction methodologies are essential, often supported by a degree in computer science or a related field. Experience with web scraping tools like BeautifulSoup, Scrapy, or Selenium, and knowledge of APIs are commonly required; relevant certifications in data analysis or automation can be advantageous. Strong problem-solving skills, attention to detail, and effective communication help professionals successfully tackle complex data challenges and collaborate with teams. These competencies are vital to reliably gather, process, and deliver valuable data while adhering to legal and ethical standards.

What are the typical challenges faced in data scraping, and how are they addressed?

Data Scraping professionals often encounter challenges such as dealing with anti-scraping measures, frequent changes to website structures, and managing large volumes of unstructured data. These obstacles require adaptability and creative problem-solving, as well as continuous learning to update scripts and leverage the latest tools. Team collaboration is common—data scrapers may work closely with data analysts, engineers, or business stakeholders to ensure the data collected meets the project's needs. Proactively communicating and staying current with industry best practices helps data scraping specialists efficiently navigate these challenges and deliver high-quality results.

What are the most commonly searched types of Data Scraping jobs in California?

The most popular types of Data Scraping jobs in California are:

What job categories do people searching Data Scraping jobs in California look for?

The top searched job categories for Data Scraping jobs in California are:

Infographic showing various Data Scraping job openings in California as of August 2026, with employment types broken down into 89% Full Time, and 11% Part Time. Highlights an 78% In-person, 11% Hybrid, and 11% Remote job distribution.

AI Engineer -- Research Agents (Full-Stack)

Socket.dev

San Francisco, CA • On-site

$180 - $260/hr

Other

Posted 3 days ago

New


Job description

What you’ll do
  • Design and ship agentic systems (tool calling, multi-agent workflows, structured outputs) that reliably fetch, extract, and normalize data across the web and APIs.
  • Build and operate search/indexing pipelines on OpenSearch/Elasticsearch (schema design, analyzers, reindex/data migration strategies, relevance tuning).
  • Own robust web scraping: directory crawling, CAPTCHA handling, headless browsers, rotating proxies, anti-bot evasion, and backoff/retry policies.
  • Develop backend services in Python + FastAPI with clean contracts and strong observability.
  • Scale workloads on AWS + Temporal (batch/queue workers, autoscaling, fault tolerance, cost control).
  • Parallelize external API requests safely (rate limits, idempotency, circuit breakers, retries, dedupe).
  • Integrate third‑party APIs for enrichment and search; model and cache responses; manage schema evolution.
  • Transform and analyze data using Pandas (or similar) for normalization, QA, and reporting.
  • Pitch in across the stack: billing (Stripe), and occasional front‑end changes to ship end‑to‑end features.
Minimum requirements
  • Hands-on experience with agentic architectures (tool calling, structured outputs/JSON, planning/execution loops) and prompt engineering.
  • Deep knowledge of OpenSearch/Elasticsearch: index design, analyzers, ingestion pipelines, snapshots, rolling upgrades, and zero-downtime reindexing/data migrations.
  • Proven web scraping expertise: solving CAPTCHAs, session/auth flows, proxy rotation, stealth techniques, and legal/ethical constraints.
  • AWS + Temporal in production (at least two of: ECS/EKS, Lambda, SQS/SNS, Batch, Step Functions, CloudWatch).
  • Building high-throughput data/IO pipelines with concurrency (asyncio/multiprocessing), resilient retries, and rate‑limit aware scheduling.
  • Integrating diverse external APIs (auth patterns, pagination, webhooks); designing stable interfaces and backfills.
  • Strong data wrangling with Pandas or equivalent; comfort with large CSV/Parquet workflows and memory/perf tuning.
  • Familiarity with Stripe (subscriptions, metered billing, webhooks) and basic front‑end changes (React/TypeScript or similar).
  • Excellent ownership, product sense, and pragmatic debugging.
Nice to have
  • Entity resolution/record linkage at scale (probabilistic matching, blocking, deduping).
  • Experience with Langfuse, OpenTelemetry, or similar for tracing/evals; task queues (Celery/RQ), Redis, Postgres.
  • Search relevance (BM25/vector/hybrid), embeddings, and retrieval pipelines.
  • Playwright/Selenium, stealth browsers, anti‑bot frameworks, CAPTCHA providers.
  • CI/CD, infrastructure as code (Terraform), and cost/perf observability.
  • Security & compliance basics for data handling and PII.
#J-18808-Ljbffr