1

Freelance Web Data Scraping Jobs in California (NOW HIRING)

Data Engineer

San Francisco, CA · On-site +1

$145K/yr

Experience with web scraping and cleaning unstructured data * Knowledge of data science and machine learning concepts * Professional working experience with MLB or NBA data Salary: Starting at $145 ...

Experience with web scraping and cleaning unstructured data * Knowledge of data science and machine learning concepts * Professional working experience with MLB or NBA data Salary: Starting at $145 ...

Data Engineer

San Francisco, CA · On-site +1

$160K/yr

Experience with web scraping and cleaning unstructured data * Knowledge of data science and machine learning concepts * A strong interest in sports and sports betting, with an emphasis on Tennis. An ...

Experience with web scraping and cleaning unstructured data * Knowledge of data science and machine learning concepts * A strong interest in sports and sports betting, with an emphasis on Tennis. An ...

Data Engineer with Java & Scala

San Jose, CA · On-site

$134K - $161K/yr

... web scraping, calling APIs, write SQL queries, etc.). Work closely with our engineering team to integrate and build algorithms Process unstructured data into a form suitable for analysis - and then ...

Staff Engineer

San Francisco, CA · On-site

$250K - $300K/yr

Experience with agent orchestration, LLMs, or web scraping. * Familiarity with distributed systems, cloud infrastructure (AWS or GCP), microservices, and data pipelines. * Background at a high-growth ...

Staff Engineer

San Francisco, CA · On-site

$250K - $300K/yr

Experience with agent orchestration, LLMs, or web scraping. * Familiarity with distributed systems, cloud infrastructure (AWS or GCP), microservices, and data pipelines. * Background at a high-growth ...

Integration Engineer

Oakland, CA · On-site

$119K - $160K/yr

... of data collection. You will use a custom Python-based macro system and write HTML, text, and OCR parsers to interface with it. OCR and web-scraping experience required . This role can be remote.

Integration Engineer

Oakland, CA · On-site +1

$119K - $160K/yr

... of data collection. You will use a custom Python-based macro system and write HTML, text, and OCR parsers to interface with it. OCR and web-scraping experience required . This role can be remote.

Showing results 21-40

Freelance Web Data Scraping information

What is freelance web data scraping?

Freelance web data scraping involves extracting information from websites on behalf of clients, usually for purposes like research, business intelligence, or data analysis. Freelancers use programming languages and tools, such as Python with libraries like BeautifulSoup or Scrapy, to collect and organize data efficiently. They must ensure their methods comply with legal and ethical guidelines, as not all websites permit data scraping. Clients typically hire freelance web scrapers for specific projects or ongoing data collection needs.

What are the key skills and qualifications needed to thrive as a freelance web data scraper, and why are they important?

To thrive as a Freelance Web Data Scraper, you need strong programming skills in languages like Python, familiarity with HTML/CSS, and a solid understanding of data extraction techniques. Proficiency with tools such as Beautiful Soup, Scrapy, Selenium, and knowledge of APIs is essential, and certifications in data science or web development can be advantageous. Attention to detail, problem-solving abilities, and effective communication help freelancers navigate complex websites and manage client expectations. These skills ensure efficient, ethical, and accurate data extraction, which is crucial for delivering value and maintaining client trust.

What are some common challenges faced by freelance web data scraping professionals, and how can they be addressed?

Freelance web data scraping professionals often encounter challenges such as website structure changes, anti-scraping measures (like CAPTCHAs or IP blocking), and varying data formats. Staying updated with the latest scraping tools and techniques, using rotating proxies, and employing headless browsers can help overcome these obstacles. Building strong communication with clients to clarify data requirements and proactively addressing potential legal or ethical concerns also contributes to successful project outcomes.

What is the difference between Freelance Web Data Scraping vs Freelance Data Mining?

AspectFreelance Web Data ScrapingFreelance Data Mining
CredentialsBasic programming skills, knowledge of web scraping toolsData analysis skills, knowledge of data mining techniques
Work EnvironmentPrimarily remote, project-basedRemote or on-site, often involves data analysis platforms
Industry UsageWebsites, e-commerce, researchBusiness intelligence, market research, analytics
Search & Comparison IntentFocus on extracting data from websitesFocus on analyzing large datasets for insights

Freelance Web Data Scraping involves extracting data directly from websites using programming tools, while Freelance Data Mining focuses on analyzing large datasets to uncover patterns and insights. Both roles require technical skills but serve different purposes in data collection and analysis.

What are the most commonly searched types of Web Data Scraping jobs in California?

The most popular types of Web Data Scraping jobs in California are:

What job categories do people searching Freelance Web Data Scraping jobs in California look for?

The top searched job categories for Freelance Web Data Scraping jobs in California are:

What cities in California are hiring for Freelance Web Data Scraping jobs?

Cities in California with the most Freelance Web Data Scraping job openings:

Senior AI Engineer, Agentic Data Enrichment

Baselayer

San Francisco, CA • On-site

$124K - $169K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 17 days ago


Job description

ABOUT BASELAYER


Every business in America needs a bank account to exist. The system that decides whether they're real, who's behind them, and whether they're a risk, runs on infrastructure from the 1980s. We're rebuilding that layer from scratch.

Baselayer is the identity layer for institutions across the United States - the most complete business graph in America and every human tied to it. We fuse public records, IRS data, sanctions lists, web signals, and fraud telemetry from 2,200+ financial institutions into a single graph that resolves any business and the humans behind it in milliseconds. The legacy credit bureaus took 50 years to build something that gets 60% match rates. We've built something that gets 98% in under two years.

Today we're trusted by over 20% of financial institutions in America - including FIS, Rho, Socure and leading loan infrastructure providers. But the graph is becoming infrastructure for anyone who needs to know if a business is real and worth trusting: gig platforms, marketplaces, AI companies, and commerce infrastructure at scale.

Trust is the substrate of every financial transaction. We're rebuilding it.

ABOUT THE TEAM


We're solving real-time entity resolution at a scale no one else has cracked - fusing dozens of data sources into a single business identity graph and resolving any entity in milliseconds. It's a graph AI problem, a retrieval problem, and a fraud-modeling problem stacked on top of each other. The technical depth is real.

You'd be joining a small team where the data moat is defensible, the research problems are open, and the infrastructure you build becomes load-bearing for businesses. Ownership is real. Velocity is real. There's no layer of process between an idea and shipping it.

We're at an inflection point - the graph is built, the match rates speak for themselves, and the hardest problems are still ahead: graph embeddings, fraud propagation models across the business network, real-time traversal at sub-100ms latency, and expanding the identity layer beyond finance into every platform that needs to trust a business.

If you want to work on something foundational - the kind of infrastructure that gets built once and everything else runs on top of - this is it.

ABOUT THE ROLE


Baselayer answers questions the loan application didn't ask. For every business that crosses our queues, we need to know things that aren't on the form: what the business actually does, where it actually lives on the web, whether the people it names match the public record, and whether anything across the open web contradicts the story we were told. We answer those questions with LLM-driven agents that crawl, click, search, and extract structured evidence from across the web - and we treat this as a production data pipeline, not a research demo. We're hiring a Senior AI Engineer to own a slice of this enrichment surface end-to-end.

WHAT YOU'LL DO


  • Own industry/category classification of businesses from heterogeneous signals (name, website, directory presence, reviews).
  • Build and maintain discovery and verification systems for a business's real web presence - filtering aggregators, parked domains, brand collisions, and impersonators.
  • Link individuals to businesses via public web evidence (e.g. confirming a named officer or employee genuinely works there).
  • Develop risk/legitimacy scoring derived from web-presence signals, fed back into downstream underwriting.
  • Build and evolve the shared agent infrastructure: provider-agnostic base agents, shared toolset registry (browser navigation, search, scraping, structured database lookups, scoring), eval harness, and instrumentation surface for token-and-tool tracing.
  • Own model selection, agent design, prompt and tool engineering, eval methodology, and cost control across your enrichment surface.

MINIMUM REQUIREMENTS


  • Shipped LLM-driven agents to production - not notebooks, not demos. Real users, real cost, real failure modes, real on-call.
  • Strong async Python including structured-data libraries, modern web frameworks, and relational databases.
  • Experience across multiple frontier LLM providers and at least one agent framework, with deep knowledge of failure modes.
  • Built or maintained eval methodology: curated golden datasets, scoring functions, labelling guidelines, regression diagnostics.
  • Browser automation experience: headless browsers, anti-bot evasion, authenticated flows.
  • Holds informed opinions on structured-output reliability - when to use JSON-schema mode vs. function calling vs. extractor-on-top-of-text.

WHAT SETS YOU APART


  • Web scraping at scale: anti-bot evasion, residential proxies, request fingerprinting, authenticated flows, CDN defeats.
  • Eval-framework experience (e.g., LangSmith, Braintrust, Evals, or custom).
  • Entity resolution / record linkage / fuzzy matching at scale.
  • Browser-automation experience at the devtools-protocol level.
  • Built a tool registry or toolset abstraction over multiple LLM providers.
  • Cost/latency optimization: response caching, semantic caching, model routing (cheap-first then escalate), thinking-budget tuning, prompt-cache hit-rate work.

WORK LOCATION


  • Based in SF; hybrid - 4 days per week in office.

COMPENSATION


  • Salary Range: $230,000 - $340,000 + Equity

BENEFITS


  • Time off when you need it: Flexible PTO so you can recharge without red tape.
  • In-person energy: We're based in SF and meet in the office 4 days a week.
  • Competitive compensation: We pay well and back it with equity. We want you to think and act like an owner.
  • Career rocket fuel: You'll help build the foundation of a high-growth startup, working side by side with experienced founders and team members who've done it before.
  • Benefits on us: We cover 100% of your health, dental, and vision premiums. No surprise deductions from your paycheck.
  • 401(k) with company match: We match your contributions so your future self benefits too
  • HSA contributions included: We contribute to your HSA on applicable plans, so your coverage works as hard as you do
  • Stay healthy, stay sharp: A $250 monthly gym stipend to help you bring your best self to work, and everywhere else
  • A seat at the table: We believe in transparency, radical candor, and giving every team member a voice