Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture - fairness, bias detection ...
Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture - fairness, bias detection ...
Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture - fairness, bias detection ...
Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture - fairness, bias detection ...
Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture -- fairness, bias detection ...
Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy ... Responsible AI (RAI) & Governance (10%) * Integrate RAI signal capture -- fairness, bias detection ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Internet Rater - English (United States)
Dallas, TX · Remote
$21 - $28.25/hr
... AI and deep learning. We are beyond proud to serve our clients, our employees, and the end-users of ... As an Internet Rater, you will review and grade internet content to help online engines work better.
Internet Rater - English (United States)
Dallas, TX · Remote
$21 - $28.25/hr
... AI and deep learning. We are beyond proud to serve our clients, our employees, and the end-users of ... As an Internet Rater, you will review and grade internet content to help online engines work better.
AI Governance Coordinator - Remote / Telecommute
Dallas, TX · Remote
$43/hr
Review and validate risk ratings for active and pending AI use cases; adjust ratings as needed based on updated context, control posture, or governance criteria. * Issue and collect supplemental ...
Quick apply
AI Governance Coordinator - Remote / Telecommute
Dallas, TX · Remote
$43/hr
Review and validate risk ratings for active and pending AI use cases; adjust ratings as needed based on updated context, control posture, or governance criteria. * Issue and collect supplemental ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Principal AI Architect/Engineer
Plano, TX · On-site
... false-positive rates. * Participate in on-call rotations and incident response for the ... Responsible AI (RAI) & Governance Signal Instrumentation (10%): * Implement RAI signal collectors ...
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Quick apply
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Define and track success metrics including monthly active users, asset reuse rates, time-to-publish ... Familiarity with AI concepts including hallucination and output reliability, LLM-based tooling ...
Define and track success metrics including monthly active users, asset reuse rates, time-to-publish ... Familiarity with AI concepts including hallucination and output reliability, LLM-based tooling ...
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Cost Accounting Compliance Analyst (R5254)
Dallas, TX · On-site
$100K - $150K/yr
Shield AI is a venture-backed defense-tech company with the mission of protecting service members ... Lead the preparation of incurred cost submissions (ICS), provisional billing rate (PBR), and ...
Define and track success metrics including monthly active users, asset reuse rates, time-to-publish ... Familiarity with AI concepts including hallucination and output reliability, LLM-based tooling ...
Define and track success metrics including monthly active users, asset reuse rates, time-to-publish ... Familiarity with AI concepts including hallucination and output reliability, LLM-based tooling ...
... rates. • Partner with Security and RAI teams to embed threat modeling, zero-trust agent ... for agentic AI systems. • Integrate RAI signal capture -- fairness, bias detection ...
... rates. • Partner with Security and RAI teams to embed threat modeling, zero-trust agent ... for agentic AI systems. • Integrate RAI signal capture -- fairness, bias detection ...
AI Architect - Remote / Telecommute
Dallas, TX · Remote
$57 - $62/hr
Experience leading enterprise AI programs from strategy through production deployment ... Benefits eligibility, accrual rates, and usage limits may vary based on employment status, length ...
Quick apply
AI Architect - Remote / Telecommute
Dallas, TX · Remote
$57 - $62/hr
Experience leading enterprise AI programs from strategy through production deployment ... Benefits eligibility, accrual rates, and usage limits may vary based on employment status, length ...
AI Solution Architect
Plano, TX · On-site
$150 - $190/hr
Pursuit win rate for AI deals where the AI Solution Architect supported. * Utilization or billability against target. * Estimate quality, solution feasibility, commercial accuracy, and successful ...
AI Solution Architect
Plano, TX · On-site
$150 - $190/hr
Pursuit win rate for AI deals where the AI Solution Architect supported. * Utilization or billability against target. * Estimate quality, solution feasibility, commercial accuracy, and successful ...
Senior AI Engineer
Mesquite, TX · On-site
$150K - $170K/yr
Senior AI Engineer Location: (1 days per week onsite/4 days work from home) Employment Type ... Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy ...
Senior AI Engineer
Mesquite, TX · On-site
$150K - $170K/yr
Senior AI Engineer Location: (1 days per week onsite/4 days work from home) Employment Type ... Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy ...
AI Solutions Engineer
Plano, TX · On-site
... rates, clinical/workflow impact measures) * Build LLM/NLP-based comparison and analysis frameworks that systematically evaluate AI outputs against radiologist ground truth or other clinically ...
AI Solutions Engineer
Plano, TX · On-site
... rates, clinical/workflow impact measures) * Build LLM/NLP-based comparison and analysis frameworks that systematically evaluate AI outputs against radiologist ground truth or other clinically ...
Sr Azure AI Engineer
Coppell, TX · On-site
$70/hr
Contract-to-Hire Pay Rate: $70 / Hour Benefits: This position is eligible for medical, dental, vision, and 401(k). Position Overview: We're looking for a senior engineer to help design and scale AI ...
Quick apply
Sr Azure AI Engineer
Coppell, TX · On-site
$70/hr
Contract-to-Hire Pay Rate: $70 / Hour Benefits: This position is eligible for medical, dental, vision, and 401(k). Position Overview: We're looking for a senior engineer to help design and scale AI ...
AI Solution Architect
Plano, TX · On-site +1
$60.25 - $79.50/hr
Success Measures: - Pursuit deal volume, measured as total TCV value of AI deals where the AI Solution Architect supported. - Pursuit win rate for AI deals where the AI Solution Architect supported ...
AI Solution Architect
Plano, TX · On-site +1
$60.25 - $79.50/hr
Success Measures: - Pursuit deal volume, measured as total TCV value of AI deals where the AI Solution Architect supported. - Pursuit win rate for AI deals where the AI Solution Architect supported ...
Ai Rater information
What is an AI Rater?
An AI Rater evaluates and provides feedback on artificial intelligence models, typically improving search engines, chatbots, or recommendation systems. They assess the relevance, accuracy, and quality of AI-generated content based on specific guidelines. This role requires strong analytical skills, attention to detail, and familiarity with the subject matter being reviewed. AI Raters often work remotely and on a flexible schedule.
What skills and qualifications are needed to thrive as an AI Rater?
To thrive as an AI Rater, you generally need strong attention to detail, analytical thinking, and proficiency in English, often supported by formal education such as a high school diploma or higher. Familiarity with web browsers, online research, and company-specific rating platforms or guidelines is essential. Excellent time management, adaptability, and effective written communication help individuals excel in this position. These skills and qualities ensure accurate and consistent evaluations of AI-generated content, directly impacting the improvement of artificial intelligence systems.
What does an AI Rater do?
A typical day for an AI Rater involves reviewing and evaluating various types of content, such as search engine results, social media posts, advertisements, or chatbot responses, to ensure they meet quality and relevancy standards. You may follow detailed guidelines to rate or annotate content, complete assigned tasks in a web-based platform, and provide feedback to help improve AI performance. Most positions are remote and offer flexible schedules, allowing you to plan your workload around personal commitments. Collaboration is generally limited, as most work is performed independently, but periodic communication with team leads for training or updates is common.

Full-time
Medical, Dental, Vision, Life, Retirement, PTO
Re-posted 9 days ago
PepsiCo rating
7.5
Based on 894 frontline employees who took The Breakroom Quiz
154th of 436 rated food and drinks producers
Job description
The AI Observability Architect is a senior technical leader responsible for designing, deploying, and operating an enterprise-grade, production-ready AI observability platform that spans the full spectrum of modern agentic AI - from large language model (LLM) workflows and multi-agent orchestration to physical AI systems, reinforcement learning harnesses, multi-modal pipelines, and agentic marketplaces. This role serves as the strategic and engineering authority for end-to-end telemetry, tracing, safety, and quality signals across heterogeneous agent frameworks and platforms.
The architect leads the convergence of AI observability with safety & security (including red teaming), Responsible AI (RAI), data science, physical AI, memory/skills engineering, agent fleet management, self-evolving harnesses, reinforcement learning, agent-to-agent protocols (A2A, UCP, AP2), and continuous quality engineering - making this a uniquely broad and high-impact role within the AI Solutions & Platforms organization.
The role also owns OpenTelemetry (OTEL) integration across third-party agentic platforms (Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and others), enabling unified observability and governance at enterprise scale.
Responsibilities
Agentic AI Observability Architecture at Scale (30%)
- Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems - covering planner/executor loops, tool/function calls, RAG retrieval chains, and memory/state transitions.
- Build and operate unified telemetry pipelines incorporating metrics, logs, distributed traces, semantic/vector signals, and real-time event streaming (Kafka) at enterprise scale.
- Instrument OpenTelemetry (OTEL) across heterogeneous platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and internal frameworks - delivering protocol-level observability for agent ecosystems including MCP, A2A, UCP, and AP2.
- Design and implement observability for Agent Fleets, multi-modal pipelines, physical AI systems, and self-evolving reinforcement learning harnesses - including signal capture for reward shaping and policy evaluation.
- Deliver dashboards, alerting, SLO/SLA management, incident runbook automation, and RCA tooling that drive measurable reliability improvements and reduce MTTR across agentic services.
- Establish cost telemetry and FinOps observability for AI workloads - token consumption, inference cost allocation, and GPU/compute efficiency across cloud environments (Azure, AWS, GCP).
Safety, Security & Red Teaming (15%)
- Lead observability-driven red team exercises targeting agentic AI systems - instrumenting attack surfaces, adversarial prompt injection vectors, model evasion attempts, and multi-agent trust boundary failures.
- Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy violation events, PII exposure risks, prompt leakage, and agent hallucination rates.
- Partner with Security and RAI teams to embed threat modeling, zero-trust agent authentication, and behavioral anomaly detection into the observability platform.
- Instrument secure policy enforcement layers across agent-to-agent communication protocols (A2A, UCP, AP2) and maintain audit-ready traceability for all AI decision events.
- Develop and maintain a Security Observability Playbook covering incident classification, escalation paths, and forensic trace retention policies for agentic AI systems.
Responsible AI (RAI) & Governance (10%)
- Integrate RAI signal capture - fairness, bias detection, explainability, and safety metrics - directly into observability pipelines, making compliance measurable and audit-ready.
- Deliver governance dashboards that surface RAI compliance posture across all active AI agents and LLM deployments, aligned with global regulatory standards.
- Support risk assessments, gap analyses, and governance frameworks with real-time observability insights - enabling proactive risk mitigation rather than reactive audit responses.
- Collaborate with RAI CoE and Legal/Compliance teams to define data retention, consent logging, and model decision traceability standards embedded in the telemetry architecture.
Quality Engineering for Agentic Solutions - Post Go-Live & Continuous QE (10%)
- Own the Continuous Quality Engineering (CQE) framework for post-production agentic solutions - defining and tracking quality metrics across accuracy, latency, agent success rate, tool-call fidelity, and user outcome measures.
- Build automated quality gates within CI/CD pipelines that leverage observability data to detect regressions, drift, and degradation in agent performance - preventing silent failures in production.
- Instrument and monitor Skill Evaluations (evals) across the Memory, Skills, and MCP harness stack - providing traceability from eval results to production behavior.
- Partner with product and business stakeholders to define SLA-backed quality benchmarks and deliver automated alerting when quality thresholds are breached.
- Drive root-cause analysis for quality failures using distributed trace data, enabling rapid iteration and continuous improvement cycles for agentic solutions.
Memory, Skills, MCP & Harness Engineering Observability (10%)
- Design and implement observability for the agent memory layer - episodic, semantic, and working memory read/write operations - providing latency, accuracy, and drift monitoring across memory backends.
- Instrument MCP (Model Context Protocol) server interactions, tool registrations, skill invocations, and context injection pipelines with full trace propagation and semantic tagging.
- Own observability for self-evolving harness and reinforcement learning (RL) systems - capturing reward signals, policy update events, environment state transitions, and learning convergence metrics.
- Monitor harness execution fidelity, skill eval pass/fail rates, and regression signals across training, fine-tuning, and inference workflows - feeding data back into the quality engineering loop.
Data Science Observability & Hardcore Python Engineering (5%)
- Lead a team of senior Python engineers building high-performance, production-grade observability tooling - including custom OTEL exporters, semantic trace enrichers, signal aggregators, and anomaly detection pipelines.
- Apply data science methods - statistical process control, time-series anomaly detection, clustering, and causal inference - to transform raw telemetry into actionable AI operational intelligence.
- Build and maintain Python-native SDKs and libraries that simplify observability onboarding for agent developers across the organization.
- Establish code quality standards, testing frameworks, and peer review practices for the observability engineering team - embedding software craftsmanship into the team culture.
Agentic Marketplace, Registry & Ecosystem Observability (5%)
- Instrument the Agentic Marketplace and Agent Registry platforms - providing usage telemetry, adoption metrics, capability health scores, and dependency mapping for registered agents and skills.
- Design observability APIs and SDK hooks that allow marketplace-registered agents to self-report health, performance, and behavioral signals into the central observability platform.
- Monitor inter-agent communication patterns across the marketplace ecosystem - identifying latency hotspots, circular dependencies, and protocol mismatches in agent-to-agent (A2A) workflows.
- Deliver a Marketplace Observability Dashboard surfacing agent catalog health, adoption trends, quality scores, and incident history - supporting marketplace governance and curation decisions.
Integration, Deployment & CI/CD Automation (5%)
- Build and maintain CI/CD pipelines for observability services and agent operations center components, incorporating automated testing, deployment gates, and rollback mechanisms.
- Automate onboarding for new agent use cases using templates, scaffolding, and configuration validation - reducing time-to-observability from weeks to hours.
- Drive infrastructure-as-code (IaC) practices for observability platform components across Azure, AWS, and GCP - ensuring reproducible, version-controlled, and auditable deployments.
Product Delivery & Stakeholder Collaboration (10%)
- Operate with a product mindset - defining observability platform roadmaps, OKRs, adoption playbooks, and release milestones in partnership with AI platform and business teams.
- Collaborate with transformation teams, enterprise architects, security, and business stakeholders to tailor observability solutions to domain-specific requirements.
- Serve as the technical authority in executive and governance forums - translating complex observability data into business-relevant insights on risk, cost, and AI performance.
- Partner with SRE, AI platform, and product teams to drive standard adoption and reduce integration friction across the agentic AI ecosystem.
People Leadership & Team Development (5%)
- Build, mentor, and lead a high-performing observability engineering team - spanning Python developers, data scientists, and platform engineers - with talent initially based in India.
- Define career paths, skills development plans, and leveling criteria aligned with PepsiCo job architecture - fostering an inclusive, high-accountability team culture.
- Drive hiring, coaching, performance management, and succession planning across the observability function.
Decision-Making Autonomy
- High - Owns architecture decisions, platform roadmap, and engineering standards. Strategic alignment sought from AI Solutions Director on enterprise-level commitments.
Supervision Required
- Low to Moderate - Operates independently with periodic alignment reviews. Proactively escalates cross-organizational dependencies and risk trade-offs.
Role Complexity
- Very High - Spans observability, safety/security, RL harnesses, physical AI, multi-modal systems, agent protocols, quality engineering, and marketplace governance simultaneously
Compensation and Benefits:
- The expected compensation range for this position is between $123,500 - $206,750.
- Location, confirmed job-related skills, experience, and education will be considered in setting actual starting salary. Your recruiter can share more about the specific salary range during the hiring process.
- Bonus based on performance and eligibility target payout is 15% of annual salary paid out annually.
- Paid time off subject to eligibility, including paid parental leave, vacation, sick, and bereavement.
- In addition to salary, PepsiCo offers a comprehensive benefits package to support our employees and their families, subject to elections and eligibility: Medical, Dental, Vision, Disability, Health, and Dependent Care Reimbursement Accounts, Employee Assistance Program (EAP), Insurance (Accident, Group Legal, Life), Defined Contribution Retirement Plan.
Qualifications
Minimum Education & Experience:
- Bachelor's or Master's degree in Computer Science, AI/ML, Data Science, Software Engineering, or a related field (PhD a plus for research-heavy domains).
- 12+ years in technology with deep experience in enterprise observability, distributed systems, platform engineering, or AI/ML infrastructure.
- 5+ years in a senior/principal or architect-level role with demonstrated ownership of complex, cross-functional technical programs.
Core Technical Qualifications
- AI Observability & Distributed Systems: Expert-level knowledge of observability primitives (metrics, logs, traces, events) applied to LLM/ML/agentic systems; hands-on OpenTelemetry (OTEL) instrumentation including custom exporters, semantic conventions, and trace propagation across agent/tool boundaries.
- Agentic AI Frameworks: Direct experience with agentic AI platforms, multi-agent orchestration, LLM-based workflow design, and agent lifecycle management at production scale.
- Safety, Security & Red Teaming: Demonstrated experience conducting red team exercises against AI systems; knowledge of adversarial attack patterns, prompt injection, model evasion, and multi-agent trust boundary failures; ability to design safety telemetry pipelines.
- Memory, Skills & MCP: Working knowledge of agent memory architectures (episodic, semantic, working memory), Model Context Protocol (MCP), skill registries, and context injection patterns - with ability to design observability for these layers.
- Agent-to-Agent Protocols: Familiarity with A2A (Agent-to-Agent), UCP (Universal Communication Protocol), and AP2 patterns; ability to implement protocol-level observability and policy enforcement.
- Reinforcement Learning & Self-Evolving Harnesses: Understanding of RL training loops, reward signal capture, policy evaluation, and harness instrumentation for continuously improving agent systems.
- Physical AI & Multi-Modal Systems: Experience or strong familiarity with observability for physical AI pipelines (robotics, edge inference, sensor fusion) and multi-modal models (vision, audio, text).
- Data Science & Python Engineering: Proficiency in Python at a senior engineering level; experience with statistical anomaly detection, time-series analysis, and data pipeline design applied to observability data at scale.
- Platform Integrations (OTEL / Enterprise): Hands-on experience integrating OTEL with enterprise agentic platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, or similar; strong understanding of enterprise integration patterns and API design.
- Cloud & Infrastructure: Cloud fluency across Azure, AWS, and GCP; proficiency in Kubernetes, service mesh, IaC (Terraform/Bicep), and CI/CD tooling; exper
About PepsiCo
Sourced by ZipRecruiter
PepsiCo products are enjoyed by consumers more than one billion times a day in more than 200 countries and territories around the world. PepsiCo generated $86 billion in net revenue in 2022, driven by a complementary beverage and convenient foods portfolio that includes Lay's, Doritos, Cheetos, Gatorade, Pepsi-Cola, Mountain Dew, Quaker, and SodaStream. PepsiCo's product portfolio includes a wide range of enjoyable foods and beverages, including many iconic brands that generate more than $1 billion each in estimated annual retail sales.
Industry
Food and drink manufacturing
Company size
10,000+ Employees
Headquarters location
Purchase, NY, US
Year founded
1965