Company Description
Welcome to Our Worldย We've been leading the charge in the affiliate industry from day one-establishing performance marketing and paving the way for future innovations. We're known for maintaining one of the largest, most reliable partnership platforms with impeccable, personalized service.ย ย Founded in Santa Barbara, California in 1998, CJ (formerly Commission Junction) stands as the most trusted name in performance marketing. We specialize in building partnerships between top brands and reputable publishers to drive revenue and business growth. CJ's industry-leading solutions make us the platform of choice for over 3,800 global brands across sectors like retail, travel, finance, technology, and home services. As part of Publicis Groupe, our savvy data capabilities, cutting-edge tech, and strategic expertise facilitate genuine connections, allowing brands to reach consumers wherever they are.ย ย A Quick Peek at Affiliate Marketingย Think back to your last online purchase. Did an influencer tip you off about a great product and offer a discount? Or perhaps you relied on a trusted review site to make your decision? Whatever path you took, affiliate publishers likely played a role by influencing, informing, or helping you find the best deal. CJ connects brands with these publishers, creating valuable resources for shoppers like you.ย
Job Description
You must be work authorized in the United States without the need for employer sponsorship.ย
This is a hybrid role requiring 3 days a week in a CJ office location (www.cj.com)
Aboutย CJย Engineeringย
Areย we a good fit for you?ย
At CJ, we are passionate about software engineering. We build exciting software, with quality andย maintainabilityย in mind. Weย believe in common sense,ย simplicity,ย and efficiency.ย Weย practiceย criticalย thinking, challenge each other no matter the title, and believe in the wisdom of theย team.ย
That makes usย Engineers and not just developers.ย Here are some of the principles thatย setย CJ engineersย apart.ย ย
Fullย stack:ย Expect to be involvedย andย toย gain competenceย in every aspect of software engineering, from frontend, to database, to requirements analysis,ย to testing,ย to helping chooseย technologies.ย
Ownership:ย Engineers own the full lifecycle of what they build-from design and implementation to deployment, monitoring, production support, and on-call rotations. If we build it, weย supportย it.ย
Operational Excellence:ย We embrace Infrastructure as Code, CI/CD, automation, and observability to build reliable systems and deliver software safely, efficiently, and at scale.ย
We believe in Agile values, andย incremental development. We constantly experiment, retrospect, and adjust.ย
If this sounds exciting,ย we want to hear from you!ย
As a Senior Software Engineer on the Engineering Experience (EngExp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable.ย EngExpย owns all of it, and this role touches most of it. This is not just an infrastructure role - your value is in engineering judgment. You are a leader on the team: you own the reliability and operability of the platform, act as a critical reviewer of systems and changes, set standards, and mentor other engineers. We especially want someone with real depth in the systems below - we have made platform decisions we later had to reverse because the team lacked deepย expertiseย in aย componentย we owned, and that depth is exactly what this role brings.ย
Responsibilities
The Systems You Work On:
EngExpย owns the systems below.ย You'llย own their reliability, be the critical reviewer of changes to them, and raise the team's depth in them:ย
- Observability & monitoring -ย Prometheus,ย Alertmanager, Grafana, andย OpenTelemetryย across every production region. This is not dashboard-building: you own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store strategy), and ownย Alertmanagerย HA and the blast radius of alert-routing config. Deep Prometheus and Alertmanagerย expertiseย isย a core requirement.ย
- Kubernetes & cloud infrastructure -ย multi-region EKS clusters: upgrades, node group andย Karpenterย management, controller lifecycle, and add-on / configuration management.ย Identifyย failure modes before they happen (subnet IP exhaustion, API server latency,ย ArgoCDย reconciliation lag, Prometheus cardinality,ย Karpenterย consolidation disruption).ย
- AWS networking -ย VPC design, subnetย allocationย and CIDR management, VPC peering, Transit Gateway, security groups, Route53, and NAT gateway topology across multiple accounts and regions, plus 24/7 networking alarms for all prod networking between clusters and squad resources.ย
- CI/CD & artifact management -ย GitLab administration (runner fleet, AMI updates, cache, access - not just pipeline authoring),ย GitOpsย delivery throughย ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows.ย
- Access & identity -ย Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management (including cost alerts and reporting).ย
- Cost observability -ย OpenCost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution across teams, so waste is attributable rather than shared overhead.ย
- Internal tools & delivery -ย the container image build pipeline and base image standards, code audit tooling,ย HedgeDoc, and the UI CDN (S3 + CloudFront), plus adopted applications with no other owner.ย
What You'll Do:
- Own the reliability and operability of the systems above - focused on what is happening and whyย
- Establish and enforce platform standards: RBAC, admission webhooks, resource limits,ย LimitRanges, policy-as-codeย
- Manage infrastructure-as-code with Terraform across AWS accountsย
- Act as a high-quality reviewer of infrastructure changes - Terraform, Kubernetes configs, CI/CD pipelines, observability config - catching subtle issues and long-term risks before they shipย
- Turn recurring requests (ingress, DNS, service accounts) into self-service workflows that are hard to misuseย
- Drive resolution of platform incidents with a focus on learning and lasting system improvementย
- Evaluate new patterns (Gateway API /ย Kgateway, claim-based self-service) on tradeoffs, notย hypeย
- Mentor less-senior engineers and raise the team's depth in the components we ownย
Technologies We Use:
- Kubernetes / EKS (multi-cluster, multi-region),ย Karpenter, cert-manager, external-dnsย
- Prometheus,ย Alertmanager, Grafana,ย OpenTelemetryย (and long-term storage / sharding for Prometheus)ย
- AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions)ย
- Terraform, AWS (IAM, EKS, S3, EBS)ย
- ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelinesย
- Vault,ย OpenCostย
- Gateway API /ย Kgatewayย
- Kubernetes controllers/operators (reconciliation patterns, restart safety) - Go experience is a plus, notย requiredย
Qualifications
What We Look For:
- 6+ years of experience in software and/or infrastructure engineeringย
- Bachelor's degree or equivalent experienceย
- Deep, hands-on production experience operating Kubernetes and AWS at scale, across multiple accounts and regionsย
- Real operational depth in at least one system we own beyond the cluster - most importantly the observability stack (Prometheus/Alertmanagerย at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We are filtering for people who have run these systems, not justย usedย them.ย
- Strong AWS networking judgment (VPC, peering, Transit Gateway, subnet/CIDR design)ย
- A track recordย as a critical reviewer - spotting subtle infrastructure issues and long-term risks before they shipย
- Experience leading technical work and mentoring engineers; can manage, clarify, and plan around uncertaintyย
- Effective communication and the ability to influence design in a product-focused wayย
Nice to Have:
Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent) run in productionย
- Experience owning a container image / base image pipelineย
- Policy-as-code (Kyvernoย / OPA) and admission webhook designย
- Building claim-based self-service platform capabilitiesย
- ย
What Success Looks Like:
- Engineers can deploy and debug services without needing platform interventionย
- The systems we own are understood deeply enough that we stop making decisions weย have toย reverseย
- Production issues are understood quickly, not just reacted toย
- Platform changes are intentional and low-risk, and standards are clear and enforcedย
Additional Information
This is a hybrid role requiring 3 days a week in office.CJ is the leader in Performance Marketing. We take pride in our innovative technology, comprehensive data solutions and our people. We equip our teams with advanced tools, training and career development opportunities all to provide modern solutions, strategies and support to deliver high quality results for our clients. We work in an enthusiastic, collaborative team setting that values outstanding performance.
We're a community of creative and passionate problem solvers who go the distance to tackle the tough questions, think creatively, and drive resourceful growth, for our clients-and ourselves. We foster and embody an inclusive and collaborative culture where diverse perspectives are sought, relationships are valued, and people feel accepted with a sense of belonging in expressing themselves authentically. We pride ourselves in having a workplace environment that values both work and play.
Why Our Workplace Stands Out Apart from offering competitive salaries, 401K matching, wellness programs, and comprehensive medical, dental, and vision coverage, we provide: Flexible time off without the hassle of accrual A generous number of paid holidays Company-sponsored team-building events An Employee Referral Program Annual recognition awards Hybrid work arrangements for optimal work-life balance Parental bonding leave Backup care options for children and elders An employee discount program International SOS program for global support Business Resource Groups, where employees connect over shared interests to cultivate an engaging, inclusive environment
...and those are just a few of our great perks! Come join us and see what makes our company a great place to work.
If you require accommodation or assistance with the application or onboarding process specifically, please contact USMSTACompliance@publicis.com.
All your information will be kept confidential according to EEO guidelines. #LI-DT1
ย Compensation Range: USD $112,290.00 - USD $172,032.00/Annually. This is the pay range the Company believes it will pay for this position at the time of this posting. Consistent with applicable law, compensation will be determined based on the skills, qualifications, and experience of the applicant along with the requirements of the position, and the Company reserves the right to modify this pay range at any time. Temporary roles may be eligible to participate in our freelancer/temporary employee medical plan through a third-party benefits administration system once certain criteria have been met. Temporary roles may also qualify for participation in our 401(k) plan after eligibility criteria have been met. For regular roles, the Company will offer medical coverage, dental, vision, disability, 401k, and paid time off. The Company anticipates the application deadline for this job posting will be 9/27/2026.Employment Type: FULL_TIME