Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices * Support production operations across Kubernetes, queue-driven systems, and multi-cloud ...
Quick apply
Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices * Support production operations across Kubernetes, queue-driven systems, and multi-cloud ...
Quick apply
Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices * Support production operations across Kubernetes, queue-driven systems, and multi-cloud ...
Glastonbury, CT · On-site +1
$180K - $250K/yr
Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices * Support production operations across Kubernetes, queue-driven systems, and multi-cloud ...
Glastonbury, CT · On-site +1
$180K - $250K/yr
Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices * Support production operations across Kubernetes, queue-driven systems, and multi-cloud ...
Glastonbury, CT · Remote
$57 - $75.75/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Glastonbury, CT · Remote
$57 - $75.75/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Glastonbury, CT · On-site +1
$57 - $75.75/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Glastonbury, CT · On-site +1
$57 - $75.75/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Glastonbury, CT · Remote
$58.25 - $77.50/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Quick apply
Glastonbury, CT · Remote
$58.25 - $77.50/hr
Hands-on Kubernetes operations experience, including scalability and workload reliability * Experience with observability tools, monitoring strategies, and incident management practices * Proficiency ...
Berlin, CT · Hybrid
$52.25 - $71.50/hr
Implement GitOps and advanced container orchestration techniques across Azure Kubernetes Serv ice ( AKS) , OpenShift , and Docker . * Ensure end-to-end automation, observability, and security ...
Berlin, CT · Hybrid
$52.25 - $71.50/hr
Implement GitOps and advanced container orchestration techniques across Azure Kubernetes Serv ice ( AKS) , OpenShift , and Docker . * Ensure end-to-end automation, observability, and security ...
Berlin, CT · On-site
$52.25 - $71.50/hr
Implement GitOps and advanced container orchestration techniques across Azure Kubernetes Serv ice ( AKS) , OpenShift , and Docker . * Ensure end-to-end automation, observability, and security ...
Berlin, CT · On-site
$52.25 - $71.50/hr
Implement GitOps and advanced container orchestration techniques across Azure Kubernetes Serv ice ( AKS) , OpenShift , and Docker . * Ensure end-to-end automation, observability, and security ...
Shelton, CT · On-site
$53 - $72.50/hr
Hands-on experience with containerization and orchestration, including Docker, Kubernetes, or ECS * Multi-cloud experience across both AWS and Azure * Experience with observability and monitoring ...
Shelton, CT · On-site
$53 - $72.50/hr
Hands-on experience with containerization and orchestration, including Docker, Kubernetes, or ECS * Multi-cloud experience across both AWS and Azure * Experience with observability and monitoring ...
Newington, CT · On-site
Experience with MLOps/LLMOps ecosystems, including tools such as MLflow, Kubernetes, LangChain ... detection, observability, and rollback processes. * Drive continuous improvement of model ...
Newington, CT · On-site
Experience with MLOps/LLMOps ecosystems, including tools such as MLflow, Kubernetes, LangChain ... detection, observability, and rollback processes. * Drive continuous improvement of model ...
... monitoring, observability, and continuous integration/deployment Portfolio & Program Management ... Kubernetes) Understanding of agentic automation frameworks such as LangChain, LangGraph, or similar ...
... monitoring, observability, and continuous integration/deployment Portfolio & Program Management ... Kubernetes) Understanding of agentic automation frameworks such as LangChain, LangGraph, or similar ...
Middlebury, CT · On-site +1
Champion best practices in AI/ML operations, including monitoring, observability, and continuous ... Familiarity with cloud platforms (AWS) and container orchestration (OpenShift, Kubernetes)
Middlebury, CT · On-site +1
Champion best practices in AI/ML operations, including monitoring, observability, and continuous ... Familiarity with cloud platforms (AWS) and container orchestration (OpenShift, Kubernetes)
Enhancing systems' observability with proper metrics, monitors, and alerts. * Reading and ... Modern DevOps tools such as Terraform, Docker, and Kubernetes. * Data Streaming Mechanisms such as ...
Enhancing systems' observability with proper metrics, monitors, and alerts. * Reading and ... Modern DevOps tools such as Terraform, Docker, and Kubernetes. * Data Streaming Mechanisms such as ...
Enhancing systems' observability with proper metrics, monitors, and alerts. * Reading and ... Modern DevOps tools such as Terraform, Docker, and Kubernetes. * Data Streaming Mechanisms such as ...
Enhancing systems' observability with proper metrics, monitors, and alerts. * Reading and ... Modern DevOps tools such as Terraform, Docker, and Kubernetes. * Data Streaming Mechanisms such as ...
Full-time
Re-posted 16 days ago
WHO WE ARE
Finalsite is the most valued partner for K–12 schools to build trust, strengthen community, and grow enrollment. Ranked among the best EdTech Companies in America, Finalsite supports more than 7,000 schools and districts worldwide with an integrated platform for websites, communications, mobile apps, enrollment, and marketing services.Headquartered in Glastonbury, Connecticut, Finalsite is a global company with employees working remotely across nearly every U.S. state, as well as throughout Europe, South America, and Asia.
We believe people do their best work when they feel supported, connected, and empowered to grow. That’s why we invest in our employees through competitive benefits, professional development opportunities, and a collaborative culture built on partnership and purpose. Whether you’re looking to expand your skills, take on new challenges, or make a meaningful impact in education, Finalsite offers the opportunity to grow your career while helping schools thrive.
At Finalsite, every interaction matters — with our clients, with each other, and with the schools and families we serve. Join us and help shape stronger school communities around the world
SUMMARYThe Staff Engineer, Communications Platform will provide technical leadership for Finalsite’s high-scale communications infrastructure supporting emergency alerts, district notifications, and parent communications. This role will lead platform modernization, reliability improvements, observability initiatives, and long-term architectural strategy across a multi-cloud environment.
The Staff Engineer will inherit a mission-critical platform that is functional but fragile — a legacy messaging core operating across a fragmented multi-cloud environment with limited visibility into emerging issues before they impact customers. Initial priorities include strengthening observability, improving operational discipline, and increasing platform reliability.
In parallel, the Staff Engineer will lead a modernization initiative that previous efforts were unable to complete. Responsibilities include evaluating existing architectural work, defining the future-state migration strategy, and driving execution through completion while maintaining platform stability and performance.
The role will also serve as a key driver of engineering excellence by establishing stronger architecture practices, improving cross-functional coordination, and ensuring technical decisions are made proactively, strategically, and with long-term scalability in mind.
LOCATION100% Remote - Anywhere within the US
RESPONSIBILITIESLead technical strategy, modernization planning, and operational reliability for the communications platform
Define and implement scalable architecture, migration, and infrastructure improvement initiatives
Improve observability through monitoring, alerting, dashboards, SLOs, and incident management practices
Support production operations across Kubernetes, queue-driven systems, and multi-cloud environments
Establish architecture review standards and collaborate with cross-functional engineering teams
Provide technical leadership during incidents, root cause analysis, and operational reviews
Promote Infrastructure-as-Code (IaC) and operational best practices across the engineering organization
Utilize AI-assisted development tools to support codebase navigation and platform evolution
BASIC QUALIFICATIONS
At least 12 years of experience as a Staff Engineer or senior-level individual contributor supporting high-scale communications or messaging platforms
Strong background in Kubernetes, cloud infrastructure, and distributed systems operations
Experience leading modernization or migration initiatives for production systems
Proficiency with observability tools, monitoring frameworks, and incident response practices
Experience working with Infrastructure-as-Code (IaC) methodologies and cloud-native technologies
Strong Python development and operational experience
Experience using AI-assisted development tools such as Claude Code, Codex, or similar technologies
Strong communication, technical leadership, and cross-functional collaboration skills
PREFERRED QUALIFICATIONS
Experience supporting C/C++ production systems
Familiarity with messaging and notification infrastructure, including email or SMS delivery systems
Experience with AWS-native messaging technologies such as SQS, Lambda, DynamoDB, or KEDA
Prior experience supporting large-scale modernization or migration initiatives
Background supporting EdTech, SaaS, or other highly available customer-facing platforms
Finalsite offers 100% fully remote employment opportunities, however, these opportunities are limited to permanent residents of the United States. Current residency, as well as continued residency, within the United States is required to obtain (and retain) employment with Finalsite.
DISCLOSURESFinalsite is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. EEO is the Law. If you have a disability or special need that requires accommodation, please contact Finalsite's People Operations Team. Finalsite is committed to the full inclusion of all qualified individuals. As part of this commitment, Finalsite will ensure that persons with disabilities or special needs are provided a reasonable accommodation. Ensure your Finalsite job offer is legitimate and don't fall victim to fraud. Ask your recruiter for a phone call or other type of verbal communication and ensure all email correspondence is from a finalsite.com email address. For added security, where possible, apply through our company website at finalsite.com/jobs.
Sourced by ZipRecruiter
Software development
51 - 200 Employees
Glastonbury, CT, US
1998