Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust) "Friends of ... Equal Opportunity Employer The Voleon Group is an Equal Opportunity employer. Applicants are ...
Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust) "Friends of ... Equal Opportunity Employer The Voleon Group is an Equal Opportunity employer. Applicants are ...
Senior Cluster Site Reliability Engineer
Berkeley, CA · On-site +1
$205K - $235K/yr
Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust) "Friends of ... Equal Opportunity Employer The Voleon Group is an Equal Opportunity employer. Applicants are ...
Senior Cluster Site Reliability Engineer
Berkeley, CA · On-site +1
$205K - $235K/yr
Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust) "Friends of ... Equal Opportunity Employer The Voleon Group is an Equal Opportunity employer. Applicants are ...
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Principal HPC Architect
Milpitas, CA · On-site
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Principal HPC Architect
Milpitas, CA · On-site
Group/Division The Information Technology (IT) group at KLA is involved in every aspect of the ... Implement security best practices, patching, and compliance controls. * Ensure high availability ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
... APP) group and will work to design and deliver hybrid cloud solutions with a focus on high ... S. Government security clearance. This full-time, direct hire position offers a full benefits ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
... APP) group and will work to design and deliver hybrid cloud solutions with a focus on high ... S. Government security clearance. This full-time, direct hire position offers a full benefits ...
Singularity Security Group information
What are the typical responsibilities of professionals working at Singularity Security Group?
Professionals at Singularity Security Group are responsible for monitoring network traffic, identifying and mitigating threats, conducting vulnerability assessments, and implementing security protocols. They often work as part of a dynamic team that collaborates with IT departments, management, and clients to ensure comprehensive protection of digital assets. Regular tasks include responding to incidents, analyzing security data, maintaining compliance with industry standards, and contributing to security strategy development. This role offers ongoing learning opportunities due to the constantly evolving landscape of cybersecurity threats.
What is a Singularity Security Group job?
A Singularity Security Group job typically involves working in cybersecurity, focusing on protecting organizations from advanced threats using AI-driven security solutions. Professionals in this group may analyze cyber risks, develop security protocols, and respond to incidents using cutting-edge technology. They often work with machine learning, automation, and threat intelligence to stay ahead of evolving cyber threats.
What are the key skills and qualifications needed to thrive in the Singularity Security Group position, and why are they important?
To thrive in a Singularity Security Group position, candidates typically need expertise in cybersecurity, risk assessment, and incident response, along with a background in IT or information security. Familiarity with security operations tools, SIEM platforms, and certifications like CISSP or CEH are highly valued in this role. Strong analytical thinking, effective communication, and teamwork skills help professionals address threats and collaborate with stakeholders. These competencies are critical to safeguarding organizational assets, ensuring compliance, and maintaining a resilient security posture.

$15K/mo
Full-time
Re-posted yesterday
Job description
Voleon is a technology company that applies state-of-the-art AI and machine learning techniques to real-world problems in finance. For nearly two decades, we have led our industry and worked at the frontier of applying AI/ML to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.
Your colleagues will include internationally recognized experts in artificial intelligence and machine learning research as well as highly experienced finance and technology professionals. In addition to our enriching and collegial working environment, we offer highly competitive compensation and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.
As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale. You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.
The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs.
Responsibilities-
Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise
-
Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability
-
Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams
-
Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won\'t do
-
Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies
-
Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability
-
5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead
-
Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)
-
Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
-
Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)
-
Experience with cloud infrastructure (AWS or GCP)
-
Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)
-
Experience with distributed storage technologies (Lustre, Ceph, S3)
-
Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation
-
Bachelor degree in computer science
-
Hands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark)
-
Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed)
-
Familiarity with hybrid/on-prem environments
-
Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments
-
Experience with HPC networking (InfiniBand, RDMA)
-
Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust)
If you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Voleon Referral Bonus Program.
Equal Opportunity EmployerThe Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
About Voleon Group
Sourced by ZipRecruiter
Industry
Investment management and consulting services
Company size
11 - 50 Employees
Headquarters location
Berkeley, CA, US
Year founded
2007