1

Buildkite Jobs in Missouri (NOW HIRING)

Buildkite information

What is Buildkite and what does a Buildkite engineer do?

Buildkite is a continuous integration and continuous deployment (CI/CD) platform that helps software teams automate the building, testing, and deployment of their code. A Buildkite Engineer is responsible for configuring and optimizing Buildkite pipelines, integrating Buildkite with source control systems, and ensuring efficient, reliable automation for software delivery. They often collaborate with developers and DevOps teams to troubleshoot build issues, improve workflow efficiency, and maintain secure, scalable CI/CD processes.

What are the key skills and qualifications needed to thrive as a Buildkite engineer, and why are they important?

To thrive as a Buildkite Engineer, you need strong experience in CI/CD pipelines, scripting languages (like Bash, Python, or Ruby), and a background in software development or DevOps. Familiarity with Buildkite's platform, cloud infrastructure (such as AWS or GCP), and tools like Docker, Git, and Kubernetes is typically required. Excellent problem-solving, communication, and collaboration skills help you work effectively with development and operations teams. These abilities ensure efficient automation, smooth deployments, and robust software delivery processes.

What are some typical challenges faced by engineers working with Buildkite pipelines, and how can they be addressed?

Engineers working with Buildkite pipelines often encounter challenges related to pipeline configuration complexity, managing secrets securely, and optimizing build times. To address these, it's important to modularize pipeline steps for maintainability, use environment variables or secret management plugins for sensitive data, and leverage Buildkite's parallelism and agent scalability features to speed up builds. Collaborating closely with development and DevOps teams ensures that best practices are shared and pipelines remain efficient and secure.

What is the difference between Buildkite vs Jenkins?

AspectBuildkiteJenkins
Required CredentialsCloud account, API accessServer setup, Java knowledge
Work EnvironmentCloud-based, SaaS platformSelf-hosted or cloud, open-source
Industry UsageDevOps, CI/CD pipelinesDevOps, CI/CD, automation
Common Search/ComparisonYesYes

Buildkite and Jenkins are both popular CI/CD tools used to automate software testing and deployment. Buildkite offers a cloud-based, user-friendly platform with minimal setup, ideal for teams seeking quick deployment. Jenkins, on the other hand, is an open-source, self-hosted solution with extensive plugin support, suitable for organizations needing customizable pipelines. While Buildkite emphasizes ease of use and cloud integration, Jenkins provides more control and flexibility for complex workflows.

What are popular job titles related to Buildkite jobs in Missouri?

For Buildkite jobs in Missouri, the most frequently searched job titles are:

Infographic showing various Buildkite job openings in Missouri as of August 2026, with employment types broken down into 98% Full Time, and 2% Contract. Highlights an 58% Physical, 8% Hybrid, and 34% Remote job distribution.

Member of Technical Staff - ML Operations

Veeda

California, MO • On-site

$120 - $190/hr

Other

Posted 4 days ago


Job description

ABOUT US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

RESPONSIBILITIES
  • Run Lifecycle & Launch Tooling: Own how a run is defined, launched, resumed, and killed, from typed configs (Hydra, OmegaConf) to pinned container digests to a relaunch that takes one command.

  • Experiment Tracking & Provenance: Bind every checkpoint to its code commit, config hash, dataset version, and container digest in Weights & Biases or MLflow, so an old run rebuilds from its manifest, not from memory.

  • Checkpoint Registry & Lineage: Own retention and garbage-collection policy, PyTorch DCP resharding and format conversion, and promotion from raw checkpoint to evaluated artifact that simulation and robotics can safely build on.

  • Evaluation in CI: Gate each checkpoint on seeded rollout and policy-success suites, run per-change and nightly as Slurm arrays under Argo Workflows, with confidence intervals wide enough to separate regression from eval noise.

  • Goodput & Incident Response: Carry the pager for live runs (loss spikes, throughput cliffs, data loader stalls), and report goodput against allocated GPU-hours as the number capacity decisions actually run on.

REQUIREMENTS
  • You have a Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or a related technical field.

  • You have strong Python and software engineering skills and real CI/CD experience (e.g., GitHub Actions, Buildkite), and you have shipped internal tooling that other engineers chose to keep using.

  • You have operated multi-node training jobs, carried the pager for them, and decided from telemetry whether to kill, requeue, or let a degraded run ride.

  • You have built reproducible pipelines end to end, and can say precisely which parts of a training run are bit-reproducible, which are not, and why.

  • You are rigorous about evaluation methodology, from seeds and sample sizes to confidence intervals, and can tell a genuine regression from a flaky harness.

NICE TO HAVE
  • You have run experiment tracking at scale, logging video, 3D, and trajectory artifacts rather than only scalars.

  • You have built evaluation harnesses for generative or embodied models, where quality is a distribution rather than a pass/fail.

  • You have orchestrated ML workflows with Argo Workflows, Flyte, or Ray, and know where each one breaks.

  • You are fluent with Prometheus, Grafana, and OpenTelemetry, and instrument a training job before its first outage.

  • You have built GPU-hour attribution that maps cluster spend back to specific experiments and teams.

  • You have written a postmortem that changed how a team ran jobs, not just what it logged.

  • You have contributed to open-source ML tooling, or published on evaluation or reproducibility methodology.

#J-18808-Ljbffr