Staff Site Reliability Engineer
Fuld tid
Garner Health
Responsibilities
- Own the end-to-end reliability, performance, and resilience strategy for Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads.
- Architect the SLO framework and set technical direction for production quality and reliability at organizational scale.
- Lead the incident response program, complex escalations, on-call improvements, root-cause analysis, and post-incident corrective actions.
- Own the monitoring, alerting, and observability platform across engineering teams.
- Convert ambiguous scaling and reliability requirements into automated, composable Terraform infrastructure-as-code deliverables.
- Drive cloud cost-efficiency, performance optimization, technical debt reduction, and operational automation.
- Establish deployment and observability standards, mentor engineers, and raise operational rigor across the organization.
- Ensure infrastructure and operations meet security and HIPAA compliance obligations.
Requirements
- 7+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering roles.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment, preferably AWS.
- Experience designing reliability practices including SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
- Strong Python or Go skills applied to infrastructure automation; Kubernetes API experience is a plus.
- Demonstrated experience optimizing cloud costs and performance across compute, storage, and networking.
- Experience mentoring engineers and setting technical direction as a senior reliability leader.
- Strong communication skills with both technical and non-technical stakeholders.
- Fluency with AI tools such as Claude in engineering and operations workflows, or strong motivation to develop that capability.
- Experience supporting AI/ML or data-intensive production workloads is preferred.
- Experience in security-conscious or regulated environments, including HIPAA or SOC 2, is preferred.
Benefits
- Remote work is available for candidates comfortable with occasional travel to Garner’s New York City headquarters.
- Flexible paid time off.
- Medical, dental, and vision plan options.
- 401(k) plan.
- Teladoc Health and other competitive benefits.
- Eligibility for equity incentive participation.
Ledig stilling opslået Indrykket for 24 dage siden
Tilsvarende jobs, som kan have interesseBaseret på Staff Site Reliability Engineer vedrørende den ledige stilling i Fjernarbejde
- ...How You'll Make an Impact: As a Staff Software Engineer in Revenue Intelligence, you will shape... ...establishing architecture, improving reliability and security, reducing systemic technical... ...and Recruiting Agencies: Our Careers Site is only for individuals seeking a job...
- ...Make an Impact: As a Staff Machine Learning Engineer , you will play a key role... ...be focused on ensuring the reliability, scalability, and high... ...—ensuring hotels have the reliable, accurate insights they need... ...Recruiting Agencies: Our Careers Site is only for individuals...
- ...improve large-scale production systems through maintainable code, reliability practices, on-call participation, and incident response.... ...Improve customer and internal user experience through performance engineering, observability adoption, and automation. Apply technical...
- ...opportunities and communicate model performance in terms of business impact. Own and prioritize the ML roadmap with product, engineering, and risk operations. Mentor ML engineers and raise technical standards through influence, example, and durable engineering practices...
185000 $ - 240000 $ om året
...robust, scalable, and well-tested full-stack code. Establish engineering best practices and help build a high-performing, collaborative... ...across unit, integration, and end-to-end tests. ~ Demonstrated staff-level technical leadership, including architecture ownership...2000 $ om året
...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's leading... ...profitable, and growing. This is a general track for Senior+ (Senior/Staff/Principal) Engineers in any team at Canonical. After the first...- ...budgets. Prototype, validate, and build critical perception components hands-on. Partner with functional safety and systems engineering on safety-case strategy and required evidence. Support system-level failure mode and hazard analysis for autonomous robotic...
- ...drive architectural decisions with lasting platform impact. Turn ambiguous technical challenges into scalable solutions that engineering teams can execute. Build shared systems, patterns, and platform capabilities that improve engineering delivery. Partner with...
2000 $ om året
...This is the general track for Engineering Director at Canonical, apply here if you are confident... ...and managing engineering managers and staff engineers. Canonical’s largest software... ...is to make open source easier, more reliable and more secure for deployment and development...- Site Activation Specialist - Sponsor-dedicated; 0.8 FTE Syneos Health® is a leading fully-integrated life sciences services organization... ...skills Ability to mentor, lead and motivate more junior staff Demonstrate an ability to provide quality feedback and...
2000 $ om året
...such as public cloud, data science, AI, engineering innovation and IoT. Our customers include... .../UI Engineer to develop a data-rich and reliable user experience. These frontends are... ...create consistency across our products and sites, we have a central team that builds an open...- ...What we're looking for: We're hiring a Backend Engineer to make Enode's EV platform more reliable, observable, and efficient. As a Mid Level Backend... ...receiving a co-working pass on demand. Three annual off-sites to connect with the team in exciting & fun places....
- ...including Amazon EKS and Google Kubernetes Engine. Provision and manage secure, scalable... ...issues, and implement proactive reliability and efficiency improvements. Apply SRE... ...Requirements ~7+ years of experience in DevOps, Site Reliability Engineering, or cloud...
- ...solutions that help ensure the quality, reliability, and scalability of our platforms. You... ...ll work closely with developers, DevOps engineers, and fellow QA engineers to embed... ...Company Amazon book account! ~ Kodify off-sites, on-sites, events, and team activities!...
- ...with external threat intelligence feeds. Instrument systems for reliability and observability through logging, monitoring, alerting, and SLAs. Partner with threat research and detection engineering teams to translate analytical needs into scalable systems....
- ...organization’s AI/ML roadmap. Requirements ~5+ years of cloud engineering and/or solutions architecture experience with significant AWS... ...~ Primarily remote work within the U.S.; some travel or on-site work may be required for certain positions. ~ Up to 10% travel...
- ...isolation. Build highly available signing and transaction infrastructure with reliable developer-facing APIs. Set technical direction for wallet security, threat modeling, policy engines, secrets management, account recovery, authentication, and programmable...
- ...As a Backend Engineer you have the opportunity to work on our entire software stack and expand your knowledge about state-of-the art software... ...our solution worldwide; Experience in building scalable and reliable systems; Experience in big data infrastructures is a plus;...
2000 $ om året
...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's leading... ...tests that validate software behaviour Build and maintain reliable, fault-tolerant applications and services Collaborate proactively...- ...Frontend Engineer – Expert Workflows for Wind Turbine Diagnostics Do you want to build... ...for wind turbine monitoring, using an AI engine that already analyzes sensor data from... ...concepts into software that feels clear and reliable Contribute to the evolution of our...Fjernjob
- ...for: We're hiring a Senior Software Engineer to help build the software layer that makes... ...fast. Strong skills in designing for reliability, scalability, and performance.... ...working pass on demand. Three annual off-sites to connect with the team in exciting & fun...
- ...Responsibilities Design, build, and maintain reliable backend services and distributed systems. Develop core cybersecurity and... ...participate in architecture discussions. Review code, help establish engineering practices, collaborate with security researchers and product...
- ...employees. The Role We're looking for a Senior Fullstack Engineer who thrives on owning products end-to-end. You'll work across the... ...over bureaucracy. We care about writing clean code, building reliable systems, and shipping work we're proud of. How We Work Everyone...
- ...security, infrastructure security, container security, or security engineering. Deep practical experience with AWS, Kubernetes, containers,... ...80% for dependents, wherever applicable. Annual company off-site held in a new city for one week. Flexible asynchronous work...
2000 $ om året
...here if you are an exceptional software engineer who wants to work on both stable and cutting... ...virtualization enablement, security, reliability and performance. There are a number of... ...patchsets Virtualisation or abstraction engines Container technology Security with...2000 $ om året
...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's leading... ..., high quality software You understand the importance of reliable operations in an agile world You have sound knowledge of cloud...- ...Zencore is a fast-growing company founded by former Google Cloud leaders, architects, and engineers. We are seeking candidates with significant experience in Google Cloud to join our team. Our engagements aim to eliminate obstacles, reduce risk, and accelerate timelines...
- ...Identify repeatable patterns and share insights with Product and Engineering teams. Requirements Strong programming skills in Python... ...non-technical partners. Benefits Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles...
- ...Responsibilities Design and implement robust, highly reliable software for safety-critical systems using functional safety and deep... ...and safety in edge-case scenarios. Mentor junior engineers on safety-critical coding standards and best practices. Requirements...
- ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's... ...believe that automated tests are the key to higher velocity and reliability, you'll fit right in. What you’ll do Collaborate remotely...
Vil du modtage flere ledige stillinger?
Abonnér og modtag tilsvarende ledige stillinger for Staff Site Reliability Engineer. Bliv den første ansøger!
