Remote Otter LogoRemoteOtter

Site Reliability Engineer - Remote

Posted 15 weeks ago
DevOps / Sysadmin
Full Time
Australia

Overview

Join Avetta as a Site Reliability Engineer in Australia! Site Reliability Engineers are pioneers of the production systems, we believe in proactive discovery and analysis of our entire stack, continually optimizing, tuning, and scaling the system for maximal end-user experience on a globally distributed cloud-based SaaS platform. Downtime is not within the SRE’s vocabulary. The ability to maintain highly resilient and distributed systems, while integrating uptime monitors using programmatic APIs and developing intelligent scaling algorithms are important skills for the SRE. In addition, the SRE needs to be able to communicate effectively with both development and product teams to drive technical discovery and help prioritize features that maintain and exceed uptime goals and end-user experience.

In Short

  • Lead the management and monitoring of highly available replicated cloud systems.
  • Oversee 24/7 Network Operations Center (NOC) operations, maintaining a minimum 99.9% annual uptime.
  • Define golden signals for all services in our core SaaS application.
  • Manage NOC engineer teams, including scheduling and responsibilities.
  • Design PagerDuty escalation policies across various teams.
  • Expertise in AWS technologies and building dashboards with leading observability platforms.
  • Automate monitors and dashboards using modern programmatic methods.
  • Provide regular reports to Engineering leadership and executive teams for continuous improvement.

Requirements

  • Minimum B.S. or B.A. in Computer Science.
  • Minimum of 5 years of experience as a Site Reliability Engineer, including some experience in managing teams and leading projects.
  • Stellar communication and interpersonal skills for effective collaboration with Development & Product teams.
  • Proficiency in monitoring the networking stack using distributed tracing and profiling tools.
  • Proficient with building dashboards with NewRelic, Kibana, Grafana, Prometheus and other observability platforms.
  • Proficient with AWS technologies.
  • Working knowledge in monitoring RESTful microservices and basic HTTP protocols.
  • Able to automate monitors and dashboards using REST APIs, GraphQL, and other modern programmatic methods.
  • Working knowledge of profiling tools for measuring CPU, Memory, I/O, Disk, and process threads dumps.
  • Experience in managing, integrating, and automating alerting and escalation tools.

Benefits

  • Join us at Avetta and be at the forefront of driving technical excellence and ensuring a seamless experience for our users across the globe.
Avetta logo

Avetta

Avetta is a forward-thinking company dedicated to transforming the way businesses connect and collaborate through its flagship SaaS platform, Connect. The company focuses on empowering organizations to streamline their operations and enhance customer experiences. Avetta is committed to innovation and customer-centric solutions, seeking passionate individuals to lead product development and drive industry change. With a dynamic and creative team, Avetta values continuous learning and professional growth, offering comprehensive benefits to its employees.

Share This Job!

Save This Job!

Similar Jobs:

Software Mind logo

Site Reliability Engineer - Remote

Software Mind

6 weeks ago

Software Mind is looking for a Site Reliability Engineer to enhance the reliability of their software systems in a flexible and supportive work environment.

LATAM
Full-time
DevOps / Sysadmin
Jackbox Games logo

Site Reliability Engineer - Remote

Jackbox Games

7 weeks ago

Join Jackbox Games as a Site Reliability Engineer to maintain AWS infrastructure and develop applications in Go.

USA
Full-time
DevOps / Sysadmin
$103,326 - $190,465/year
Pinterest logo

Site Reliability Engineer - Remote

Pinterest

7 weeks ago

Pinterest is seeking a Site Reliability Engineer to ensure the reliability of its large-scale distributed systems.

USA
Full-time
Software Development
Printify logo

Site Reliability Engineer - Remote

Printify

7 weeks ago

Join our team as a Site Reliability Engineer, responsible for ensuring the reliability of our distributed systems and platforms in a dynamic international environment.

Worldwide
Full-time
DevOps / Sysadmin
Zepz logo

Site Reliability Engineer - Remote

Zepz

7 weeks ago

Join Zepz as a Site Reliability Engineer to enhance service stability and resilience through innovative automation and observability practices.

South Africa
Full-time
DevOps / Sysadmin