Partnering with expansion-stage, enterprise software companies that we believe can become category leaders.

Limited partner investing in exceptional early-stage venture fund managers.

Partnering with early-stage companies at the nexus of technology and culture.

Portfolio Jobs

Looking for your next role? Take a look at these exciting jobs at Sapphire Ventures’ portfolio companies. Our Talent team is passionate about connecting you to your dream job!

My job alerts

Site Reliability Engineer

Catchpoint

This job is no longer accepting applications

See open jobs at Catchpoint.See open jobs similar to "Site Reliability Engineer" Sapphire Ventures.

Software Engineering

Indiana, USA · Remote

Posted on Oct 31, 2024

Who monitors the monitoring system? A Site Reliability Engineer at Catchpoint is responsible for supporting the systems that run Catchpoint’s global monitoring platform. In this role, you will interact directly with operations and development teams on building and maintaining automation and monitoring to ensure Catchpoint has a scalable and highly reliable system for our customers.

The role requires an operational mindset and a love of solving problems on a global scale with solutions that maintain high reliability and availability. You’ll be exploring and making sense of systems telemetry, logs, passive monitoring and our own synthetic monitors to create an automation that controls, rolls out, and maintains our platform.

This position reports to an SRE manager.

Responsibilities:

Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement
Maintain services once they are live by measuring and monitoring availability, latency and overall system health. Establish performance baselines, define actions and automation correlating data from multiple sources
Design, build, and maintain logging and telemetry systems that are used to manage all services.
Design, code, test, and deliver software to automate manual operational work.
Troubleshoot priority incidents, facilitate blameless post-mortems and ensure permanent closure of incidents.
Identify application patterns and analytics in support of better service level objectives.
Deploy and maintain systems that run on multiple cloud providers (AWS, GCP, Azure, Alibaba, Tencent, Oracle, IBM) and physical systems around the world.
Be part of an on-call rotation to support production systems

Desired Skills & Experience:

Strong Experience/knowledge of administering application servers, web servers, and databases
Familiarity with Automation and configuration management tools (preferably terraform)
Good networking knowledge and experience with Internet Architecture (BGP, peering, DNS).
2+ years of incident resolution experience in a large-scale operations environment.
Hands-on experience with cloud deployment, monitoring, and ops analysis tools such as Prometheus, Elasticsearch, Grafana, Kibana, Splunk, Terraform, Jenkins, etc.
3+ years with python, bash, PowerShell, C, etc.
Virtualization experience required.
BS degree in Computer Science or related technical field involving coding or equivalent practical experience.
Appreciation of the value of diversity of opinions

Overview

Catchpoint is the Internet Resilience Company™. The top online retailers, Global2000, CDNs, cloud service providers, and xSPs in the world rely on Catchpoint to increase their resilience by catching any issues in the Internet stack before they impact their business. The Catchpoint platform offers synthetics, RUM, performance optimization, high fidelity data and flexible visualizations with advanced analytics. It leverages thousands of global vantage points (including inside wireless networks, BGP, backbone, last mile, endpoint, enterprise, ISPs and more) to provide unparalleled observability into anything that impacts your customers, workforce, networks, website performance, applications and APIs.

Catchpoint is an equal opportunity employer that strongly prohibits Discrimination and Harassment of any kind. We celebrate diversity and are committed to creating an inclusive and engaging environment for all employees. We welcome applications from all candidates and look forward to receiving yours!

#LI-REMOTE

This job is no longer accepting applications

See open jobs at Catchpoint.See open jobs similar to "Site Reliability Engineer" Sapphire Ventures.

See more open positions at Catchpoint

Privacy policy Cookie policy

Portfolio Jobs

Site Reliability Engineer

Austin, TX

San Francisco, CA

Menlo Park, CA

London