Job Description
We are looking for a Senior Site Reliability Engineer with Cloud platform experience. This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring & alerting, and ensuring overall platform reliability. Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.
Responsibilities
- Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation.
- First response for incidents, contribute to problem management and root cause analysis.
- Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle.
- Develo...