Site Reliability Engineer

Budapest
IT
Távmunka
Short Description

We are seeking an experienced Site Reliability Engineer to join our partner's managed services team and support the reliability, stability, and performance of enterprise production environments. In this hands-on operations role, you will be responsible for monitoring production systems, triaging and resolving incidents, performing approved remediation activities, and ensuring service continuity across cloud and application infrastructure. Working closely with SRE teams and technical stakeholders, you will play a key role in maintaining high availability and operational excellence within a 24/7 service environment.

Description

  • Production Monitoring & Alert Management
  • Monitor production infrastructure, cloud platforms, and application environments using industry-standard observability tools.
  • Acknowledge, assess, and triage alerts within established SLAs.
  • Evaluate service impact and determine appropriate operational responses.
  • Analyze logs, metrics, and monitoring data to support troubleshooting and incident resolution.
  • Manage alert filtering and maintenance activities according to approved procedures.
  • Identify recurring issues and recommend monitoring improvements.
  • Incident Management & Resolution
  • Create, update, and maintain incident records in ServiceNow or similar ITSM platforms.
  • Perform first-line troubleshooting and incident diagnosis.
  • Execute approved remediation activities and service recovery procedures.
  • Maintain detailed incident timelines, investigation notes, and resolution records.
  • Escalate complex issues to the SRE on-call team when necessary.
  • Support rapid restoration of production services while adhering to operational processes.
  • Linux Administration & Troubleshooting
  • Perform routine Linux administration and production support activities.
  • Investigate system, application, and infrastructure issues using command-line tools.
  • Analyze system logs, resource utilization, process status, and service health.
  • Execute operational recovery procedures and node-level remediation.
  • Escalate advanced technical issues requiring deeper investigation.
  • Cloud Operations & Automation
  • Support production workloads running on public cloud platforms.
  • Monitor cloud-hosted environments and infrastructure services.
  • Assist with troubleshooting performance, availability, and infrastructure-related issues.
  • Develop and maintain operational scripts using Bash, Python, or similar technologies.
  • Automate repetitive operational tasks and contribute to efficiency improvements.
  • Documentation & Continuous Improvement
  • Execute operational activities following runbooks and standard operating procedures.
  • Maintain accurate technical documentation and incident records.
  • Identify opportunities to improve runbooks, monitoring capabilities, and operational processes.
  • Participate in knowledge sharing, incident reviews, and continuous improvement initiatives.
  • Ensure effective communication during shift handovers and ongoing incident management.

Requirements

  • Technical Skills & Experience
  • Minimum 3 years of experience in DevOps, Site Reliability Engineering, Production Operations, or Infrastructure Support.
  • At least 2 years of experience supporting public cloud environments.
  • Strong Linux administration and troubleshooting skills.
  • Experience with monitoring and observability platforms such as: Splunk, Dynatrace, Prometheus, Grafana, ELK Stack or equivalent tools.
  • Experience in alert management, incident triage, prioritization, and escalation.
  • Hands-on experience with ITSM tools such as ServiceNow.
  • Ability to analyze logs and troubleshoot complex production issues.
  • Scripting and automation experience using Bash, Python, or similar languages.
  • Familiarity with cloud infrastructure, virtual machines, containers, and enterprise production environments.
  • Experience working with runbooks, operational procedures, and escalation frameworks.
  • Nice to Have: Knowledge of Kubernetes and containerized environments.
  • Understanding of networking, load balancers, and storage fundamentals is an advantage.
  • Experience with automated operational workflows.
  • Exposure to enterprise production support or managed services environments.
  • Familiarity with incident lifecycle management and SLA-driven support models.
  • Experience working with SAP-related production environments.

Offer

  • The opportunity to work with enterprise-scale infrastructure and business-critical systems.
  • An international, collaborative engineering environment.
  • Opportunities to develop your technical expertise across infrastructure, automation, cloud, and reliability engineering.
  • Exposure to modern technologies and structured operational practices.
  • Responsibilities and technical ownership aligned with your experience and seniority.

Érdekli ez a pozíció?
Értesüljön hasonló lehetőségekről is.

Jelentkezzen karrier-adatbázisunkba, és értesítjük, amikor olyan új lehetőség nyílik, amely érdekelheti.

Jelentkezem
Site Reliability Engineer
Jelentkezés
Engedélyezett fájlkiterjesztések: doc, docx, pdf, txt. Maximális fájlméret: 50 MB.
Hajlandó költözni?
CAPTCHA
Írja be a képen látható karaktereket.
This question is for testing whether or not you are a human visitor and to prevent automated spam submissions.
loading-gif