JobSnoutLooking for top dogs across Europe

All jobs · Spain

Senior Site Reliability Engineer

Lodgify

Apply at LodgifyOpens the employer's own posting.
Where
Remote · Spain
Level
Senior
Salary
not stated by the employer
Technologies
PythonKubernetesDistributed systems
First sniffed
27 Aug 2026 12:00 · open 5 days
Last verified
01 Sept 2026 02:00 · still on the employer's site
Source
Company careers system

⭐ Who we are

Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology.

Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals.
⭐ How will you make an impact?
-
Define meaningful SLIs, SLOs, and reliability targets for the platform.
-
Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability.
-
Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness.
-
Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures.
-
Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana.
-
Implement operational and security best practices through guidelines, policies and automation.
-
Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly.
-
Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks.
-
Implement self-service Internal Developer Platform features via APIs and Kubernetes operators.
-
Improve deployment safety, rollbackability, and release observability.
-
Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms.
-
Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements.
-
Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability.
⭐ What makes you a great fit?
-
You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure.
-
You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery.
-
You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements.
-
You can write maintainable software to automate operational tasks and reduce manual intervention.
-
You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms.
-
You know how to balance reliability, performance, cost, and delivery speed pragmatically.
-
You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support.
-
You collaborate effectively with Engineering, Platform, Security, and Product stakeholders.
-
You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership.
-
You model initiative and accountability, raising risks early and driving improvements through to completion.
⭐ What does success look like?
-
Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage.
-
Reliability targets are consistently met across critical infrastructure and services.
-
Operational toil and manual intervention are measurably reduced through automation and safer workflows.
-
MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths.
-
Post-incident actions are tracked, completed, and used to reduce repeat incidents.
-
Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations.
-
Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience.