Observability engineer

  • Stockholm
  • Full-time
  • POSTED 29 DAYS AGO

About the job

At FDJ UNITED, we don't just follow the game, we reinvent it.

FDJ UNITED is one of Europe’s leading betting and gaming operators, with a vast portfolio of iconic brands and a reputation for technological excellence. With more than 5,000 employees and a presence in around fifteen regulated markets, the Group offers a diversified, responsible range of games, both under exclusive rights and open to competition. We set new standards, proving that entertainment and safety can go hand in hand. Here, you’ll work alongside a team of passionate individuals dedicated to delivering the best and safest entertaining experiences for our customers every day.

We’re looking for bold people who are eager to succeed and ready to level-up the game. If you thrive on innovation, embrace challenges, and want to make a real impact at all levels, FDJ UNITED is your playing field.

Join us in shaping the future of gaming. Are you ready to LEVEL-UP THE GAME?

We are looking for a talented and driven Site Reliability Engineer (Observability Engineer) to join our team and play a key role in building the next generation of our high-performance observability framework.

In this role, you will be at the heart of designing, developing, and operating the systems that give our engineering organisation deep visibility into the health, performance, and reliability of our products and infrastructure. You will work closely with engineering teams across the business to ensure our observability stack is scalable, reliable, and fit for purpose in a fast-moving, high-stakes environment.

The current focus of the framework is built around VictoriaLogs and VictoriaMetrics, with a strong emphasis on OpenTelemetry for metrics and log ingestion. You will help shape how this framework evolves and be a hands-on contributor to its ongoing development, as well as help maintain this in a high traffic production site.

What You'll Be Doing Designing and developing a next-generation, high-performance observability platform using modern tooling and best practices.

Driving the adoption and implementation of OpenTelemetry (OTEL/OTLP) for metrics and log ingestion across the organisation.

Building and maintaining Infrastructure as Code (IaC) and configuration management pipelines to automate platform operations.

Contributing to CI/CD pipelines and deployment workflows to support continuous delivery of observability components.

Owning and improving the reliability, scalability, and performance of the observability stack.

Collaborating with engineering teams to define and track Service Level Indicators (SLIs) and establish reliability standards.

Participating in code reviews, architectural discussions, and technical planning sessions.

Supporting capacity planning, release processes, and change management practices.

Championing an automation-first mindset and contributing to a culture of operational excellence.

What We're Looking For Infrastructure & Configuration Management Proven experience with Infrastructure as Code, specifically Terraform

Configuration management experience using Ansible and/or Chef

Coding Skills Strong scripting and development capabilities in Bash, Python, and Go

Experience with Rust and/or Java is a nice to have

CI/CD & Version Control Solid understanding and hands-on experience with Git and GitLab

Experience with ArgoCD and/or Jenkins for CI/CD pipeline management

Monitoring & Observability Metrics platforms: Prometheus, Mimir, InfluxDB, VictoriaMetrics

Log aggregation: Splunk, Loki, Elasticsearch, VictoriaLogs

Visualisation: Grafana

Log forwarding: FluentBit, Vector.dev

OpenTelemetry (OTEL/OTLP): Metrics and log pipeline experience

Containers & Orchestration Strong experience with Docker and/or Podman

Kubernetes (including workload management and cluster operations)

Helm for Kubernetes package management

Cloud Hands-on experience with AWS

Core Concepts & Principles You'll need a solid understanding and proven experience of the following:

Systems Architecture: ability to design and reason about complex distributed systems

Networking principles: TCP/IP, DNS, load balancing, proxies, and beyond

TSDB (Time Series Database) principles: understanding of storage, retention, cardinality, and query patterns

Logging principles: structured logging, log pipelines, and log management at scale

Messaging queues: pub/sub patterns, event-driven architectures

Clustering principles: high availability, fault tolerance, distributed coordination

Load balancing: strategies, implementations, and traffic management

API / JSON principles: RESTful APIs, data serialisation, and integration patterns

Caching mechanisms: In-memory caching, external caches, and cache invalidation strategies

Code reviews: Contributing to and leading thoughtful, constructive reviews

Testing and documentation: Writing quality tests and clear technical documentation

Agile mindset: Comfortable working in iterative, collaborative delivery environments

Automation mindset: Defaulting to automation and repeatability in everything you build

Nice to Have Experience with AI concepts and tooling, including the creation of MCP's and agent Skills

Microservices architecture: Design patterns, service meshes, inter-service communication

Operational Experience Change Management: Structured approaches to managing infrastructure and service changes

Service Level Indicators (SLIs): Defining, measuring, and acting on reliability metrics

Capacity Planning: Forecasting infrastructure needs based on usage trends and growth

Release Processes: Experience with progressive delivery, feature flags, and rollback strategies

We believe talent knows no boundaries. Our hiring process focuses solely on your skills, experience, and potential to contribute to our team. We welcome applicants from all backgrounds and evaluate each candidate based on merit, regardless of personal characteristics as the age, gender, origin, religion, sexual orientation, neurodiversity or disability.