Senior Site Reliability Engineer (Compute Node Team)

Nebius is a cutting-edge AI cloud platform that offers scalable infrastructure for developing and deploying AI solutions.

Nebius

Employee count: 201-500

Netherlands only

Stay safe on Himalayas

Never send money to companies. Jobs on Himalayas will never require payment from applicants.

Why work at Nebius Nebius is leading a new era in cloud computing to serve the global AI economy. We create the tools and resources our customers need to solve real-world challenges and transform industries, without massive infrastructure costs or the need to build large in-house AI/ML teams. Our employees work at the cutting edge of AI cloud infrastructure alongside some of the most experienced and innovative leaders and engineers in the field.

Where we workHeadquartered in Amsterdam and listed on Nasdaq, Nebius has a global footprint with R&D hubs across Europe, North America, and Israel. The team of over 800 employees includes more than 400 highly skilled engineers with deep expertise across hardware and software engineering, as well as an in-house AI R&D team.

The Role

We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud. The Compute Node team is responsible for creating and operating the cluster scheduling layers and compute nodes that schedule and manage virtual machines in the cloud across all regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform.

Your responsibilities will include:

Ensure reliability, availability and performance of compute nodes running VMs
Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer
Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling
Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies
Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs
Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements
Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability

We expect you to have:

Strong Linux expertise:
deep understanding of Linux user space and kernel space
knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces)
clear understanding of system boundaries and constraints at different layers
Virtualization experience:
hands-on experience with QEMU/KVM
understanding of VM lifecycle, performance characteristics and failure modes
Containerization knowledge:
practical experience with containers, namespaces and cgroups
strong understanding of resource isolation and control
Strong debugging skills:
ability to reason about complex system failures
structured, hypothesis-driven approach to incident analysis
SRE mindset:
clear understanding of the SRE role in system design and operations
experience building and operating observability stacks, not just consuming them
ability to turn system behavior into actionable reliability signals

Nice to Have / Optional:

Experience with Kubernetes internals or node-level components
Hands-on experience with low-level Linux debugging tools (e.g. perf, eBPF, ftrace, strace, kernel crash dumps)
Familiarity with large-scale compute or bare-metal platforms
Contributions to open-source infrastructure or system software
Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand

What we offer

Competitive salary and comprehensive benefits package.
Opportunities for professional growth within Nebius.
Flexible working arrangements.
A dynamic and collaborative work environment that values initiative and innovation.

We’re growing and expanding our products every day. If you’re up to the challenge and are excited about AI and ML as much as we are, join us!

Apply now

Please let Nebius know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Apr 28, 2026

Posted on

Jan 28, 2026

Hiring timezones

Netherlands +/- 0 hours

Browse similar jobs

Remote Senior Site-Reliability-Engineer Jobs Remote Full Time Site-Reliability-Engineer Jobs Remote Senior Site-Reliability-Engineer Jobs in Netherlands Remote Full Time Jobs in Netherlands Remote Site-Reliability-Engineer Jobs in Netherlands

About Nebius

Learn more about Nebius and their company culture.

View company profile

At Nebius, we offer an advanced AI cloud platform designed for those who wish to develop, tune, and deploy their AI models with the most efficient infrastructure available. Our platform utilizes cutting-edge NVIDIA GPU clusters, including the H100 and H200, optimized for maximum performance with InfiniBand. One of the standout features of Nebius is our comprehensive fine-tuning ecosystem that includes on-demand GPUs and tools necessary for robust dataset processing, ensuring that AI teams can efficiently manage their computational resources according to demand.

We recognize the importance of AI inference in deploying real-world applications. Hence, we provide a resilient and cost-effective infrastructure that has been optimized for rapid deployment of Generative AI applications. Our services span the entire lifecycle of AI solutions, from model training to inference, making Nebius not just a GPU cloud but a full-stack AI platform. Additionally, we pride ourselves on supporting our clients with 24/7 expert guidance, offering resources to help architects and engineers harness our AI-optimized data centers to build scalable solutions.

Apply now

Please let Nebius know you found this job on Himalayas. This helps us grow!

Apply now

About the job

Apply before

Apr 28, 2026

Posted on

Jan 28, 2026

Job type

Full Time

Experience level

Senior

Location requirements

Netherlands

Hiring timezones

Netherlands +/- 0 hours

Browse similar jobs

Claim this profile

Nebius

Company size

201-500 employees

Markets

AI Cloud Computing Artificial Intelligence Machine Learning Infrastructure GPU Cloud Services Generative AI High Performance Computing Cloud Infrastructure Deep Learning Platforms AI Model Training Data Center Services

Employees live in

Netherlands

View company profile

Similar remote jobs

Here are other jobs you might want to apply for.

View all remote jobs

AL, AD + 48 more

Site Reliability Engineer

Strike

Full Time

Engineering

AF, AL + 146 more

Senior Site Reliability Engineer

Customer.io

Employee count: 51-200

Salary: 140k-180k USD

Full Time

Site Reliability Engineering

AU, CA + 5 more

Senior Web Engineer

Canonical

Employee count: 501-1000

Full Time

Web engineering

AU, BR + 17 more

Site Reliability Engineer

Canonical

Employee count: 501-1000

Full Time

Site Reliability Engineering

AU, AT + 29 more

Site Reliability Engineering Manager

Canonical

Employee count: 501-1000

Full Time

Site Reliability Engineering

AU, CA + 4 more

Senior Site Reliability / Gitops Engineer

Canonical

Employee count: 501-1000

Full Time

Information Systems

158 remote jobs at Nebius

Explore the variety of open remote roles at Nebius, offering flexible work options across multiple disciplines and skill levels.

View all jobs at Nebius

AU, CA + 4 more

Senior Technical Project Manager (External Solutions)

Nebius

Employee count: 201-500

Full Time

& Technology

AU, CA + 6 more

Senior Technical Project Manager (Base Infra)

Nebius

Employee count: 201-500

Full Time

Technical Project Management

United States only

Enterprise Applications Engineer

Nebius

Employee count: 201-500

Full Time

Enterprise Applications Engineering

CZ, DE + 4 more

System Engineer (Token Factory)

Nebius

Employee count: 201-500

Full Time

System Engineering

United States only

Head of Builder Growth

Nebius

Employee count: 201-500

Salary: 225k-315k USD

Full Time

Growth Marketing

Netherlands only

Enterprise Applications Engineer

Nebius

Employee count: 201-500

Full Time

Enterprise Applications Engineer

Top remote companies

Remote companies like Nebius

Find your next opportunity by exploring profiles of companies that are similar to Nebius. Compare culture, benefits, and job openings on Himalayas.

View all companies

Lambda

Benefits Tech stack

Lambda Labs is an AI infrastructure company providing GPU cloud services, servers, and workstations designed to accelerate deep learning and machine learning processes.

Developer Tools Edge AI

CO1 job

CoreWeave

Salaries Benefits Tech stack

AE1 job

Aethir

Tech stack

Aethir is a decentralized cloud infrastructure (DCI) provider focused on delivering enterprise-grade GPU-as-a-Service for AI and cloud gaming applications.

Decentralized Cloud Infrastructure GPU Computing

TE1 job

TensorWave

TensorWave is a pioneering AI-focused cloud platform that leverages AMD's MI300X accelerators, enabling organizations to optimize AI workloads with enhanced performance and lower costs.

Cloud Computing Artificial Intelligence

DA2 jobs

DataCrunch

DataCrunch is a cloud service provider specializing in high-performance GPU servers and clusters for machine learning, powered by renewable energy.

Cloud Computing Artificial Intelligence

IN8 jobs

InfraCloud

Benefits Tech stack

InfraCloud Technologies provides cutting-edge cloud-native solutions, specializing in AI cloud infrastructure and GPU enablement.

Cloud Computing

Top remote companies

Remote companies like Nebius

Find your next opportunity by exploring profiles of companies that are similar to Nebius. Compare culture, benefits, and job openings on Himalayas.

View all companies

Find your dream job

Sign up now and join over 100,000 remote workers who receive personalized job alerts, curated job matches, and more for free!

Find your dream job

Sign up now and join over 100,000 remote workers who receive personalized job alerts, curated job matches, and more for free!

Senior Site Reliability Engineer (Compute Node Team)

The Role

We expect you to have:

Nice to Have / Optional:

What we offer

Apply now

About the job

Apply before

Posted on

Job type

Experience level

Location requirements

Hiring timezones

Job categories

Skills

Browse similar jobs

About Nebius

Apply now

About the job

Apply before

Posted on

Job type

Experience level

Location requirements

Hiring timezones

Job categories

Skills

Browse similar jobs

Nebius

Company size

Markets

Employees live in

Similar remote jobs

Site Reliability Engineer

Senior Site Reliability Engineer

Senior Web Engineer

Site Reliability Engineer

Site Reliability Engineering Manager

Senior Site Reliability / Gitops Engineer

158 remote jobs at Nebius

Senior Technical Project Manager (External Solutions)

Senior Technical Project Manager (Base Infra)

Enterprise Applications Engineer

System Engineer (Token Factory)

Head of Builder Growth

Enterprise Applications Engineer

Remote companies like Nebius

Remote companies like Nebius

Find your dream job

Find your dream job

Apply now

Apply now

Site Reliability Engineer

Senior Site Reliability Engineer

Senior Web Engineer

Site Reliability Engineer

Site Reliability Engineering Manager

Senior Site Reliability / Gitops Engineer

Senior Technical Project Manager (External Solutions)

Senior Technical Project Manager (Base Infra)

Enterprise Applications Engineer

System Engineer (Token Factory)

Head of Builder Growth

Enterprise Applications Engineer

Find your dream job

Remote companies like Nebius