NextPath Jobs

Senior Site Reliability Engineer, Vulcan (AI Security product)

AIFT ยท UAE

Posted 2026-09-08 ยท Verified live 2026-09-21

See how your real experience scores against this role โ€” our AI drafts an honest, verifiable resume tailor. Nothing invented, ever. You approve everything.

Get matched โ€” join the waitlist Apply on company site โ†—

About this role

<h3><strong>Job Overview</strong></h3>

<p>We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of&nbsp;on-premise&nbsp;Kubernetes environments for enterprise and government clients,&nbsp;including&nbsp;airgapped, high-security data&nbsp;center&nbsp;environments where remote access is not possible.&nbsp;&nbsp;</p>

<p>This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure&nbsp;and communicating clearly with client stakeholders throughout.&nbsp;</p>

<p>This role carries real ownership; you will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.</p>

<p>In this role, you will play a vital part in supporting our Cybersecurity business, Vulcan.&nbsp;Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security.&nbsp;</p>

<p>Learn more about us ๐Ÿ‘‰</p>

<ul>

<li>Vulcan product:&nbsp;<a href="https://vulcanlab.ai/" rel="noopener noreferrer">https://vulcanlab.ai/</a></li>

<li>Vulcan LinkedIn:&nbsp;<a href="https://www.linkedin.com/company/vulcanlab-ai/" rel="noopener noreferrer">https://www.linkedin.com/company/vulcanlab-ai/</a></li>

<li>AIFT group:&nbsp;<a href="https://aift.io/" rel="noopener noreferrer">https://aift.io/</a></li>

</ul>

<p>&nbsp;</p>

<h3><strong>Responsibilities</strong></h3>

<ul>

<li>Plan and execute on-prem Kubernetes cluster deployments,&nbsp;upgrades&nbsp;and infrastructure migrations (including IP re-addressing, certificate&nbsp;rotation&nbsp;and cluster reconfiguration) in production and&nbsp;airgapped&nbsp;environments&nbsp;</li>

<li>Diagnose and resolve failures independently on-site&nbsp;</li>

<li>Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g.&nbsp;SeaweedFS/Ceph/similar), private container&nbsp;registries&nbsp;and centralized logging (ELK or equivalent)</li>

<li>Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution&nbsp;</li>

<li>Represent the technical work directly to client stakeholders on-site: explain status,&nbsp;failures&nbsp;and remediation plans clearly&nbsp;</li>

<li>Travel to client data&nbsp;centers&nbsp;(including&nbsp;airgapped/restricted-access sites) as&nbsp;required, sometimes on short notice, for deployment and go-live support</li>

<li>Write clear, structured runbooks, decision&nbsp;trees&nbsp;and incident reports that others (including less experienced engineers) can follow under pressure&nbsp;</li>

<li>Escalate risks proactively to internal leadership, not just after something has gone wrong&nbsp;</li>

</ul>

<h3><strong>Requirements</strong></h3>

<p><strong>Technical:</strong>&nbsp;</p>

<ul>

<li>5-6 years of hands-on experience with Kubernetes in production, including at least one&nbsp;on-premise&nbsp;(not purely cloud-managed) deployment&nbsp;</li>

<li>Solid understanding of&nbsp;etcd&nbsp;internals. Quorum, peer membership, failure recovery,&nbsp;not just&nbsp;kubectl-level familiarity&nbsp;</li>

<li>Experience with&nbsp;kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal)&nbsp;</li>

<li>Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls)&nbsp;and container registries (Docker Distribution or similar)&nbsp;</li>

<li>Comfortable working entirely from the Linux command line, writing and debugging bash scripts&nbsp;and reading unfamiliar automation tooling under time pressure&nbsp;</li>

<li>Experience with at least one distributed storage system (SeaweedFS, Ceph,&nbsp;MinIO&nbsp;or similar) is a strong plus&nbsp;</li>

<li>GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not&nbsp;required</li></ul>

One of thousands of fresh listings refreshed nightly, built for cleared & defense careers.

Browse all jobs