NextPath Jobs

Lead Engineer, Issue Management & Triage

Diligent Robotics · Austin, Texas, United States

Posted 2026-09-11 · Verified live 2026-09-26

See how your real experience scores against this role — our AI drafts an honest, verifiable resume tailor. Nothing invented, ever. You approve everything.

Get matched — join the waitlist Apply on company site ↗

About this role

<p><strong>What we’re doing isn’t easy, but nothing worth doing ever is.&nbsp;</strong></p>

<p>Diligent builds helpful robots that work safely and autonomously in real world environments. We move quickly, solve messy problems, and care deeply about reliability at scale. As a Fleet Engineer, you'll own the reliability and continuous improvement of our deployed robotic fleet — leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You'll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors.</p>

<p>&nbsp;We are hiring a <strong>Lead Engineer, Issue Management &amp; Triage</strong> to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering.&nbsp;This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet.&nbsp;You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale.&nbsp;</p>

<p>Location: Austin preferred, Remote possible (U.S.)<br>Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days)</p>

<p><strong>What You’ll Do:</strong></p>

<p><em>Own Issue Management &amp; Triage Systems</em></p>

<ul>

<li>Design and own end-to-end systems for issue intake, triage, and escalation.</li>

<li>Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization.</li>

</ul>

<p><em>Build Tools &amp; Automation (Hands-On)</em></p>

<ul>

<li>Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort.</li>

<li>Contribute directly to codebases (Python, backend services) and partner with Engineering on system integrations (logs, telemetry, alerts).</li>

</ul>

<p><em>Bridge Operations &amp; Engineering</em></p>

<ul>

<li>Act as the primary technical interface between the Remote Operations Center (ROC) and Engineering.</li>

<li>Translate real-world issues into prioritized, categorized technical problems for resolution alignment.</li>

</ul>

<p><em>Performance Measurement &amp; Classification Frameworks</em></p>

<ul>

<li>Develop systems and taxonomies to systematically measure and classify robot performance, failure modes, and degradation across the fleet.</li>

<li>Build dashboards and reporting systems to track trends, severity, and impact.</li>

</ul>

<p><em>Root Cause Analysis &amp; Continuous Improvement</em></p>

<ul>

<li>Establish best practices for Root Cause Analysis (RCA) and identify systemic issues.</li>

<li>Drive long-term fixes and create feedback loops to influence improvements in hardware, software, and autonomy.</li>

</ul>

<p><strong>What We’re Looking For:</strong></p>

<ul>

<li>7+ years in relevant technical or program management roles (e.g., engineering, incident management)</li>

<li>3+ years of people management</li>

<li>Experience with complex, real-world systems (robotics, autonomous/distributed systems, or hardware-software products)</li>

<li>Proven track record building operational tools, systems, or infrastructure for workflows</li>

</ul>

<p><strong>Technical Skills</strong></p>

<ul>

<li>Strong programming experience (Python preferred; backend or data systems experience a plus)</li>

<li>Experience with:</li>

<ul>

<li>Data pipelines and telemetry systems</li>

<li>Monitoring, alerting, and logging infrastructure</li>

<li>Internal tools and automation systems</li>

</ul>

<li>Ability to design scalable systems for classification, prioritization, and workflow automation</li>

<li>Familiarity with platforms like Jira, Zendesk, SQL, Looker, Foxglove, or similar</li>

</ul>

<p><strong>Systems &amp; Product Thinking</strong></p>

<ul>

<li>Strong systems thinker, translating ambiguous operational problems into structured technical solutions</li>

<li>Experience defining metrics, taxonomies, and performance frameworks</li>

<li>Data-driven approach to prioritization and decision-making</li>

</ul>

<p><strong>Mindset</strong></p>

<ul>

<li>Hands-on and willing to dive into technical problems when needed</li>

<li>Strong ownership and bias toward action</li>

<li>Comfortable operating in a fast-paced, scaling environment</li>

<li>Passion for improving real-world system performance and reliability</li>

</ul>

One of thousands of fresh listings refreshed nightly, built for cleared & defense careers.

Browse all jobs