Fleet Software Lead

etched

Location not listedSeniorEngineeringPosted 1h ago

View the original posting ↗

Apply at etched

Opens the original listing in a new tab.

Full description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.


Job Summary

We’re hiring a Fleet Systems Lead to join our Supercomputing organization. This team builds the software that enables Etched to deploy inference clusters at gigawatt-scale. This role presents an opportunity to shape how frontier inference hardware is deployed, managed, and orchestrated for our customers.

We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads. Fleet Software enables deployments of these systems at scale through configuration management, safe software and firmware rollouts, and software recovery, while partnering with other engineering teams to drive product improvements as we ramp production at unprecedented speed.

The team also builds the monitoring and analysis systems that connect signals across the rack. Engineers and their agents use this data to investigate failures, spot patterns and continually improve the performance and reliability of the inference fleet. We’re looking for a leader to drive technical direction while building the best team in the industry. Someone who has built infrastructure and owned it through real deployments in the field. This team’s work will shape how Etched scales to a computing platform that customers depend on at massive scale.

Key Responsibilities

  • Lead and develop the Fleet Software team. Set priorities, hire and mentor engineers, and give team members clear ownership.

  • Set the architecture and roadmap for fleet provisioning, configuration, updates and recovery across labs and customer deployments.

  • Establish the team’s approach to safe software and firmware rollouts, including validation, staged deployment and handling partial failures.

  • Guide the design of telemetry and log pipelines, including decisions around storage, retention and query performance.

  • Lead the development of monitoring, alerting and analysis tools that help engineers and customers understand fleet behavior and investigate failures.

  • Partner with our Node Systems Systems team to define the interfaces and validated configurations needed to manage hardware across the fleet.

  • Work with deployment teams to support customer environments, including local monitoring and customer telemetry analysis.

  • Turn lessons from deployments and incidents into engineering priorities, and decide where to adopt existing tools versus building our own.

  • Stay involved through design reviews, code contributions and hands-on debugging, particularly on the team’s hardest technical problems.

You may be a good fit if you have (Must-have qualifications)

  • 5+ years of software engineering experience, including leading projects or managing engineers.

  • Experience building infrastructure software and owning it in production, including deployments, incidents and the changes that followed.

  • A strong background in fleet management, provisioning, configuration or observability for compute infrastructure or connected devices.

  • Solid distributed systems and Linux fundamentals. You understand how systems behave under load, during upgrades and when components fail.

  • Strong coding and debugging skills. You can work through a difficult problem with an engineer and contribute to the implementation.

  • Experience setting technical direction, mentoring engineers and getting work shipped across team boundaries.


Strong candidates may also have experience with (Nice-to-have qualifications)

  • Telemetry and log pipelines, including storage, retention and query performance.

  • Bare metal servers, accelerator clusters, robotics or other hardware fleets.

  • Firmware updates, hardware diagnostics and automated recovery.

  • Software that runs in customer environments as well as internal infrastructure.


Benefits

  • Medical, dental, and vision packages with generous premium coverage

    • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2,500 per month for those living within walking distance of the office

  • Relocation support for those moving to San Jose (Santana Row)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch and dinner in our office

  • Unlimited compute budget subject to ROI justification

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Other fresh roles

More Engineering jobs →