Stateside Spec Engineering jobs in the United States

Etched

Platform Product Reliability

San Jose · On-site · Senior

Discipline
Hardware
Industry
Semiconductor

Spec sheet

Automatically summarized from the employer’s posting. Check the original description below for full requirements.

Platform Product Reliability Engineer for AI hardware infrastructure, owning reliability strategy and fleet-level validation from design to field deployment, partnering across mechanical, thermal, firmware, manufacturing, and operations to ensure scalable reliability.

Day shift · Relocation offered · 5+ years · BS, MS, or PhD in Electrical/Mechanical/Reliability Engineering or related field

Required

  • reliability engineering for hardware systems
  • FMEA
  • Weibull analysis
  • HALT/HASS
  • qualification planning
  • failure analysis methodologies
  • reliability statistics and modeling
  • cross-functional root-cause investigations

Preferred

  • liquid-cooled systems
  • thermal management for high-power AI infrastructure
  • hyperscale/cloud datacenter deployments
  • leading a reliability organization
  • fleet telemetry and analytics

Benefits

  • Medical
  • dental
  • vision
  • 401k
  • PTO

Employer description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking a highly technical and execution-focused Platform Product Reliability Engineer to lead reliability engineering across Etched’s server, rack, and datacenter platform products.

This role owns system-level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes, and long-term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems — ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross-functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations, and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities

Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment

Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations

Build and institutionalize reliability engineering processes spanning the full product lifecycle:

EVT / DVT / PVT qualification gates and exit criteria

Accelerated life testing (ALT) and accelerated stress testing (AST)

Environmental testing: temperature, humidity, altitude, contamination

HALT / HASS programs for design margin and production screening

Vibration, shock, and transportation stress testing

Power cycling, thermal cycling, and long-duration soak testing

Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains

Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking

Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change

Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments

Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale

Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams

Build and lead a high-performing product reliability engineering organization — hiring, developing, and retaining technical talent as the company scales

You may be a good fit if you have (Must-have qualifications)

BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field

5+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work

Experience leading reliability programs for one or more of:

AI accelerator or GPU-class compute systems

Hyperscale or cloud server infrastructure

Networking platforms, storage systems, or rack-scale infrastructure

Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability

Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling

A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high

Strong technical judgment — capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations

Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders

Strong candidates may also have experience with (Nice-to-have qualifications)

Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure

Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer-facing reliability commitments and SLA management

Demonstrated experience building a reliability organization from early-stage — establishing processes, tooling, and team norms in environments without established infrastructure

Familiarity with fleet telemetry systems, large-scale field reliability analytics, and data-driven approaches to proactive reliability management

Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on-site qualification engagement

Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures — with an understanding of how these affect system-level reliability behavior

Benefits

Medical, dental, and vision packages with generous premium coverage

$500 per month credit for waiving medical benefits

Housing subsidy of $2,500 per month for those living within walking distance of the office

Relocation support for those moving to San Jose (Santana Row)

Various wellness benefits covering fitness, mental health, and more

Daily lunch and dinner in our office

Unlimited compute budget subject to ROI justification

How we’re different

Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system. Our addressable market is the entirety of inference, unlike many of our competitors.

 We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Stateside Spec lists roles for engineering teams in the United States. Work-authorization notes come from the employer. We do not guarantee visa policy. Apply off-site. Report an issue with this listing View employer source