Applications & FAE
$132,000 - $253,000 USD yearly
Originally posted 18 September 2026 by the employer — open 1 day.
About the role
This role involves evaluating the feasibility of new builds and managing new qualifications and certifications for NVIDIA data center products.
What you'll do
- Serve as the technical lead for sustaining engineering related to NVIDIA data center products used by OEM customers.
- Manage complex issues from initial triage to root-cause analysis, corrective action, and resolution.
- Perform system- and board-level failure analysis to isolate hardware, firmware, software, thermal, power, and integration-related problems.
- Facilitate the complete RMA process, including failure-data collection, return authorization, material tracking, engineering disposition, and communication of findings.
- Reproduce field failures in laboratory environments, analyze failure trends, and drive preventive improvements.
- Provide on-site technical support and deliver failure-analysis reports, corrective-action updates, troubleshooting procedures, and executive-level communications.
What you'll need
- BS or MS in Electrical Engineering, Computer Engineering, Computer Science, or a related discipline, or equivalent experience.
- 5+ years of relevant experience in Field Application Engineering, sustaining engineering, failure analysis, systems engineering, or product engineering.
- Hands-on experience debugging complex server or computing platforms at both the system and board levels.
- Demonstrated experience managing field failures or RMA cases through root-cause analysis and corrective-action closure using methods such as 8D, FRACAS, or fault-tree analysis.
- Strong understanding of server architecture, including CPUs, GPUs, memory, storage, power delivery, cooling, and high-speed interconnects.
- Knowledge of PCIe and experience troubleshooting BMC/IPMI or Redfish, BIOS/UEFI, Linux, drivers, firmware, system logs, and hardware telemetry.
- Experience with hardware bring-up, board validation, system qualification, and L1–L10 server integration and manufacturing test stages.
- Ability to interpret schematics, block diagrams, diagnostic data, manufacturing test results, and thermal or electrical measurements.
Skills: sustaining engineering, failure analysis, root-cause analysis, RMA process, server architecture, debugging