Remote Otter LogoRemoteOtter

Senior Hardware Engineer, GPU - Remote

Posted 2 weeks ago

Overview

CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024.

As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you’re someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry.

CoreWeave powers the creation and delivery of the intelligence that drives innovation.

In Short

  • Troubleshoot complex GPU and PCIe related failures
  • Partner with external vendors on failure analysis
  • Track component RMAs
  • Develop and maintain hardware/firmware management services
  • Automate all aspects of the server hardware lifecycle
  • Serve as the senior point of contact for hardware escalation and troubleshooting
  • Collaborate with cross-functional teams to define hardware requirements, specifications, and system architecture
  • Create and maintain accurate documentation of hardware designs, specifications, test procedures, and results
  • Analyze and optimize the performance of hardware systems, identify bottlenecks, and propose improvements for enhanced efficiency
  • Establish processes for internal hardware testing, deployment, and performance optimization

Requirements

  • Prior experience supporting and troubleshooting data center class GPUs (preferably A100 or newer)
  • Proficiency in ansible/python and experience with programmatically interacting with server BMCs, using IPMI or Redfish (preferably Redfish)
  • Experience using, integrating and automating data center class GPU diagnostics and troubleshooting tools
  • In-depth knowledge of server hardware, components, and management technologies, particularly GPUs and PCIe devices
  • Proven ability to stay updated with the latest industry technologies and trends
  • Previous experience collaborating with hardware vendors
  • Strong passion for automation, with a commitment to automating processes comprehensively
  • Excellent documentation skills and attention to detail
  • Strong analytical and problem-solving abilities

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

Similar Jobs:

Oowlish logo

Senior Hardware Engineer - Remote

Oowlish

4 weeks ago

Oowlish is seeking a Senior Hardware Engineer to develop embedded hardware solutions for wearables and IoT devices in a remote work environment.

Embedded Hardware
PCB Design
RF Performance Optimization
Power Efficiency
Worldwide
Full-time
Software Development
Oowlish logo

Senior Hardware Engineer - Remote

Oowlish

4 weeks ago

Oowlish is seeking a Senior Hardware Engineer to lead the development of embedded hardware solutions for wearables and IoT devices in a remote work environment.

Embedded Hardware
PCB Design
RF Performance Optimization
Power Efficiency
Brazil
Full-time
Software Development
Oowlish logo

Senior Hardware Engineer - Remote

Oowlish

4 weeks ago

Join Oowlish as a Senior Hardware Engineer to lead the development of embedded hardware solutions for wearables and IoT devices in a remote work environment.

Embedded Hardware
Wearables
IOT Devices
PCB Design
Worldwide
Full-time
Software Development
Oowlish logo

Senior Hardware Engineer - Remote

Oowlish

4 weeks ago

Oowlish is seeking a Senior Hardware Engineer to lead the development of embedded hardware solutions for wearables and IoT devices in a remote work environment.

Embedded Hardware
PCB Design
RF Performance Optimization
Power Efficiency
Worldwide
Full-time
Software Development
Oowlish logo

Senior Hardware Engineer - Remote

Oowlish

4 weeks ago

Oowlish is seeking a Senior Hardware Engineer to lead the development of embedded hardware solutions for wearables and IoT devices in a remote work environment.

Embedded Hardware
PCB Design
RF Performance Optimization
Power Efficiency
Brazil
Full-time
Software Development