Lenovo
Cluster Infrastructure Engineer
About the role
Key Responsibilities:
Design and validate AI cluster infrastructure management solutions.
Design and develop methodologies, processes, software frameworks and tools for AI cluster deployment and end-to-end performance evaluation and tuning.
Design and develop hardware and software product features that improve AI cluster solutions in performance, manageability, serviceability, and so forth.
Collaborate with cross-functional teams and contribute to business projects and customer engagements.
Basic Qualifications BS/MS in Computer Science, Computer Engineering, Electrical Engineering, or other related fields.
2+ years of experience in design and implementation of datacenter hardware and software systems, with expertise in at least one of the following areas: server, network, storage, performance benchmarking.
Strong programming skills (e.g., C/C++, Python, Go, shell scripting).
Understanding of Machine Learning and Deep Learning principles.
Experience with Machine Learning frameworks such as PyTorch, LangChain, LangGraph, Autogen, etc.
Excellent communication and leadership skills.
Ability to learn new technologies and concepts quickly.
Innovative mindset with the ability to think outside the box.
Preferred Qualifications: PhD in Computer Science or Computer Engineering.
Experience with AI infrastructure solutions in production environments.
Experience with AI hardware and software stack (e.g., CPU, GPU, network switch, ML platforms, libraries, runtimes).
Ability to innovate and publish at top conferences.
We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.
Before you apply
