Full-time

ML Systems Performance Engineer

Posted by Cerebras • toronto, on, Canada

📍 toronto, on 🕒 June 01, 2026

Apply for this Job Similar Jobs

About the Role

About The Role Engineers on the inference performance team operate at the intersection of hardware and software, driving end-to-end model inference speed and throughput. Their work spans low-level kernel performance debugging and optimization, system-level performance analysis, performance modeling and estimation, and the development of tooling for performance projection and diagnostics. 
Responsibilities Build performance models (kernel-level, end-to-end) to estimate the performance of state of the art and customer ML models. 
Optimize and debug our kernel micro code and compiler algorithms to elevate ML model inference speed, throughput and compute utilization on the Cerebras WSE. 
Debug and understand runtime performance on the system and cluster. 
Develop tools and infrastructure to help visualize performance data collected from the Wafer Scale Engine and our compute cluster. 
Requirements  ...
                

Job Details

Location toronto, on
Job Type Full-time
Category Other-General
Posted June 01, 2026
Deadline July 11, 2026

Ready to Apply?

Submit your application today and take the next step in your career journey with Cerebras.

Apply Now