Skip to main content
Updated daily · Sep 13, 2026
All open roles

Open role

Sr. Inference Optimization Engineer (local / edge runtime)

IntelPhoenix, AZFull-time$195,200.00-361,200.00 USD

This engineer will optimize local inference performance across CPUs, integrated GPUs, and Vulkan-based systems, balancing latency, throughput, memory use, and model quality. The work includes improving KV-cache behavior, batching, scheduling, startup and model-loading paths, as well as reducing CPU overhead in interactive agent applications. The engineer will benchmark multiple hardware tiers, document credible results, and contribute useful improvements to open-source inference projects. The position sits within Intel’s effort to deliver private, efficient AI on PCs and edge devices, with collaboration across inference and post-training functions. It suits an experienced systems or performance engineer who enjoys low-level debugging, accelerator programming, and practical LLM optimization.

Requirements

BS/MS in CS, EE, Math or related STEM field8+ years software development experienceStrong C++ and/or Python skills with systems-level code reading abilityProven experience profiling and optimizing CPU or GPU performance with measurable improvementsDeep understanding of LLM inference mechanics (attention, KV cache, decoding strategies)Linux expertise and low-level debugging capabilities
Apply for this role

Listed July 24, 2026 · Verify details with the employer before applying.

About Intel

Four decades of chip manufacturing in Chandler, with the Ocotillo campus among Intel's largest fab sites and ongoing expansion.

More roles at Intel

In the news