Open role
Sr. Inference Optimization Engineer (local / edge runtime)
This engineer will optimize local inference performance across CPUs, integrated GPUs, and Vulkan-based systems, balancing latency, throughput, memory use, and model quality. The work includes improving KV-cache behavior, batching, scheduling, startup and model-loading paths, as well as reducing CPU overhead in interactive agent applications. The engineer will benchmark multiple hardware tiers, document credible results, and contribute useful improvements to open-source inference projects. The position sits within Intel’s effort to deliver private, efficient AI on PCs and edge devices, with collaboration across inference and post-training functions. It suits an experienced systems or performance engineer who enjoys low-level debugging, accelerator programming, and practical LLM optimization.
Requirements
Listed July 24, 2026 · Verify details with the employer before applying.
About Intel
Four decades of chip manufacturing in Chandler, with the Ocotillo campus among Intel's largest fab sites and ongoing expansion.