← Back to jobs
A

ML Software Engineer, Data Plane

·Israel
On-siteFull-timeBackend DevelopmentIT Services and IT Consulting

Join a team building the inference data plane for large language models on custom hardware. You will develop high-performance compute kernels, integrate serving frameworks, and drive model execution from validation to production. This role requires strong C/C++ skills, Linux systems knowledge, and experience with ML accelerators.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end, from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks such as vLLM and PyTorch, including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads, identifying bottlenecks and driving latency and throughput improvements from simulation through hardware bringup.
  • Own features end-to-end, from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.

Requirements

  • Bachelor's degree or equivalent.
  • 4+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations.
  • Knowledge of computer architecture, operating systems, and parallel computing.
  • Strong proficiency in C/C++.
  • Strong Linux systems knowledge.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features end-to-end.

Nice to have

  • Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.
  • Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques.
  • Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU, or other AI acceleration hardware.
  • Familiarity with speculative decoding, KV cache optimization, or other LLM serving optimizations.
  • Experience with distributed systems, collective communication, RDMA, or high-speed interconnect programming.
  • Demonstrated early adopter of AI-assisted development tools, using LLMs or code-generation agents as part of daily workflow.

Relevance

More opportunities

Similar jobs

Finding the best alternatives for you…

Questions, answered

Frequently asked questions