Senior AI Engineer – LLM Systems
Evollabs Tech
DescriptionWe are a technology company focused on designing and developing advanced, customized server hardware solutions optimized for artificial intelligence workloads. Our mission is to accelerate AI innovation by delivering high-performance, scalable, and energy-efficient infrastructure for datacenter-scale inference.Our chips in development are purpose-built for large-scale AI inference and will be deployed in rack-level systems where multiple devices collaborate to deliver optimal latency, throughput, and efficiency. We are building the next generation of AI infrastructure and are looking for engineers who deeply understand how modern large language models behave at scale.This role is focused on LLM systems, architecture, and performance. You will work at the intersection of model internals and hardware, ensuring that state-of-the-art models run efficiently on our platform.ResponsibilitiesAnalyze, profile and optimize large language model (LLM) inference performance across distributed, multi-chip systems Bring a deep understanding of transformer architectures, including dense and Mixture-of-Experts (MoE) models Evaluate and benchmark different LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek) on custom hardware Design and implement optimizations for attention mechanisms (e.g., Flash Attention, grouped-query attention, sliding window attention) Work on model-level optimizations such as quantization (INT8/FP8), KV-cache management, batching, and parallelism strategies Collaborate with hardware and compiler teams to co-design efficient inference pipelines Build and maintain benchmarking frameworks to evaluate latency, throughput, and scaling behavior Analyze trade-offs between model architecture choices and system-level performance Contribute to model deployment strategies for large-scale datacenter environments Stay up to date with the latest research in LLM architectures and inference optimization RequirementsStrong understanding of transformer architectures and LLM internals Hands-on experience working with multiple modern LLMs (e.g., LLaMA, Mistral, Qwen, DeepSeek, etc.) Deep knowledge of different LLM architectures such as dense models and Mixture-of-Experts (MoE) architectures Familiarity with attention mechanisms and their optimizations Experience with LLM inference optimization techniques (quantization, pruning, KV caching, batching, etc.) Strong Python skills and experience with ML frameworks (PyTorch, JAX, or similar) Experience with distributed systems and large-scale inference workloads Ability to profile and debug performance bottlenecks across hardware and software stacks Strong systems thinking and ability to work across model, runtime, and hardware layers Preferred Qualifications8+ years of relevant experience in deep learning, AI systems, or performance engineering Experience working close to hardware (GPU, TPU, or custom accelerators) Experience with parallelism strategies (tensor parallelism, pipeline parallelism, expert parallelism) Familiarity with datacenter-scale deployment, orchestration and inference servers (e.g., vLLM) Background in performance engineering or systems optimization What We're Not Looking ForThis role is not focused on prompt engineering, or application-layer GenAI development. Instead, it is centered on deep LLM internals, architecture, and inference performance at scale.Why Join Us?Work on cutting-edge AI hardware designed specifically for LLM inference Solve challenging problems at the intersection of AI models and systems Collaborate with a team pushing the boundaries of datacenter-scale AI performance Make a direct impact on the future of AI infrastructure
Create a free account to apply for Senior AI Engineer – LLM Systems and track your applications.
Tailor your CV
Highlight your most relevant AI/ML experience
Research Evollabs Tech
Check their AI products and latest news
Show impact
Use metrics to quantify your achievements
Lead AI Engineer
DataRobot · Dubai, Dubai, United Arab Emirates
AI Engineer
SUPERBOT (BOT) · Dubai, Dubai, United Arab Emirates
Machine Learning Engineer, Applied AI - Deployed
Brain Co. · Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates
Software Development Engineer II
Esri · Sharjah, Sharjah Emirate, United Arab Emirates