AI Inference Engineer

    Fuse Energy

    AIJobsUAE
    Find JobsCompaniesEventsResources
    Home/Jobs/Fuse Energy/AI Inference Engineer
    Fuse Energy

    AI Inference Engineer

    Fuse Energy
    Dubai, Dubai, United Arab Emirates Full-timePosted 5 Aug 2026
    Services for Renewable Energy
    Job Description

    Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.The OpportunityFuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.ResponsibilitiesDefine Fuse's inference serving strategy and architecture from first principlesDesign and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloadsOwn model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plansAct as a direct technical owner of inference performance and reliabilityWork closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layerSet the standards, tooling, and benchmarks this function will run on as it growsRequirements4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experienceDeep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)Strong systems thinking - able to reason about the full path from incoming request to served response across a large clusterComfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving systemA track record of making high-stakes architecture calls and owning the outcomeComfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established oneNice to HaveExperience with Triton or custom ML inference/training frameworksExperience with autoscaling or capacity planning for large-scale inference workloadsExposure to multi-tenant serving or SLA-driven infrastructureBackground at a hyperscaler, frontier AI lab, or large-scale distributed inference systemFamiliarity with Kubernetes/Slurm for cluster orchestrationInterest or experience in energy markets, grid systems, or sustainability-focused computeBenefitsCompetitive salary and an equity sign-on bonusBiannual bonus schemeFully expensed tech to match your needsBreakfast and dinner allowance for office based employees

    Sign in to apply

    Create a free account to apply for AI Inference Engineer and track your applications.

    One-click apply and track your application status
    Save jobs and build your shortlist
    Get alerts for new AI & ML jobs in UAE
    About Fuse Energy

    Industry

    Services for Renewable Energy

    Application Tips

    Tailor your CV

    Highlight your most relevant AI/ML experience

    Research Fuse Energy

    Check their AI products and latest news

    Show impact

    Use metrics to quantify your achievements

    Similar Jobs

    Applied AI Engineer

    Fuse Energy · Dubai, Dubai, United Arab Emirates

    Artificial Intelligence Engineer (UAE National)

    emaratech · Dubai, Dubai, United Arab Emirates

    Data Scientist

    TalentOne · Dubai, Dubai, United Arab Emirates

    AI SME (Subject Matter Expert)

    Max Accelerate Technology Group · Dubai, Dubai, United Arab Emirates

    AIJobsUAE

    Connecting AI talent with opportunities across the United Arab Emirates.

    For Candidates

    • Browse Jobs
    • Find Events
    • Application Tracker

    For Employers

    • Post Jobs
    • Company Profiles

    Support

    • Contact Us
    • Privacy Policy

    The Sunday Brief

    Your week in UAE AI — new roles, hiring trends, and one skill to learn. Free, every Sunday.

    Browse AI Jobs

    AI Jobs in DubaiMachine Learning Jobs in DubaiData Scientist Jobs in UAERemote AI Jobs in UAEAI Jobs in Abu Dhabi

    © 2026 AIJobsUAE. All rights reserved. Empowering AI careers across the Emirates.

    Powered by ArtisanAI