AI Engineer – Agent Learning & Evaluation

Confidential Careers

Confidential Careers

AI Engineer – Agent Learning & Evaluation

Confidential Careers
United Arab Emirates Full-timeFirst posted: 6 Oct 2026Last updated: 6 Oct 2026
IT Services and IT Consulting
Job Description

We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs inside customers' own environments, including on-premises and air-gapped sites. We are a small, senior team working in two-week sprints towards a first production release in December 2026.You will own two things that decide whether our agents can be trusted:How agents learn a new application. They learn from human demonstrations and from exploring the application themselves, and they keep that knowledge up to date when the application changes.How we prove they work. You will build the evaluation framework that measures agent reliability and catches regressions before they reach users.What you will doBuild ways for agents to learn an application from human demonstrations and from self-exploration, and turn what they learn into reusable, testable procedures (playbooks).Represent what agents know about an application as structured knowledge, for example a knowledge graph of screens, actions, and data.Build simulated and recorded environments where agents can practise and be tested safely before acting on live systems.Make agents self-correcting: detect when an application changes or behaves unexpectedly, and update their knowledge and playbooks without breaking production tasks.Build the evaluation framework:task suites and success metricsLLM-as-judge with calibration against human labelsregression testing whenever a prompt, model, or tool changesonline monitoring of agent reliabilityImprove agent performance through prompt and program optimisation (e.g. DSPy, TextGrad, Ax), and through fine-tuning or reinforcement learning where it pays off.Work closely with the AI Engineer – Agent Harness, and mentor the L4 AI engineers on evaluation practice.What we are looking forMust have10+ years of software engineering, including 2+ years building LLM-powered applications.Has designed and run rigorous evaluations for LLM or agent systems, and used the results to make decisions.Experience with at least one of: learning from demonstration, agent memory, or prompt or program optimisation.Strong Python. Comfortable with statistics and experiment design.Deep knowledge of the core AI topics below.A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.Nice to haveReinforcement learning or fine-tuning for LLMs (LoRA / PEFT, SFT, preference tuning).Graph databases (e.g. Memgraph, Neo4j) used as agent memory.LLM observability and evaluation tools (e.g. Langfuse, Arize Phoenix).Working with recorded browser sessions (HAR / WARC, rrweb) as test environments.Published research, writing, or open-source work on agent evaluation or learning.Core AI knowledgeWe expect every AI Engineer on our team to be able to discuss these topics with confidence. For this role we expect depth in all of them, especially evaluation and benchmarks.How LLMs work:the Transformer architecture (attention, positional encodings, feed-forward layers), tokenisation, autoregressive decoding and sampling, and KV cachingthe training pipeline: pre-training, SFT, preference optimisation such as RLHF and DPO, and RL with verifiable rewardsreasoning models and test-time compute; Mixture-of-Experts, quantisation, and servingfailure modes: hallucination, prompt injection, context rot, and reward hackingAgentic AI:agent patterns: ReAct, plan-and-execute, reflection, orchestrator–worker, and evaluator–optimisertool design and context engineering: retrieval, memory, compaction, and sub-agentsthe trade-offs of multi-agent systemssafety controls: permissions, sandboxing, human-in-the-loop, and prompt-injection defenceProtocols:Model Context Protocol (MCP): hosts, clients and servers; tools, resources and prompts; transports; and OAuthAgent2Agent (A2A)awareness of AG-UI and the OpenTelemetry GenAI conventionsAgent harnesses:what sits around the model: the loop, tool registry, permission layers, sandboxes, context management, hooks, and skillshands-on experience with at least one harness or SDK (e.g. Claude Agent SDK, OpenAI Agents SDK, LangGraph, Google ADK, Pydantic AI, or a custom one)why the choice of harness changes benchmark scoresBehavioural vs procedural approaches:when to let the model decide the steps (behavioural) and when to define them in code or workflows (procedural)how to combine the two, and how the balance shifts as models improvethis applies both to system design and to how instructions are writtenBenchmarks:what current benchmarks measure and what they miss:computer use and web: OSWorld, WebArena, Online-Mind2Web, ScreenSpot, BrowseCompagents and tools: SWE-bench Verified / Pro, Terminal-Bench, τ²-bench, GAIAreasoning: Humanity's Last Exam, ARC-AGI, GPQA Diamondhow to read results critically: contamination, harness effects, and costWhat we will assessEvaluation design: design an evaluation suite for an agent that operates a web application, including how you would validate an LLM-as-judge.Critical thinking: interpret a set of benchmark and evaluation results, and make a model or approach recommendation.System design: how an agent should detect that an application has changed and update what it knows, safely.Fundamentals: LLM training, evaluation, and the behavioural vs procedural trade-off.Why joinOwn the part of the platform that makes agents trustworthy: learning, self-correction, and proof that they work.Turn current research on agent learning and evaluation into production engineering.Shape the AI engineering team and its evaluation standards from day one.

Sign in to apply

Create a free account to apply for AI Engineer – Agent Learning & Evaluation — it takes about 30 seconds.

One-click apply and track your application status
Save jobs and build your shortlist
Get alerts for new AI & ML jobs in UAE
About Confidential Careers

Industry

IT Services and IT Consulting

Application Tips

Tailor your CV

Highlight your most relevant AI/ML experience

Research Confidential Careers

Check their AI products and latest news

Show impact

Use metrics to quantify your achievements