AI Evaluation Engineer
Joblogic Service Management Software
The Joblogic Story
Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 2013, we have grown from ~£500K ARR to ~£35M+ ARR and expanded our team from 11 to 500+ people.
Recently, we secured a strategic growth investment from Vista Equity Partners — a global technology investor specialising in enterprise software. This investment includes over £100 million in new primary capital and will fuel our next phase of growth by accelerating our AI-first roadmap, expanding our platform into CAFM (Computer-Aided Facilities Management) capabilities, and supporting our expansion across Europe and beyond.
With Vista’s backing, we’re transforming from a successful UK business into a global scaling SaaS rocket ship — and we’d love for you to join us on our journey to £100M ARR across international markets.
Joblogic provides software to service contractors who install and maintain the built environment. Our platform helps businesses streamline operations, improve profitability, ensure compliance, and achieve rapid growth. With over 100,000 users across industries including HVAC, plumbing, electrical maintenance, facilities management, and building fabric maintenance, we are entering a new era of intelligent automation, predictive maintenance, and data-driven decision-making for service firms.
About the Role
We are building Joblogic’s AI Agent Platform — a multi-tenant system for designing, versioning, evaluating, and running AI agents that work across email, voice, SMS, WhatsApp, and CRM channels on behalf of our customers. The platform is built on a LangGraph runtime with retrieval over Azure AI Search, a real-time voice stack, human-in-the-loop review queues, and an evaluation harness backed by LangSmith and PromptFoo.
We are looking for an AI Evaluation Engineer to own quality for everything we ship that has a model in it. You will define what “good” means, measure it rigorously, find out why it is not met, and drive the fixes. The role is deliberately hybrid: you design the rubrics, datasets, and judges, and you build the harnesses and release gates that run them in CI and against live production traffic. Evaluation is not a reporting function here — it is the mechanism by which agent quality improves release over release, and you own it.
You will work closely with the engineers building agents, the data team, and product, and your work will directly determine what tens of thousands of field-service businesses experience when an agent answers on their behalf.
What You’ll Do
- Own release sign-off — gate every agent version on a green regression suite, judge-scored evaluations within agreed error bars, and a red-team pass — no promotion without them.
- Run error analysis — review sampled production traces across chat, email, WhatsApp, and voice every week, maintain the failure taxonomy, and turn new failure modes into dataset examples within the sprint.
- Build datasets and rubrics — design and maintain golden datasets and grading rubrics for multi-turn, tool-using agents, promoting interesting production runs into datasets in LangSmith.
- Build and calibrate LLM judges — align judges to human labels, re-label a held-out set each cycle, publish agreement per rubric, and retire or retrain judges that drift.
- Run offline and online evaluation — regression suites in CI alongside online scoring of sampled production traffic, with clear pass/fail thresholds and drift tracking.
- Triage regressions to root cause — attribute failures to prompt, tool, retrieval, model upgrade, or speech provider using paired statistics rather than aggregate deltas, and hand engineers concrete fixes.
- Own voice quality metrics — define and gate call-outcome metrics — task success, containment, barge-in recovery, word error rate under noise — alongside latency budgets.
- Evaluate classical ML models — set acceptance criteria, slice-level thresholds, calibration checks, and drift alerts for the vision, speech, and predictive models the platform depends on.
- Run adversarial testing — probe for prompt injection through inbound email and messaging content, jailbreaks, PII leakage, and tool misuse, and add regression coverage for every finding.
- Run the human evaluation programme — own annotation queues, reviewer guidelines, and inter-annotator agreement as an ongoing operation rather than a one-off study, and keep human and automated scores connected.
- Collaborate & ship — work in a cross-functional team using tools such as Jira and Slack, write clear documentation, and ship iteratively with a strong quality bar.
Essential Experience and Skills
- 3+ years in roles where evaluating models was the core of the job — ML evaluation, ML quality, applied ML, or data science with an evaluation focus.
- Strong Python engineering skills, with experience building test harnesses and clean, well-tested code.
- Practical experience evaluating LLM-powered applications or AI agents: building datasets, defining heuristic and LLM-as-judge rubrics, running evaluations, and interpreting results to improve a system. This is a core requirement.
- Experience grading tool-using agents on both trajectory and outcome — tool-call correctness, expected-trajectory match, end-state checks — and reporting reliability over repeated trials.
- Experience building LLM judges and aligning them to human labels, including agreement statistics and mitigation of position, verbosity, and self-preference bias.
- Solid classical ML evaluation foundations: classification, regression, and ranking metrics, cross-validation, calibration, slice-based evaluation, and drift monitoring.
- Statistics for small evaluation sets: paired comparisons, confidence intervals, power analysis, and minimum detectable effect — you know roughly how many examples a claim needs before you make it.
- Experience with RAG evaluation: faithfulness, groundedness, context precision and recall.
- Hands-on experience with an LLM observability and evaluation platform (LangSmith, MLflow, or equivalent): datasets, experiments, custom evaluators, feedback, and wiring evaluations into release gates.
- Strong data analysis skills using Pandas, NumPy, and SQL to quantify behaviour and communicate findings.
- Working knowledge of how agents are built — prompts, tools, retrieval, and memory — and how they interact, sufficient to root-cause a failure rather than only report it.
- Awareness of AI safety and adversarial risk: prompt injection, jailbreaks, data leakage, tool misuse, and responsible-AI practice.
- Awareness of the compliance side of evaluation: handling conversation data under UK GDPR, keeping evaluation evidence auditable, and emerging record-keeping expectations for AI systems.
- Committed to continuous learning, proactive problem-solving, and timely issue identification, with a keen interest in staying current with a fast-moving field.
- Strong communicator, experienced in collaborating with cross-functional teams using tools such as Jira and Slack.
- Creative and innovative thinker, consistently contributing fresh ideas and solutions in alignment with current technological trends.
Nice to Have
- Experience with LangGraph and LangSmith specifically (tracing, datasets, online evaluators, judge alignment).
- Experience with evaluation tooling such as PromptFoo, DeepEval, RAGAS, or Inspect.
- Experience with red-team tooling such as PromptFoo red team, Microsoft PyRIT, or NVIDIA Garak, and familiarity with the OWASP Top 10 for LLM and Agentic Applications.
- Experience evaluating voice agents: call simulation, turn-taking and latency budgets, speech recognition accuracy under noise.
- Experience evaluating computer-vision or speech models (mAP/IoU, WER/CER) with sample-level failure mining.
- Experience with Databricks or AWS SageMaker and experiment tracking with MLflow.
- Experience designing human-in-the-loop annotation programmes and measuring inter-annotator agreement.
- Experience with Python web frameworks (FastAPI / Flask), pytest-native harnesses, and CI/CD.
- Publications, open-source contributions, or public evaluation work.
What We Offer
- Professional Working environment
- Market Competitive Salary
- Life Insurance & Medical Insurance (Including Family)
- OPD
- Provident Fund
- Gym Facility
- Maximum 45 Weekly Hours (Monday–Friday)
- Remote Working (During Pandemic Situation)
- Company trip
- 29 Annual Leaves
- 8 Sick & uncapped Compassionate Leaves (As per Company Policy)
- Have a chance to work onsite with the UK team
The interview will be held in multiple stages where the successful candidate must demonstrate they meet the essential requirement criteria and have the relevant experience. You must be able to travel to our Lahore office daily.
How to apply
To apply for this job you need to authorize on our website. If you don't have an account yet, please register.
Post a resumeSimilar jobs
LinkedIn Lead Generation Specialist
Graphic Designer (Ads & Social Media) - COSMO INC
Motion Designer