On-siteFull time

AI Research Engineer

EUR 100000 - 120000 / Year

Dublin city centre location

5 days onsite

Small, ambitious research team

Work on frontier models and benchmarking

Specialist AI research role focused on benchmarking, evaluation and adversarial testing of frontier LLMs within a fast-moving Dublin research team working fully onsite.

Overview

This is a deliberately specialised role for someone deeply focused on frontier AI models rather than product integrations.

You will join a small research team in Dublin working on training and evaluation environments for leading AI labs. The team moves quickly and works hard, so this role will suit someone who actively wants a start-up environment rather than a traditional 9-to-5.

This is a 5 days per week onsite role in Dublin city centre.

What You'll Be Doing

  • Designing and building benchmarks for frontier LLMs
  • Comparing how different models perform and behave
  • Building difficult tasks and datasets to expose model strengths and weaknesses
  • Adversarially testing models and finding where they break
  • Designing rigorous evaluation methodologies
  • Working with synthetic training and evaluation data
  • Investigating model reasoning, reliability and unexpected behaviour
  • Working on post-training and reinforcement learning
  • Running experiments across multiple frontier models

What You'll Need to Succeed

Experience in one or more of the following:

  • Building or contributing to recognised or public LLM benchmarks
  • LLM benchmarking or model capability evaluation
  • Adversarial LLM testing or red teaming
  • AI safety or model safety research
  • LLM post-training or reinforcement learning
  • Synthetic data generation for model training or evaluation
  • Designing evaluation frameworks, metrics or model judges
  • Research into LLM behaviour and failure modes
  • Working deeply across multiple models such as Claude, GPT, Gemini, Llama or DeepSeek

Strong research credentials are highly valued. This could include:

  • A PhD in a relevant area
  • Significant research experience
  • Publications at conferences such as NeurIPS, ICML or ICLR
  • Demonstrably strong work building benchmarks and evaluation systems

Most importantly, this role is looking for people who have actually done this work, rather than simply used the terminology.

Likely Not a Fit If Your Experience Is Mainly

  • RAG
  • LangChain
  • Vector databases
  • Prompt engineering
  • Building chatbots
  • Agent orchestration
  • Calling LLM APIs
  • Adding GenAI features to existing products

Without deeper model evaluation, benchmarking or research experience, this role is unlikely to be a fit.

What's In It for You

  • Salary of EUR 100000 - 120000 / Year
  • Opportunity to work on frontier model benchmarking and evaluation
  • Small, ambitious team environment
  • Exposure to multiple leading frontier models
  • Central Dublin location with fully onsite collaboration

If you are the kind of person who sees a new frontier model released and immediately wants to test it, break it, compare it and understand why it behaves differently, this role could be a strong match.

Get notified about your perfect job

Send us your CV and we'll message you when we find a good match.

upload CV