FIRE 2026

Forum for Information Retrieval Evaluation

Indian Statistical Institute, Kolkata

17th-20th December

Most research on search agents is driven by benchmark evaluation. Yet practitioners increasingly face a different question: how to choose a model for a task that no benchmark covers yet. This talk proposes that the search behaviour of a model offers grounds for such decisions. Using controlled experimentation on model search behaviour in product search, I show that models tend to differ in how they click, judge relevance, and reformulate, and that these differences are broadly stable across collections, forming behavioural signatures. Furthermore, models of indistinguishable effectiveness deliver largely different results to the user due to their signatures. I then discuss how behavioural signatures can inform which tasks a model suits when no benchmark is available.
For decades, IR evaluation has rested on a simple and effective idea: fix a set of queries, collect relevance judgments, and compare ranked lists. This test-collection approach gave us TREC, CLEF, NTCIR, and FIRE, and it made IR one of the most empirically rigorous areas of computing. Agentic IR systems strain nearly every assumption behind it. An agent does not simply return a ranked list. It plans, issues its own queries, calls tools, reads and synthesizes sources, and acts on a person's behalf, often over many steps. So what exactly do we judge, and who does the judging?
In this talk, I will draw on recent work from my lab and collaborators to lay out five challenges: (1) static benchmarks that cannot keep up with agents that adapt, personalize, and may have memorized the test; (2) simulated users that are far more cooperative than real people; (3) the gap between a correct answer and a useful outcome for the person's task; (4) the circularity of using LLMs to judge LLM-based agents; and (5) the "delegation paradox," where the more we hand off to agents, the more auditing, trust, and oversight we need. For each, I will share what we have tried, what worked, and what did not. I will argue that these questions get harder as agents serve people across languages and cultures, which is where FIRE's community has much to contribute, including through this year's Multilingual Agentic Search Task.


ACM SIGIR


TCS Research


To be announced soon.