GPT-5 Powered Agentic Search Smashes Baseline Scores But Hits Wall on Knowledge-Gap Tasks
GPT-5-powered agentic search systems are crushing baseline scores with an NDCG of 0.453 versus a 0.289 BM25 baseline on Amazon ESCI, but hit a hard wall when facing knowledge-gap tasks where LLMs cannot evaluate information they don't already know.