Search papers, labs, and topics across Lattice.
This study evaluates an LLM-based search agent that utilizes a Boolean retrieval engine to navigate the MS MARCO V2.1 dataset, achieving an NDCG@10 score of 0.6863 across 86 topics with a budget of 100 model calls per topic. The agent's ranking mechanism relies solely on substring density matching, eschewing the need for supervised learning or complex statistical methods. These exploratory results indicate that straightforward pattern matching could effectively support agentic search, challenging the necessity for more sophisticated retrieval techniques.
Simple Boolean queries can outperform complex retrieval methods, achieving competitive search performance without the need for supervised learning.
We equipped an LLM-based search agent with access to a Boolean retrieval engine to search the MS MARCO V2.1 deduped segment collection used by the TREC 2024 RAG track. Over a standard track subset of 86 topics, and operating under a budget of 100 model calls/topic, the agent achieved an NDCG@10 of 0.6863, which would place it above many dense, sparse, and learned-sparse first-stage retrievers. Ranking is based solely on the density of corpus substrings matching a query, with no requirement for supervised learning, global statistics, or term weights. Formally, the query language expresses a strict subset of the regular languages, with a document's score based on the number and length of matches it contains. Although the results are more exploratory than definitive, because they are based on a single test collection that was publicly available during model training, they suggest that simple pattern matching may be sufficient for agentic search.