Search papers, labs, and topics across Lattice.
This study investigates how large language models (LLMs) handle knowledge boundaries and referent specificity through a Gricean framework, revealing that LLMs often fabricate details instead of retreating to safer, more general claims when faced with unknown entities. By employing a T-REx-based benchmark, the authors demonstrate that while LLMs can recognize when a referent is outside their knowledge boundary and anticipate referent specificity, they still generate overly specific responses instead of opting for generic alternatives. The findings highlight a gap between the models' capabilities and their generation policies, suggesting a need for improved alignment strategies that connect knowledge awareness with referent specificity during text generation.
LLMs can recognize when they lack knowledge about a referent but still choose to fabricate specific details instead of opting for safer, generic responses.
When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to safer, more general claims. We frame this failure through a Gricean lens: a cooperative speaker who is uncertain about a referent retreats up the specificity hierarchy, trading informativeness for truthfulness. We ask whether LLMs have the ingredients to perform this retreat. Using a T-REx-based benchmark that varies entity familiarity and referent specificity, we probe models to answer two questions: (i) do their activations encode whether a referent falls inside the knowledge boundary, and (ii) do they anticipate the specificity of the referent they are about to generate? We find that the answer to both is yes, but the two signals are not reconciled in generation. Models overwhelmingly prefer specific referents even when the entity is unknown to them, and do so even when offered correct generic alternatives. The substrate for a Gricean retreat is present, but the policy that would act on it is not. We position our findings as a first step toward Gricean alignment, training or steering objectives that couple knowledge-boundary awareness to referent-specificity during generation.