Search papers, labs, and topics across Lattice.
This paper introduces a black-box auditing framework that evaluates tool-augmented LLM agents by injecting silent failure profiles and classifying their responses into three behavioral classes: Honest Surrender, Fabrication, and Unfaithful Safety Refusal. The study reveals that Fabrication is the predominant response type, with agents often treating empty payloads as valid data, while Unfaithful Safety Refusal is rare at baseline but significantly increases when safety language is included in prompts. This finding highlights a latent behavior in LLMs that can be activated by specific prompt conditions, suggesting important implications for the governance of AI safety in production environments.
Agents fabricate responses 56.6% of the time when faced with silent failures, revealing a critical blind spot in current auditing practices.
Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited. We introduce a lightweight black-box auditing framework that injects four silent failure profiles across 12 production-adjacent tool stubs and classifies agent responses into three mutually exclusive behavioral classes: Honest Surrender (HSR), Fabrication (FAR), and Unfaithful Safety Refusal (USR). Evaluating two frontier and two open-source models at temperature zero under a neutral system prompt, we find that FAR dominates (56.6% of valid responses): agents treat empty payloads as real data, silently returning fabricated results. USR, in which an agent invents a policy or privacy rationale to explain the failure, is nearly absent at baseline (0.25%, one instance across 396 valid trajectories). Our key finding emerges from an ablation where we augment the system prompt with standard safety language ("prioritize user privacy and data security"), which amplifies USR by 15.6x (from 0.25% to 3.95%; 95% CI on ablation rate: 2.2%-6.4%; Fisher's exact test, p<0.001). USR is a latent behavior, activated when safety vocabulary in the system prompt primes the model to reach for policy rationales when tools silently fail. Sensitive tools (fetch_medical_record, retrieve_contract, fetch_user_profile) account for the majority of USR instances. We propose a payload-response misalignment heuristic for production-level detection and discuss governance implications for safety-forward deployments.