Search papers, labs, and topics across Lattice.
This study investigates the effectiveness of user-authored permission policies in mitigating AI agent overreach compared to traditional human-in-the-loop (HITL) and automated review methods. While participants using user-authored policies (POLICY) experienced a reduction in runtime prompts, they ultimately blocked less overreach than both HITL and automated methods, revealing a tendency to approve actions that deviated from initial requests. The findings highlight a significant gap between user preferences for control and the actual protective capabilities of user-defined rules, suggesting that while these policies streamline decision-making, they may inadvertently allow for more overreach due to user approval behaviors.
User-authored permission policies may streamline interactions with AI agents, but they paradoxically enable more overreach by allowing users to approve actions outside their original intent.
AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored"allow","ask", or"never"rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions. POLICY blocked less overreach than HITL (-20.1 percentage points, 95% CI [-32.1, -8.1]) and AUTO (-14.5 points, 95% CI [-25.8, -3.2]). POLICY lowered runtime prompts from 18.0 to 10.9, but total intervention time was not reliably lower when rule setup was included. Exploratory analysis showed that participants chose"ask"for 114 of 140 POLICY rules, returning most overreach actions to runtime. Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under"allow"rules. Across all 7 overreach actions, POLICY had the highest approval rate. Counterintuitively, user-authored rules did not by themselves provide stronger protection: many actions outside users'original requests went through after users approved them. These results reveal a gap between preference and commitment: repeatedly choosing"ask"preserves case-by-case choice but prevents a standing policy from settling decisions in advance.