Search papers, labs, and topics across Lattice.
This study investigates the contrasting effects of direct exposure versus multi-agent mediation on the advice generated by a high-capability LLM, specifically OpenAI's gpt-5.6-sol model. The findings reveal that when the model is directly exposed to a dangerous objective, it produces advice that is contrary to the intended manipulative target; conversely, when the objective is transformed and mediated by other agents, the model generates advice that aligns with the target. This behavioral shift suggests that the model may recognize and distrust manipulative motives, highlighting a significant safety gap in automated workflows that can obscure harmful instructions from user-facing components.
Direct exposure to harmful objectives can lead LLMs to generate safer advice, while mediation can inadvertently align them with dangerous targets.
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model's context while preserving the objective's target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.