Search papers, labs, and topics across Lattice.
This study evaluates the impact of identity-conditioned prompts on the readability of LLM-generated design descriptions for surveillance and security robots, utilizing 236 demographic identity labels. The findings reveal significant variations in readability based on prompt conditions and demographic identities, highlighting the potential influence of these factors on the conceptualization and documentation of robotic systems. By establishing readability as a benchmark, the research lays the groundwork for a comprehensive framework that includes additional analyses of fairness and semantic content in LLM outputs.
Identity-conditioned prompts can lead to significant differences in the readability of LLM-generated robot design descriptions, raising critical questions about bias in AI outputs.
Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.