Search papers, labs, and topics across Lattice.
This paper investigates the use of human writing samples to strategically paraphrase machine-generated text, aiming to make it indistinguishable from human output. The authors demonstrate that repeated paraphrasing can effectively shift the distribution of machine-generated responses towards that of human writing, providing a convergence rate and scaling analysis for the number of samples and paraphrasing rounds needed to achieve a specific error threshold. These findings highlight the potential for advanced paraphrasing techniques to blur the lines between human and AI-generated content, raising important implications for text authenticity and detection methods.
Repeated paraphrasing can significantly enhance the indistinguishability of AI-generated text from human writing, challenging existing detection methods.
The rapid improvement of LLMs has made distinguishing AI-generated text from human writing a pressing problem. This challenge is further amplified by paraphrasing tools designed to make machine-generated text appear more"human". We study how access to human writing samples can be used to strategically paraphrase machine-generated responses toward the human distribution. Under a multi-sample setting with human and machine responses to the same prompts, we show that repeated paraphrasing moves the machine distribution toward the empirical human distribution under simple mixing and stability conditions. Our results derive an explicit convergence rate, extend the analysis to a finite-sample setting, and characterize how the required number of human samples and paraphrasing rounds scale with the desired error.