Search papers, labs, and topics across Lattice.
This study investigates the effects of human input variations鈥攕pecifically voice transcription and keyboard typing鈥攐n the performance of instruction-tuned language models using a novel tool called HIVE (Human Input-Variation Engine). The findings reveal that voice transcription perturbations significantly reduce model accuracy due to structural issues in the transcription, while keyboard perturbations have a lesser impact, allowing models to absorb more noise before performance declines. Notably, the research highlights that the detrimental effects stem from the survival rate of tokens during perturbation, with implications for how models handle input variations in real-world applications.
Voice input significantly hampers LLM accuracy due to structural transcription issues, while keyboard input is surprisingly more resilient.
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and restructuring from AI-backed dictation tools. How do they impact an LLM's performance? In this paper we present HIVE (Human Input-Variation Engine), a suite of voice transcription perturbations and QWERTY keyboard perturbations. We use HIVE to evaluate how robust models are to these perturbations. We present seven findings. (i) Voice transcription perturbations lower accuracy across every instruction-tuned model we test, and it is the structure of the transcription rather than its fillers that carries the cost. (ii) QWERTY keyboard perturbations cost less, and a model absorbs a lot of them before accuracy falls away. (iii) Both trace back to one cause, how many of the question's tokens survive the perturbation: destroying a token is what hurts, while adding new ones alongside it costs little. (iv) The gap between the two channels appears only where the answer must be constructed or deduced; on multiple choice there is none. (v) The harm does not solely come from test-set contamination. (vi) It cannot be trained away with lightweight adaptation. (vii) A thinking budget recovers the keyboard channel almost entirely but leaves the spoken registers untouched, and compressed speech is worse with it.