Creating Documents with Voice: Maybe it is not about Transcription but Reflection?

Abstract:

Traditional dictation systems implicitly treat speech as equivalent to writing, overlooking the recursive nature of composition and penalizing cognitive pauses essential for reflection. In an exploratory study (N= 10), participants dictated formal and informal emails, then compared raw transcripts, manually edited versions, and Large Language Model (LLM)-transformed variants (Llama 3.1-8B/3.3-70B). No participant preferred raw output; while LLM processing helped with formal tasks, most preferred self-edited versions for authorial control. One-shot transformation proved vulnerable to Automatic Speech Recognition (ASR) error propagation and stylistic mismatches in informal contexts. These findings motivate a thought-to-text paradigm, reconceiving dictation as collaborative composition rather than linear transcription.


Year: 2026
In session: Voice, Language and Cognition
Pages: 143 to 150