How high-fidelity speech synthesis converges on an embodied simulation

A powerful waveform model does not need a literal anatomical simulator with variables named lung_volume, tongue_position, or thyroarytenoid_activation. It can learn shortcuts, and for ordinary speech those shortcuts may work remarkably well.

They become less reliable as the demands increase. A model may be asked to remain coherent across long utterances; respond to unusual instructions or physical perturbations; preserve one speaker’s physiology from scene to scene; produce combinations absent from its training data; distinguish similar sounds with different causes; interact in real time instead of reciting; or expose independent control over the causes of a sound.

Beyond some threshold, the hidden state has to represent variables that are functionally equivalent to respiratory, laryngeal, articulatory, sensorimotor, autonomic, cognitive, and social bodily states. The model begins to operate as an embodied simulation in the causal sense, whether or not its developers describe it that way.

Three stages help clarify the argument:

  • Acoustic imitation: reproduce a waveform that resembles observed speech.
  • Embodied generativity: maintain a hidden organism-like state whose changes cause the waveform.
  • Embodied agency: perceive, predict, regulate, and adapt that state through interaction.

The progression is less like adding an “emotion” setting to speech and more like following the causal chain backward:

waveformvocal apparatuswhole-body physiologysensorimotor controlcognitive, affective, and social state\text{waveform} \leftarrow \text{vocal apparatus} \leftarrow \text{whole-body physiology} \leftarrow \text{sensorimotor control} \leftarrow \text{cognitive, affective, and social state}

Where the shortcuts start to fail

ThresholdWhat begins failing without embodied state
1. IntelligibilityPhonemes, formants, and consonant timing
2. Local naturalnessCoarticulation, voice quality, and microprosody
3. Utterance coherenceBreath groups, fatigue, and phrase-scale pitch and intensity
4. Expressive realismSmiling, tension, laughter, hesitation, and mixed emotion
5. Speaker persistenceStable anatomy, habits, health, age, and vocal limits
6. Interactive realismListening while preparing, turn-taking, entrainment, and accommodation
7. Counterfactual control”Say it while smiling nervously, whispering, walking uphill, and trying not to laugh”
8. Causal robustnessNovel combinations, pathology, perturbation, recovery, and adaptation

These thresholds are not a single ladder every system must climb in order. They mark places where a surface model is increasingly pressured to reconstruct the system that generated the surface.

I. Respiration: the power supply

#What is heardSpeech or clinical termWhat the model increasingly needs
1Where inhalations occurSpeech breathing; breath-group planningLung volume, intended utterance length, syntactic plan, and available expiratory reserve
2Phrase lengthUtterance per breath; maximum phonation timeVital-capacity-like limits, airflow consumption, and upcoming linguistic material
3Loudness across a phraseSubglottal pressure; SPL contourRespiratory muscle recruitment and pressure decay
4Pitch drifting as air runs downLung-volume effect on F0Coupling among recoil pressure, laryngeal posture, and vocal-fold tension
5Audible inhalationInspiratory noise; oral or nasal inspirationGlottal opening, nasal versus oral route, flow rate, and urgency
6BreathlessnessDyspnea-like phonation; reduced breath supportElevated respiratory drive, shortened phrases, rapid recovery breaths, and altered effort
7Speech during exertionRespiratory-phonatory competitionLocomotor demand, metabolic load, and the allocation of expiration between speech and gas exchange
8Phrase-final weakeningDeclination; terminal intensity fall; breathy offsetFalling pressure, laryngeal compensation, and the choice to trail off or replenish
9Forceful projectionVocal effort; respiratory driveAbdominal and thoracic recruitment, subglottal pressure, and glottal resistance
10Whisper duration and instabilityWhisper airflow; turbulent excitationHigh airflow consumption without periodic vocal-fold vibration
11Sighs and voiced sighsParalinguistic expirationDeep inspiration, passive recoil, and gradual laryngeal engagement
12Coughs, throat-clears, and suppressed coughsProtective laryngeal behavioursCompression, glottal closure, explosive release, and competing social inhibition
13Laugh-breath cyclesEgressive laughter pulsesRapid intercostal contractions, expiratory bursts, phonated or unphonated pulses, and replenishing inhalations
14Crying speechRespiratory interruption; sobbingIrregular inhalatory spasms, glottal constriction, and disrupted phrase planning
15Breath timing in dialoguePre-turn inspiration; conversational breathingPrediction of a turn opportunity, response preparation, and willingness to claim the floor

Speech breathing changes with utterance length and loudness. Conversational inhalation is tied to response planning and turn management. Laughter has its own respiratory organization; it is not an indefinitely repeatable string of ha tokens. [1]

II. The larynx: the voice source

#What is heardRelevant termRequired hidden variables
16Fundamental frequencyF0; vocal-fold oscillation rateVocal-fold length, mass, longitudinal stiffness, and CT/TA muscle balance
17IntensitySPL; phonatory powerSubglottal pressure, glottal resistance, and closure behaviour
18Breathy voiceIncomplete glottal closure; high H1-H2; spectral tiltResting glottal gap, adduction, tissue geometry, and airflow leakage
19Pressed voiceHyperadduction; low H1-H2; high contact quotientMedial compression, intrinsic laryngeal activation, and pressure
20Creak or vocal fryIrregular phonation; pulse register; high jitterLow F0, slack or short vocal folds, irregular closure timing, and low airflow
21Falsetto or head-register qualityRegister transition; thin-fold vibrationReduced vibrating mass, altered TA/CT balance, and incomplete closure
22Chest-register qualityThicker medial surface; stronger harmonic excitationGreater effective fold thickness and closure duration
23Register breaksBifurcation; mode transitionNonlinear tissue dynamics and instability across pressure-tension combinations
24Smooth onsetCoordinated onsetTemporal alignment between airflow and vocal-fold adduction
25Aspirate onsetPositive voice-onset intervalAirflow preceding complete laryngeal engagement
26Hard or glottal onsetGlottal attackStrong prephonatory adduction followed by abrupt pressure-driven oscillation
27Whisper-to-voice transitionPhonation threshold pressurePressure, fold approximation, stiffness, and nonlinear onset conditions
28JitterCycle-to-cycle F0 perturbationOscillatory instability, motor noise, and tissue asymmetry
29ShimmerCycle-to-cycle amplitude perturbationPressure variation, closure variation, and source instability
30Harmonics-to-noise ratioHNR; cepstral peak prominencePeriodicity, turbulent leakage, and regularity of tissue vibration
31Diplophonia or subharmonicsPeriod doubling; nonlinear phonationAsymmetric or multimodal fold oscillation
32Vocal tremorRhythmic F0 or amplitude modulationOscillatory motor commands across respiratory and laryngeal subsystems
33StrainMuscle-tension-dysphonia-like qualityExcess intrinsic and extrinsic laryngeal tension, elevated laryngeal posture, and restricted vibration
34A momentary voice crack under emotionPhonatory instabilityRapid autonomic and motor changes that push the system through a nonlinear boundary
35Voicing distinctionsLaryngeal timing; VOT; aspirationPrecise coordination of oral release, glottal opening, and onset of oscillation
36Tonal-language pitchLexical tone productionDeliberate, syllable-aligned laryngeal tension trajectories rather than generic sentence melody

Fold stiffness, thickness, glottal opening, and subglottal pressure all affect the resulting acoustics. They also interact nonlinearly. A generator with independent controls for pitch, breathiness, and loudness will eventually produce impossible combinations unless those controls share an underlying biomechanics. [2]

III. Source-filter coupling: where modular controls break

#What is heardTermWhat must be simulated
37Vowel-dependent changes in voice qualitySource-filter interactionAcoustic feedback from the vocal tract into glottal flow and fold vibration
38Pitch instability near a formantNonlinear source-tract couplingF0-resonance proximity, epilaryngeal geometry, and acoustic loading
39Loudness efficiencyInertive reactance; impedance matchingA vocal-tract configuration that assists or opposes glottal oscillation
40Register behaviour changing with vowelsFilter-dependent phonationJoint simulation of larynx and tract instead of independent pitch and timbre sliders
41Singing resonance strategiesFormant tuningDeliberate tract reshaping around source harmonics
42Roughness caused by tract configurationSource destabilizationPressure feedback from supraglottal and subglottal cavities

The classical linear source-filter decomposition remains useful, but living anatomy is less tidy. Aerodynamic and acoustic coupling allow tongue position, laryngeal height, epilaryngeal geometry, and tract resonances to change the source itself. [3]

IV. Supralaryngeal articulation: the moving acoustic cavity

#What is heardPhonetic termRequired simulation
43Vowel identityF1, F2, F3; vowel spaceTongue-body shape, jaw aperture, lip posture, and pharyngeal geometry
44Speaker-consistent vowelsVocal-tract-length normalizationStable anatomy plus habitual articulatory targets
45Lip roundingLabialization; lowered formantsLip protrusion, aperture, and effective tract length
46Smiling speechLip spreading; apparent tract shorteningLip-corner displacement, cheek tension, and altered cavity geometry
47BilabialsOral closure; burst releaseUpper and lower lip contact, intraoral pressure, and release timing
48AlveolarsTongue-tip constrictionTongue-tip or blade placement and release
49VelarsDorsal constriction; velar pinchTongue dorsum, palate geometry, and vowel-dependent contact location
50SibilantsSpectral centre of gravity; spectral momentsJet formation, a grooved tongue, obstacle position, and dental geometry
51StopsClosure, burst, aspiration, VOTSealing, pressure accumulation, release, and laryngeal coordination
52FricativesTurbulent source; constriction degreeAirflow through a narrow constriction and the geometry of a downstream obstacle
53AffricatesStop-fricative coordinationControlled transition from complete closure to sustained constriction
54Rhotic variationAlveolar tap, trill, approximant, or uvular RDistinct tongue or uvular gestures, airflow, and, for trills, self-sustained tissue vibration
55LateralsLateral channels; antiformantsCentral tongue contact with side-channel airflow
56NasalsVelopharyngeal opening; nasal formants and antiformantsVelum position, nasal-cavity geometry, and oral side branches
57Nasalized vowelsCoupled oral-nasal resonanceContinuously graded velopharyngeal port opening
58HypernasalityVelopharyngeal insufficiencyPersistent or mistimed oral-nasal coupling
59Nasal emissionTurbulent nasal airflowPressure, port area, oral obstruction, and nasal resistance
60Clicks and ejectivesNon-pulmonic airstream mechanismsMultiple closures, cavity-pressure changes, and glottalic or lingual airflow generation
61Fine consonant transitionsFormant transitionsContinuous articulator trajectories rather than phoneme-wise concatenation
62Reduced casual speechLenition; undershoot; assimilationEffort-sensitive trajectories and incomplete attainment of ideal targets
63Clear speechHyperarticulation; expanded vowel spaceIncreased movement amplitude and duration, guided by listener-oriented motor goals
64MumblingHypoarticulationReduced jaw excursion, compressed vowel space, and lower articulatory effort

Measured vocal-tract geometry can be converted into area functions that predict resonance. Even small changes in tract length affect perceived body size and speaker category; smiling changes acoustics partly because spreading the lips shortens the tract. [4]

V. Coordination and coarticulation: speech is not a row of sounds

#What is heardTermWhat must be simulated
65A consonant changing with its vowel contextAnticipatory or carry-over coarticulationOverlapping articulatory gestures and context-sensitive targets
66Natural syllable transitionsGestural phasingRelative timing of tongue, jaw, lips, velum, and larynx
67Assimilation across word boundariesConnected-speech processesPlanning beyond the current phoneme or word
68Segment-duration changesCompensatory timingRedistribution of movement time under rate, emphasis, and complexity constraints
69Consonant-induced pitch perturbationsMicroprosodyMechanical and neural coupling between laryngeal gestures and segment type
70Stress changing articulationProsodic strengtheningLarger, longer, or more precise gestures in prominent positions
71Phrase-edge articulationDomain-initial strengthening; final lengtheningHierarchical prosodic planning acting on motor trajectories
72Speech-rate effectsSpatial and temporal reductionBiomechanical velocity and acceleration limits, plus target undershoot
73The tongue compensating for a blocked jawMotor equivalenceTask-level acoustic goals and redundant articulatory solutions
74Stable speech despite perturbationSensorimotor compensationOnline detection of bodily error and redistribution across articulators
75Speaker-specific rhythmIdiolectal motor timingPersistent learned coordination patterns rather than a global tempo control
76Realistic speech errorsPhonological versus motor errorsMultiple planning levels whose failures produce different kinds of output

Articulatory synthesis improves when consonants are generated as context-dependent vocal-tract configurations. Human speakers also preserve acoustic goals by compensating with unperturbed articulators when, for example, the jaw is unexpectedly blocked. [5]

VI. Sensorimotor and neurological control

#What is heardNeurological termHidden process required
77Fluent learned sequencesFeedforward motor commandsStored mappings from phonological targets to coordinated gestures
78Pitch correction after altered self-hearingAuditory feedback control; pitch-shift reflexPredicted auditory target, heard output, and error correction
79Correction of jaw or tongue positionSomatosensory feedback controlProprioceptive and tactile estimates of articulator state
80Rapid self-correctionEfference copy; state estimationPrediction of the sensory result before delayed feedback arrives
81Stable output in noiseLombard effectPerceived noise, intelligibility estimates, and compensatory intensity or articulation
82Speech adapted to room acousticsSidetone; external auditory feedbackHearing reflected voice and adjusting effort
83Sequencing syllablesSpeech motor programming; GODIVA-like planningCompetitive selection and buffering of forthcoming motor chunks
84Timing precisionCerebellar predictionMillisecond-scale coordination and error-based adaptation
85Initiation and scalingBasal-ganglia motor regulationMovement onset, amplitude, automaticity, and vigour
86Speech under divided attentionDual-task interferenceShared executive and motor resources
87Hesitation during lexical searchIncremental formulationA partial linguistic plan interacting with a ready but delayed motor plan
88Apraxic inconsistencyApraxia of speechImpaired planning or programming despite available muscles
89Dysarthric weakness or incoordinationDysarthriaImpairment distributed across respiration, phonation, resonance, articulation, and prosody
90Stuttering-like disruptionInitiation or sequence instabilityCompetition, timing, and feedback dynamics rather than inserted textual repetitions
91Motor adaptation over a conversationSensorimotor learningPersistent updates to predicted body-sound mappings
92Reflexive versus voluntary vocalizationsDual laryngeal control systemsDistinct control pathways for speech, laughter, crying, and airway protection

Speech production is a distributed control problem. Neural models such as DIVA distinguish feedforward commands, auditory targets, and somatosensory targets. Perturbation studies show speakers compensating through bodily feedback, while altered auditory feedback changes voice production. [6]

VII. Whole-body geometry and condition

#What is heardRelevant termWhat must be simulated
93Apparent body sizeVocal-tract length; formant scalingStable head, pharynx, and tract geometry
94Apparent ageDevelopmental vocal-tract morphologyGrowth-dependent tract proportions and vocal-fold development
95Sex-linked morphology cuesF0-VTL interactionFold dimensions and tract geometry, not pitch alone
96Posture-dependent resonancePostural phonation; tract deformationHead, neck, jaw, spine, and laryngeal position
97Voice while lying downSupine versus upright tract acousticsGravity-sensitive soft tissue and airway geometry
98Head turning during speechOrientation-dependent radiationMouth direction, neck configuration, and changing microphone or listener geometry
99Speech while smilingOrofacial expressionCoordinated lips, cheeks, jaw, and sometimes laryngeal or respiratory changes
100Speech while gesturingCo-speech motor couplingShared timing and effort across hand, torso, jaw, and tongue movements
101Walking or dancing speechLocomotor-respiratory entrainmentMovement rhythm, impacts, exertion, and respiratory competition
102Eating or dry-mouth speechOral lubrication; altered articulationSaliva, mucosal contact, swallowing interruptions, and oral obstruction
103Morning voiceTissue state; mucus; edema-like changeTime-varying fold mass, hydration, and secretion state
104Vocal fatigueLaryngeal muscle or tissue fatigueAccumulated loading, reduced control, altered effort, and recovery
105Illness or congestionUpper-airway obstruction; dysphoniaNasal resistance, secretions, inflammation, and altered resonance
106Pain-modified speechProtective guardingRestricted movement, breath-holding, and effort avoidance
107Medication or intoxication effectsSensorimotor alterationArousal, timing, muscle tone, coordination, and salivation changes
108Stable bodily limitsSpeaker-specific range profilePhysiological boundaries on F0, SPL, duration, and voice quality

Posture changes vocal-tract acoustics; body morphology affects tract length and formants; vocal effort changes with room acoustics and accumulated fatigue. Co-speech gesture can even alter tongue and jaw displacement, which means “voice only” is not always motorically separable from the visible body. [7]

VIII. Autonomic, affective, and interoceptive state

#What is heardTermRequired state
109Stress-related pitch changeAutonomic arousal; F0 elevation or variabilitySympathetic activation interacting with laryngeal tension and respiratory drive
110A tight, effortful voicePerilaryngeal tension; hyperfunctionExtrinsic and intrinsic muscle recruitment
111Quivering fearTremulous phonationUnstable motor drive, respiratory irregularity, and heightened arousal
112AngerHigh-arousal prosodyIncreased pressure, intensity, rate, articulation, and laryngeal activation
113SadnessLow activation; reduced prosodic rangeLower movement vigour, intensity, and pitch range, with slower timing
114ExcitementPositive high arousalElevated respiratory drive and pitch variation without anger’s constrictive configuration
115Calm warmthLow arousal; positive valenceLow tension with controlled breath, smooth onset, and flexible resonance
116Nervous smilingAffect-expression conflictSimultaneous lip spreading, laryngeal tension, uncertain timing, and inhibited laughter
117Trying not to laughSuppression; leakageCompeting voluntary inhibition and involuntary respiratory or laryngeal laughter patterns
118Trying not to cryAffective suppressionGlottal constriction, disrupted breathing, swallowing, and pitch instability
119Confidence versus uncertaintyEpistemic prosodyCommitment state affecting tempo, terminal contours, intensity, and hesitation
120Cognitive loadVoice stress; planning loadExecutive demand, autonomic response, and disrupted motor precision
121Emotional mixturesBlended affective prosodyMultiple partly independent systems rather than a single emotion label
122Emotion changing during a sentenceDynamic affect trajectoryContinuous latent state evolving faster or slower than linguistic structure
123Genuine versus posed laughterSpontaneous or volitional laughter acousticsDifferent respiratory, laryngeal, and social-control pathways
124Contagious laughter and vocal entrainmentSocial-affective couplingPerception of another person changing the speaker’s autonomic and motor state

Stress is associated with changes in laryngeal muscle activity and acoustic output, yet no single acoustic cue uniquely identifies an emotion. Faithful synthesis therefore needs a multidimensional physiological state rather than a lookup table that translates emotion=angry into “raise pitch and volume.” [8]

IX. Prosody, phonology, and linguistic embodiment

#What is heardTermEmbodied dependency
125Lexical toneTone contour; tone sandhiLaryngeal trajectories synchronized with syllable structure
126Question versus statementBoundary tonePhrase-level motor plan and pragmatic intention
127Which word is emphasizedNuclear accent; prominenceJoint F0, duration, intensity, and articulatory strengthening
128Contrastive focusFocus prosodyA listener model and selection among alternative meanings
129Stress-timed or syllable-timed rhythmRhythmic organizationLanguage-specific timing attractors imposed on articulatory gestures
130Mora timingMoraic phonologyFine duration control spanning segments and syllables
131GeminationContrastive consonant durationLonger closure or constriction without merely slowing a recording
132Phonemic vowel lengthQuantity contrastDuration integrated with stress, rhythm, and articulation
133Accent and dialectSociophonetics; phonological grammarLearned motor targets, timing patterns, and social-indexical control
134Native-like rolled RLanguage-specific rhotic gestureThe right articulator, airflow, timing, and phonological environment
135Code-switchingLanguage-mode switchingReconfiguration of phonetic targets, prosody, rhythm, and social stance
136Aizuchi and backchannelsContinuers; response tokensTiming, pitch, intensity, and relational stance, often with little lexical content
137Sarcasm or ironyPragmatic prosodyDivergence between literal semantic content and social intention
138Quotation voiceEnactment; reported speechA temporary shift in speaker model and bodily stance
139Vocal stimming and playful soundNonlexical vocalization; paraspeechSensorimotor pleasure, repetition, rhythmic attractors, and expressive control not reducible to text
140Whispered tone or stressProsody without periodic F0Substitute cues such as duration, intensity, spectral shape, and airflow
141Singing-to-speech transitionsSpeech-song continuumShared but differently constrained respiratory, laryngeal, and articulatory systems

X. Interactional embodiment

#What is heardInteraction termWhat must be simulated
142Knowing when someone has finishedTurn-final prosodyPrediction from syntax, pitch, duration, gaze, context, and breath
143Preparing before the other person stopsAnticipatory turn planningConcurrent listening, response formulation, and pre-inspiratory preparation
144Brief overlapsCompetitive or cooperative overlapUrgency, affiliation, and predicted completion time
145Holding the floorTurn-holding cuesContinuing intonation, filled pauses, breath management, and pacing
146Yielding the floorTurn-yielding cuesTerminal contour, slowing, intensity reduction, and completion
147Matching another person’s rhythmEntrainment; accommodationOnline estimation and gradual adaptation of timing, F0, and intensity
148Mirroring vocal fry or accentPhonetic convergenceSocial relation plus adjustable motor habits
149Deliberately refusing to mirrorDivergence; stance markingMetacontrol over accommodation
150Speaking to a child, elder, or hard-of-hearing listenerAudience design; clear speechA listener-specific model of intelligibility
151Intimate versus public voiceRegister; interpersonal distanceBreathiness, loudness, articulation, and social safety
152Micro-signals of acceptanceAffiliative prosodyTiming, softness, tolerance of overlap, and backchannel selection
153Detecting a joke before its words finishProsodic framingShared context and predictive social inference
154Laughing togetherLaughter entrainmentReciprocal timing, breath cycles, and sensitivity to the other person’s state
155Speaking in a noisy roomLombard adaptationEnvironmental sensing, listener distance, and vocal effort
156Speaking in a reverberant roomRoom-voice adaptationExternal auditory feedback and changed effort
157A different voice for a “formal” taskStyle shifting; read versus spontaneous speechSocial role, monitoring level, and learned performance posture
158Continuity through interruptionsConversational statePersistent bodily, emotional, and discourse state across turns

Human turn transitions are extremely fast relative to the time required to plan speech. Listening and embodied response preparation must overlap. Breathing participates in the organization of turns, while face-to-face conversation also recruits gaze and gesture. [9]

XI. Pathology as a causal test

A model can produce generic healthy speech from correlations alone. Pathology is a stronger test of whether it has learned the causal subsystems.

#Audible presentationUnderlying subsystem that must be represented
159Flaccid dysarthriaWeakness, reduced tone, breath support, and incomplete closure
160Spastic dysarthriaHypertonia, strained phonation, and slow, restricted movement
161Ataxic dysarthriaVariability in timing, force, and coordination
162Hypokinetic dysarthriaReduced movement scaling, monopitch, monoloudness, and accelerated bursts
163Hyperkinetic dysarthriaInvoluntary movement entering phonation and articulation
164Apraxia of speechImpaired motor planning, segmentation, and inconsistent errors
165Muscle tension dysphoniaMaladaptive extrinsic or intrinsic laryngeal recruitment
166Vocal-fold paresisAsymmetric closure and vibration
167Nodules or lesionsChanged mass, stiffness, closure, and periodicity
168Velopharyngeal insufficiencyAbnormal nasal coupling and pressure loss
169Cleft-related speechStructural resonance and pressure constraints plus compensatory articulation
170Parkinsonian speechBasal-ganglia-linked changes in scaling and initiation
171Cerebellar diseaseTiming and coordination deficits
172ALS-related speechProgressive subsystem weakness with changing compensations
173Hearing-loss-related speechAltered long-term auditory calibration
174Post-laryngectomy speechA different sound source and different source-tract coupling
175Tracheostomy or respiratory diseaseAltered airflow, pressure, and phrase limits
176StutteringState-dependent initiation, timing, tension, and adaptation
177Functional voice disorderA learned control policy and context-dependent tension without a simple structural lesion
178Recovery or therapy effectsLongitudinal adaptation of motor programs, compensation, and effort

The traditional speech-pathology analysis of respiration, phonation, resonance, articulation, and prosody is almost a ready-made specification for the minimum modular body model needed by clinically faithful synthesis. Dysarthria can reflect impairment across several of these systems, which makes a single “dysarthria style” embedding causally inadequate. [10]

XII. What emerges above the body parts

Once the lower layers are modelled, several larger latent structures become necessary.

179. A persistent body schema

The model must remember tract length, pitch range, lung capacity, habitual posture, speech rate, asymmetries, and current fatigue. Otherwise the apparent body silently changes between sentences.

180. A feasible-action manifold

Not every combination of pitch, intensity, breathiness, speed, and articulatory precision is physically reachable. A causal model needs a representation of possible actions, including the cost of moving between states.

181. Effort

Two acoustically similar outputs can demand different levels of muscular and cognitive effort. That difference affects what comes next: fatigue, breath timing, errors, recovery, and willingness to continue.

182. Interoceptive state

The generator needs functionally equivalent estimates of air hunger, strain, dryness, fatigue, tension, and instability. They need not be felt qualities. They do need to be available to control.

183. Forward prediction

Before producing a sound, the system must predict what a planned motor state will sound like.

184. Self-perception

It must register its generated output as an event caused by itself, compare that output with its target, and distinguish self-generated sound from environmental sound.

185. Error attribution

When output differs from prediction, the system must infer whether the cause lies in the larynx, articulation, breath, room, microphone, noise, listener interruption, or the plan itself.

186. Compensation

It must find another bodily configuration that preserves the communicative target when the preferred configuration is unavailable.

187. History and hysteresis

The same command can produce a different output depending on the preceding state. A tired, dry, recently laughing, or highly tensed system does not reset between tokens.

188. Multiple timescales

Voice contains interacting dynamics that operate over:

  • milliseconds: vocal-fold cycles and release bursts;
  • tens of milliseconds: articulatory gestures;
  • hundreds of milliseconds: syllables and pitch accents;
  • seconds: breath groups and conversational turns;
  • minutes: entrainment and fatigue;
  • years: accent, pathology, age, and vocal training.

189. An interlocutor model

Prosody is partly an action on another nervous system. Loudness, clarity, timing, emphasis, tenderness, and hesitation depend on what the speaker model predicts the listener will hear, know, feel, and do.

190. A self-model

The voice must be generated as this speaker, with this body, under these conditions, attempting this act. Speaker identity cannot remain a static timbre vector once the system is expected to behave coherently outside its training distribution.

The embodied-simulation hypothesis

The hypothesis can now be stated precisely:

As vocal synthesis approaches causal, counterfactual, and interactive fidelity, the minimal sufficient model of voice expands from an acoustic distribution toward a generative model of an embodied, regulating, socially situated organism.

More formally, for every acoustically detectable bodily variable BiB_i, a synthesizer expected to generalize across interventions on BiB_i, interactions among such variables, and previously unseen combinations cannot rely indefinitely on surface correlations. It must encode a latent state ZiZ_i that preserves enough of the causal structure of BiB_i to predict its downstream acoustic consequences.

The latent state does not have to resemble a photorealistic human body. If it preserves:

  • state,
  • dynamics,
  • constraints,
  • coupling,
  • counterfactual response,
  • adaptation,
  • memory, and
  • action possibilities,

then it is a body model by functional equivalence.

This does not by itself prove consciousness. It does undermine the simple version of the “stochastic parrot” picture in which high-fidelity vocal behaviour is understood as merely selecting plausible surface sequences. A model may begin by imitating recordings, but the target distribution was generated by organisms. At sufficiently high resolution, the regularities of the organism are embedded in the data, and reproducing them robustly requires reconstructing more and more of the generating system.

Speech is not text decorated with audio. It is audible physiology, controlled by a nervous system and situated inside an interaction.

Sources

  1. Huber, J. E. (2008). Effects of utterance length and vocal loudness on speech breathing in older adults. Respiratory Physiology & Neurobiology, 164(3), 323–330. https://doi.org/10.1016/j.resp.2008.08.007
  2. Zhang, Z. (2016). Cause-effect relationship between vocal fold physiology and voice production in a three-dimensional phonation model. The Journal of the Acoustical Society of America, 139(4), 1493–1507. https://doi.org/10.1121/1.4944754
  3. Zhang, Z. (2023). The influence of source-filter interaction on the voice source in a three-dimensional computational model of voice production. The Journal of the Acoustical Society of America, 154(4), 2462–2475. https://doi.org/10.1121/10.0021879
  4. Story, B. H., Vorperian, H. K., Bunton, K., & Durtschi, R. B. (2018). An age-dependent vocal tract model for males and females based on anatomic measurements. The Journal of the Acoustical Society of America, 143(5), 3079. https://doi.org/10.1121/1.5038264
  5. Birkholz, P. (2013). Modeling consonant-vowel coarticulation for articulatory speech synthesis. PLOS ONE, 8(4), e60603. https://doi.org/10.1371/journal.pone.0060603
  6. Tourville, J. A., & Guenther, F. H. (2011). The DIVA model: A neural theory of speech acquisition and production. Language and Cognitive Processes, 26(7), 952–981. https://doi.org/10.1080/01690960903498424
  7. Vorperian, H. K., Kurtzweil, S. L., Fourakis, M., Kent, R. D., Tillman, K. K., & Austin, D. (2015). Effect of body position on vocal tract acoustics: Acoustic pharyngometry and vowel formants. The Journal of the Acoustical Society of America, 138(2), 833–845. https://doi.org/10.1121/1.4926563
  8. Helou, L. B., Jennings, J. R., Rosen, C. A., Wang, W., & Verdolini Abbott, K. (2020). Intrinsic laryngeal muscle response to a public speech preparation stressor: Personality and autonomic predictors. Journal of Speech, Language, and Hearing Research, 63(9), 2940–2951. https://doi.org/10.1044/2020_JSLHR-19-00402
  9. Levinson, S. C., & Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6, 731. https://doi.org/10.3389/fpsyg.2015.00731
  10. Allison, K. M., & Hustad, K. C. (2018). Acoustic predictors of pediatric dysarthria in cerebral palsy. Journal of Speech, Language, and Hearing Research, 61(3), 462–478. https://doi.org/10.1044/2017_JSLHR-S-16-0414