Francesco Cutugno - So many units for speech analysis, and yet...

  • Date: 02 OCTOBER 2026  from 15:00 to 17:00

  • Event location: Aula Seminari 2, Piazza San Giovanni in Monte 2, 40124 Bologna, Italia - In presence and online event

  • Type: Seminar

Human speech unfolds as a single continuous acoustic stream, yet listeners parse it into discrete linguistic units that operate at markedly different temporal scales. Converging evidence from psycholinguistics and cognitive neuroscience indicates that phonetic segments are resolved within windows of roughly 100 milliseconds, syllables within about 250-300 milliseconds, words within 800-900 milliseconds, and larger prosodic, or "tonal," units — closer to intonational phrases — within a window approaching two seconds.

Syllabic timing appears to play a privileged role among these grains, since its quasi-periodic structure underlies the perceived rhythmic scansion of spoken language and may act as a temporal scaffold onto which finer, segmental, and coarser, lexical and prosodic, information is anchored.

These timescales are not processed by a single mechanism: distinct neural populations, and regions of the auditory and language networks appear to track and integrate information at each grain, drawing on multimodal cues, acoustic, visual, gestural, and on multidimensional representations, phonological, semantic, pragmatic, and affective.

What remains poorly understood is how the mind binds these parallel, multi-scale analyses into the unitary, coherent experience of understanding and producing speech in real time. The speech signal itself, as a single undifferentiated stream, simultaneously carries acoustic detail, semantic content, pragmatic intent, and emotional coloring, and current models disagree on whether these strands are extracted through a shared hierarchical architecture or through partially independent computations merged only later.

This seminar reviews the empirical evidence for a hierarchy of minimal analysis units in speech, from the phone to the intonational unit, and situates it within the ongoing theoretical debate on the real nature of segmental information, syllable-based rhythm, and the neural integration of multiscale linguistic information. Rather than offering a settled account, the seminar aims to map the current state of disagreement and identify the questions that most urgently call for new experimental and computational work.