Achieving Sub-Frame Precision Speech Segmentation Through Computational Multimodal Analysis: Bridging Neuroscience Requirements with Engineering Solutions

Poplin, M., Hughes, M., McCrory, B., & Modyanova, N. (2026). Achieving Sub-Frame Precision Speech Segmentation Through Computational Multimodal Analysis: Bridging Neuroscience Requirements with Engineering Solutions. 2026 International Conference on Artificial Intelligence, Computer, Data Sciences and Applications (ACDSA), 1-6. https://doi.org/10.1109/acdsa67686.2026.11468095

Publication date: 5 Feb 2026 Added to AutiHub: 9 Aug 2026 Type: Other Article language: English

This publication is integrated into AutiHub through:

Authors

Publication authors
4
Publication authors identified as autistic
1 / 4 (25.0%)

Abstract

Recent neuroscience research reveals that audiovisual speech integration occurs within specific temporal windows of$100-300 \text{ms}$, with lip movements entraining brain oscillations at$\mathbf{2 - 7 H z}$corresponding to syllabic rate. Studying these phenomena, particularly in clinical populations where timing differences may serve as biomarkers, requires precise temporal control that current video processing tools cannot provide. We present a novel computational approach that combines MediaPipe-based facial landmark detection as a visual trigger with LibROSA-based audio-onset validation, achieving sub-frame temporal precision through audio onset detection at 5 ms intervals. Unlike previous neuroscience studies that used manually prepared single-word stimuli, our system processes continuous discourse while maintaining the temporal precision required for EEG/fNIRS experiments. The visual detection stage identifies mouth movements using 468 facial landmarks, triggering targeted audio analysis within a neurobiologically-grounded 300 ms search window. Evaluation on presentation videos demonstrates 92-100% segment extraction accuracy with processing speeds of$0.556 \times$real-time. The current implementation targets single-speaker, frontal-face recordings with controlled audio quality. This enables large-scale investigation of differences in audiovisual integration, which is particularly relevant to research on autism spectrum disorder, where altered temporal binding may characterize the condition.

Bibliography cited by this reference

Cited references are imported from external metadata sources when they are available. The list may be partial.

Bibliography inclusion overview

These indicators describe the cited bibliography imported for this publication. Cited-reference metrics use the cited-reference total as denominator. Cited-author metrics state whether they use all cited-author occurrences or only occurrences linked to authors already integrated in the AutiHub database. They use cached links between cited authors and authors integrated in the AutiHub database. Last computed: 16 Aug 2026 11:31.

Cited references
13
Total cited references integrated for this publication.
Cited references with an identified autistic author
0 / 13 (0.0%)
Cited author occurrences identified as autistic
0 / 58 (0.0%)
Among occurrences linked to AutiHub author records: 0 / 0 (0.0%). Distinct cited authors identified as autistic: 0 / 58 (0.0%).
Cited author occurrences linked to AutiHub author records
0 / 58 (0.0%)
Distinct linked cited authors: 0 / 58 (0.0%)
Linked cited-author occurrences not identified as autistic
0 / 0 (0.0%)
Among linked cited-author occurrences only. Across all cited-author occurrences: 0 / 58 (0.0%). Distinct linked cited authors not identified as autistic: 0 / 0 (0.0%).
  1. HARRY MCGURK , JOHN MACDONALD (1976). Hearing lips and seeing voices . Nature, 264(5588), 746-748. Springer Science and Business Media LLC.
    Type: Article DOI: 10.1038/264746a0 OpenAlex: https://openalex.org/W2015394094
    Crossref OpenAlex
  2. Chandramouli Chandrasekaran , Andrea Trubanova , Sébastien Stillittano , Alice Caplier , Asif A. Ghazanfar (2009). The Natural Statistics of Audiovisual Speech . PLoS Computational Biology, 5(7), e1000436. Public Library of Science (PLoS).
    Crossref OpenAlex
  3. Ayaz A. Shaikh , Dinesh K. Kumar , Wai C. Yau , M. Z. Che Azemin , Jayavardhana Gubbi (2010). Lip reading using optical flow and support vector machines . 2010 3rd International Congress on Image and Signal Processing, 327-330. IEEE.
    Crossref OpenAlex
  4. Hyojin Park , Christoph Kayser , Gregor Thut , Joachim Gross (2016). Lip movements entrain the observers’ low-frequency brain oscillations to facilitate speech intelligibility . eLife, 5. eLife Sciences Publications, Ltd.
    Crossref OpenAlex
  5. Patrick Reisinger , Marlies Gillis , Nina Suess , Jonas Vanthornhout , Chandra Leon Haider , Thomas Hartmann et al. (2025). Neural Speech Tracking Contribution of Lip Movements Predicts Behavioral Deterioration When the Speaker's Mouth Is Occluded . eneuro, 12(2), ENEURO.0368-24.2024. Society for Neuroscience.
    Crossref OpenAlex
  6. Jongseo Sohn , Nam Soo Kim , Wonyong Sung (1999). A statistical model-based voice activity detection . IEEE Signal Processing Letters, 6(1), 1-3. Institute of Electrical and Electronics Engineers (IEEE).
    Crossref OpenAlex
  7. Kazuki Sekine , Christina Schoechl , Kimberley Mulder , Judith Holler , Spencer Kelly , Reyhan Furman et al. (2020). Evidence for children’s online integration of simultaneous information from speech and iconic gestures: an ERP study . Language, Cognition and Neuroscience, 35(10), 1283-1294. Informa UK Limited.
    Crossref OpenAlex
  8. Brian McFee , Colin Raffel , Dawen Liang , Daniel Ellis , Matt McVicar , Eric Battenberg et al. (2015). librosa: Audio and Music Signal Analysis in Python . Proceedings of the Python in Science Conference, 18-24. SciPy.
    Crossref OpenAlex
  9. Kartynnik (2019). Real-time facial surface geometry from monocular video on mobile GPUs . CVPR Workshop on Computer Vision for Augmented and Virtual Reality.
    Type: Other
    Crossref
  10. Frédéric Roux , Blair C. Armstrong , Manuel Carreiras (2017). Chronset: An automated tool for detecting speech onset . Behavior Research Methods, 49(5), 1864-1881. Springer Science and Business Media LLC.
    Crossref OpenAlex
  11. Arnab Dey , Kousik Dasgupta (2022). Emotion Recognition Using Deep Learning in Pandemic with Real-time Email Alert . In Lecture Notes in Electrical Engineering (pp. 175-190). Springer Singapore.
    Crossref OpenAlex
  12. Diego Resende Faria , Abraham Itzhak Weinberg , Pedro Paulo Ayrosa (2024). Multimodal Affective Communication Analysis: Fusing Speech Emotion and Text Sentiment Using Machine Learning . Applied Sciences, 14(15), 6631. MDPI AG.
    Crossref OpenAlex
  13. Stavros Petridis , Themos Stafylakis , Pingehuan Ma , Feipeng Cai , Georgios Tzimiropoulos , Maja Pantic (2018). End-to-End Audiovisual Speech Recognition . 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 6548-6552. IEEE.
    Crossref