End-to-End Deep Lip-reading: A Preliminary Study

Masters Thesis


Thapa, K. (2023). End-to-End Deep Lip-reading: A Preliminary Study. Masters Thesis London South Bank University School of Engineering https://doi.org/10.18744/lsbu.97xz0
AuthorsThapa, K.
TypeMasters Thesis
Abstract

Deep lip-reading is the use of deep neural networks to extract speech from silent videos. Most works in lip-reading use a multi staged training approach due to the complex nature of the task. A single stage, end-to-end, unified training approach, which is an ideal of machine learning, is also the goal in lip-reading. However, pure end-to-end systems have so far failed to perform as good as non-end-to-end systems. Some exceptions to this are the very recent Temporal Convolutional Network (TCN) based architectures (Martinez et al., 2020; Martinez et al., 2021). This work lays out preliminary study of deep lip-reading, with a special focus on various end-to-end approaches. The research aims to test whether a purely end-to-end approach is justifiable for a task as complex as deep lip-reading. To achieve this, the meaning of pure end-to-end is first defined and several lip-reading systems that follow the definition are analysed. The system that most closely matches the definition is then adapted for pure end-to-end experiments. We make four main contributions: i) An analysis of 9 different end-to-end deep lip-reading systems, ii) Creation and public release of a pipeline to adapt sentence level Lipreading Sentences in the Wild 3 (LRS3) dataset into word level, iii) Pure end-to-end training of a TCN based network and evaluation on LRS3 word-level dataset as a proof of concept, iv) a public online portal to analyse visemes and experiment live end-to-end lip-reading inference. The study is able to verify that pure end-to-end is a sensible approach and an achievable goal for deep machine lip-reading.

Keywordsend-to-end lip-reading ; deep lip-reading
Year2023
PublisherLondon South Bank University
Digital Object Identifier (DOI)https://doi.org/10.18744/lsbu.97xz0
File
License
File Access Level
Open
Publication dates
Print04 Jan 2023
Publication process dates
Deposited08 Aug 2024
Funder/ClientLondon South Bank University
Permalink -

https://openresearch.lsbu.ac.uk/item/97xz0

Download files


File
MRes_Thesis_Revised.pdf
License: CC BY 4.0
File access level: Open

  • 5
    total views
  • 0
    total downloads
  • 0
    views this month
  • 0
    downloads this month

Export as

Related outputs

End-to-end Lip-reading: A Preliminary Study
Thapa, K. (2023). End-to-end Lip-reading: A Preliminary Study. Masters Thesis London South Bank University School of Engineering https://doi.org/10.18744/lsbu.92zq5