Anda belum login :: 06 Jul 2025 16:04 WIB
Detail
ArtikelStatistical Multimodal Integration for Audio-Visual Speech Processing (Invited Paper)  
Oleh: Nakamura, S.
Jenis: Article from Journal - ilmiah internasional
Dalam koleksi: IEEE Transactions on Neural Networks vol. 13 no. 4 (2002), page 854-866.
Topik: audio visual; statistical multimodal; integration; audio - visual speech
Ketersediaan
  • Perpustakaan Pusat (Semanggi)
    • Nomor Panggil: II36.7A
    • Non-tandon: 1 (dapat dipinjam: 0)
    • Tandon: tidak ada
    Lihat Detail Induk
Isi artikelSensory information is indispensable for living things. It is also important for living things to integrate multiple types of senses to understand their surroundings. In human communications, human beings must further integrate the multimodal senses of audition and vision to understand intention. In this paper, we describe speech related modalities since speech is the most important media to transmit human intention. To date, there have been a lot of studies concerning technologies in speech communications, but performance levels still have room for improvement. For instance, although speech recognition has achieved remarkable progress, the speech recognition performance still seriously degrades in acoustically adverse environments. On the other hand, perceptual research has proved the existence of the complementary integration of audio speech and visual face movements in human perception mechanisms. Such research has stimulated attempts to apply visual face information to speech recognition and synthesis. This paper introduces works on audio - visual speech recognition, speech to lip movement mapping for audio - visual speech synthesis, and audio - visual speech translation.
Opini AndaKlik untuk menuliskan opini Anda tentang koleksi ini!

Kembali
design
 
Process time: 0 second(s)