Perpustakaan Unika Atma Jaya

Anda belum login :: 17 May 2026 06:52 WIB

Home

Logon

» »

Detail

Statistical Multimodal Integration for Audio-Visual Speech Processing (Invited Paper)

Oleh:

Nakamura, S.

Jenis: Article from Journal - ilmiah internasional
Dalam koleksi: IEEE Transactions on Neural Networks vol. 13 no. 4 (2002), page 854-866.
Topik: audio visual; statistical multimodal; integration; audio - visual speech

Ketersediaan

Perpustakaan Pusat (Semanggi)
- Nomor Panggil: II36.7A
- Non-tandon: 1 (dapat dipinjam: 0)
- Tandon: tidak ada

Lihat Detail Induk

Isi artikelSensory information is indispensable for living things. It is also important for living things to integrate multiple types of senses to understand their surroundings. In human communications, human beings must further integrate the multimodal senses of audition and vision to understand intention. In this paper, we describe speech related modalities since speech is the most important media to transmit human intention. To date, there have been a lot of studies concerning technologies in speech communications, but performance levels still have room for improvement. For instance, although speech recognition has achieved remarkable progress, the speech recognition performance still seriously degrades in acoustically adverse environments. On the other hand, perceptual research has proved the existence of the complementary integration of audio speech and visual face movements in human perception mechanisms. Such research has stimulated attempts to apply visual face information to speech recognition and synthesis. This paper introduces works on audio - visual speech recognition, speech to lip movement mapping for audio - visual speech synthesis, and audio - visual speech translation.

Opini AndaKlik untuk menuliskan opini Anda tentang koleksi ini!

Kembali

Process time: 0.015625 second(s)