『徳島大学 教育・研究者情報データベース (EDB)』---[学外] /
ID: Pass:

登録内容 (EID=315662)

EID=315662EID:315662, Map:0, LastModified:2022年5月5日(木) 21:15:47, Operator:[[ADMIN]], Avail:TRUE, Censor:承認済, Owner:[北岡 教英], Read:継承, Write:継承, Delete:継承.
種別 (必須): 学術論文 (審査論文) [継承]
言語 (必須): 英語 [継承]
招待 (推奨):
審査 (推奨): Peer Review [継承]
カテゴリ (推奨):
共著種別 (推奨):
学究種別 (推奨):
組織 (推奨):
著者 (必須): 1. (英) Tamura Satoshi (日) (読)
役割 (任意):
貢献度 (任意):
学籍番号 (推奨):
[継承]
2. (英) Ninomiya Hiroshi (日) (読)
役割 (任意):
貢献度 (任意):
学籍番号 (推奨):
[継承]
3.北岡 教英
役割 (任意): (英)   (日) 共同研究者として手法考案の一部を行った.   [継承]
貢献度 (任意):
学籍番号 (推奨):
[継承]
4. (英) Osuga Shin (日) (読)
役割 (任意):
貢献度 (任意):
学籍番号 (推奨):
[継承]
5. (英) Iribe Yurie (日) (読)
役割 (任意):
貢献度 (任意):
学籍番号 (推奨):
[継承]
6. (英) Takeda Kazuya (日) (読)
役割 (任意):
貢献度 (任意):
学籍番号 (推奨):
[継承]
題名 (必須): (英) Investigation of DNN-based audio-visual speech recognition  (日)    [継承]
副題 (任意):
要約 (任意): (英) <p>Audio-Visual Speech Recognition (AVSR) is one of techniques to enhance robustness of speech recognizer in noisy or real environments. On the other hand, Deep Neural Networks (DNNs) have recently attracted a lot of attentions of researchers in the speech recognition field, because we can drastically improve recognition performance by using DNNs. There are two ways to employ DNN techniques for speech recognition: a hybrid approach and a tandem approach; in the hybrid approach an emission probability on each Hidden Markov Model (HMM) state is computed using a DNN, while in the tandem approach a DNN is composed into a feature extraction scheme. In this paper, we investigate and compare several DNN-based AVSR methods to mainly clarify how we should incorporate audio and visual modalities using DNNs. We carried out recognition experiments using a corpus CENSREC-1-AV, and we discuss the results to find out the best DNN-based AVSR modeling. Then it turns out that a tandem-based method using audio Deep Bottle-Neck Features (DBNFs) and visual ones with multi-stream HMMs is the most suitable, followed by a hybrid approach and another tandem scheme using audio-visual DBNFs.</p>Audio-VisualDNN-HMMAVAV  (日) 音声と口唇画像を用いたAudio-Visual(AV)音声認識を検討している.画像を用いるため,雑音に頑健となる.近年の音声認識ではDNNをHMMと組み合わせることが精度向上につながるため,DNNによりボトルネック特報を抽出して組み合わせることとした.その結果,従来のAV音声認識よりも精度が向上することが明らかとなった.   [継承]
キーワード (推奨): 1. (英) audio-visual speech recognition (日) (読) [継承]
2. (英) deep neural network (日) (読) [継承]
3. (英) Deep Bottleneck Feature (日) (読) [継承]
4. (英) multi-stream HMM (日) (読) [継承]
発行所 (推奨): 電子情報通信学会 [継承]
誌名 (必須): 電子情報通信学会英文論文誌(D) ([電子情報通信学会])
(pISSN: 0916-8532, eISSN: 1745-1361)

ISSN (任意): 0916-8532
ISSN: 0916-8532 (pISSN: 0916-8532, eISSN: 1745-1361)
Title: IEICE transactions on information and systems
Title(ISO): IEICE Trans Inf Syst
Supplier: 一般社団法人電子情報通信学会
Publisher: Oxford University Press
 (NLM Catalog  (CiNii NCID  (J-STAGE  (Scopus  (CrossRef (Scopus information is found. [need login])
[継承]
[継承]
(必須): E99-D [継承]
(必須): 10 [継承]
(必須): [継承]
都市 (任意):
年月日 (必須): 西暦 2016年 10月 初日 (平成 28年 10月 初日) [継承]
URL (任意): http://ci.nii.ac.jp/naid/130005598220/ [継承]
DOI (任意): 10.1587/transinf.2016SLP0019    (→Scopusで検索) [継承]
PMID (任意):
CRID (任意): 1390001204379815040 [継承]
NAID : 130005598220 [継承]
Scopus (任意):
機関リポジトリ : 110586 [継承]
researchmap (任意):
評価値 (任意):
被引用数 (任意):
指導教員 (推奨):
備考 (任意):

標準的な表示

和文冊子 ● Satoshi Tamura, Hiroshi Ninomiya, Norihide Kitaoka, Shin Osuga, Yurie Iribe and Kazuya Takeda : Investigation of DNN-based audio-visual speech recognition, IEICE Transactions on Information and Systems, E99-D, 10, 2016.
欧文冊子 ● Satoshi Tamura, Hiroshi Ninomiya, Norihide Kitaoka, Shin Osuga, Yurie Iribe and Kazuya Takeda : Investigation of DNN-based audio-visual speech recognition, IEICE Transactions on Information and Systems, E99-D, 10, 2016.

関連情報

Number of session users = 2, LA = 0.58, Max(EID) = 467671, Max(EOID) = 1243063.