Automatic Quality Assessment for Audio-Visual Verification Systems. The LOVe Submission to NIST SRE Challenge 2019

Grigory Antipov, Nicolas Gengembre, Olivier Le Blouch, Gaël Le Lan


Fusion of scores is a cornerstone of multimodal biometric systems composed of independent unimodal parts. In this work, we focus on quality-dependent fusion for speaker-face verification. To this end, we propose a universal model which can be trained for automatic quality assessment of both face and speaker modalities. This model estimates the quality of representations produced by unimodal systems which are then used to enhance the score-level fusion of speaker and face verification modules. We demonstrate the improvements brought by this quality-dependent fusion on the recent NIST SRE19 Audio-Visual Challenge dataset.


 DOI: 10.21437/Interspeech.2020-1434

Cite as: Antipov, G., Gengembre, N., Blouch, O.L., Lan, G.L. (2020) Automatic Quality Assessment for Audio-Visual Verification Systems. The LOVe Submission to NIST SRE Challenge 2019. Proc. Interspeech 2020, 2237-2241, DOI: 10.21437/Interspeech.2020-1434.


@inproceedings{Antipov2020,
  author={Grigory Antipov and Nicolas Gengembre and Olivier Le Blouch and Gaël Le Lan},
  title={{Automatic Quality Assessment for Audio-Visual Verification Systems. The  LOVe Submission to NIST SRE Challenge 2019}},
  year=2020,
  booktitle={Proc. Interspeech 2020},
  pages={2237--2241},
  doi={10.21437/Interspeech.2020-1434},
  url={http://dx.doi.org/10.21437/Interspeech.2020-1434}
}