{"id":1357,"date":"2021-10-13T07:50:14","date_gmt":"2021-10-13T07:50:14","guid":{"rendered":"https:\/\/sail.usc.edu:\/~ccmi\/?p=1357"},"modified":"2021-10-16T08:26:48","modified_gmt":"2021-10-16T08:26:48","slug":"active-speaker-detection","status":"publish","type":"post","link":"https:\/\/sail.usc.edu:\/ccmi\/active-speaker-detection\/","title":{"rendered":"Active speaker detection"},"content":{"rendered":"\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<h2 class=\"wp-block-heading\">Project description<\/h2>\n\n\n\n<p>&#8220;<em>When<\/em>&#8221; and &#8220;<em>where<\/em>&#8221; are the fundamental pillars of Computational media intelligence, for developing a holistic understanding of a scene which translates to locate the action of interest in the space and time. In this project, we develop a system to automatically detect the active speakers in space and time, specifically for media content.<\/p>\n<\/div><\/div>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<h2 class=\"wp-block-heading\">Methodology<\/h2>\n<\/div><\/div>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<p>This work is inspired by our preliminary work [1], where we presented a cross-modal system for the task of voice activity detection by observing just the visual frames. We performed a thorough analysis and showed that the learned embeddings can locate the human bodies and faces. The cross-modal architecture is shown below: <\/p>\n<\/div><\/div>\n\n\n\n<figure class=\"wp-block-image size-large is-resized\"><img decoding=\"async\" loading=\"lazy\" src=\"https:\/\/sail.usc.edu\/~ccmi\/wp-content\/uploads\/2021\/10\/architecture.png\" alt=\"\" class=\"wp-image-1373\" width=\"759\" height=\"290\" srcset=\"https:\/\/sail.usc.edu\/ccmi\/wp-content\/uploads\/2021\/10\/architecture.png 873w, https:\/\/sail.usc.edu\/ccmi\/wp-content\/uploads\/2021\/10\/architecture-300x115.png 300w, https:\/\/sail.usc.edu\/ccmi\/wp-content\/uploads\/2021\/10\/architecture-150x57.png 150w, https:\/\/sail.usc.edu\/ccmi\/wp-content\/uploads\/2021\/10\/architecture-768x294.png 768w\" sizes=\"(max-width: 759px) 100vw, 759px\" \/><\/figure>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<h2 class=\"wp-block-heading\">Demonstration of the system output.<\/h2>\n<\/div><\/div>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<figure class=\"wp-block-video\"><video controls src=\"https:\/\/sail.usc.edu\/~mica\/ccmi-demos\/demo_as.mp4\"><\/video><figcaption>The heatmaps represented the intermediate output of the system signifying the salient regions in the visual frames pertaining to active speakers. The active speakers are shown in green bounding boxes and all other faces in blue.<\/figcaption><\/figure>\n<\/div><\/div>\n<\/div><\/div>\n<\/div><\/div>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<h2 class=\"wp-block-heading\">References<\/h2>\n<\/div><\/div>\n\n\n\n<div class=\"wp-block-group\"><div class=\"wp-block-group__inner-container is-layout-flow wp-block-group-is-layout-flow\">\n<ul><li>R. Sharma, K. Somandepalli and S. Narayanan, &#8220;<a href=\"https:\/\/ieeexplore.ieee.org\/abstract\/document\/8803248\">Toward Visual Voice Activity Detection for Unconstrained Videos<\/a>,&#8221;&nbsp;<em>2019 IEEE International Conference on Image Processing (ICIP)<\/em>, 2019, pp. 2991-2995, doi: 10.1109\/ICIP.2019.8803248.<\/li><li>Sharma, Rahul, Krishna Somandepalli, and Shrikanth Narayanan. &#8220;<a href=\"https:\/\/arxiv.org\/pdf\/2003.04358.pdf\">Crossmodal learning for audio-visual speech event localization<\/a>.&#8221;&nbsp;<em>arXiv preprint arXiv:2003.04358<\/em>&nbsp;(2020)<\/li><\/ul>\n<\/div><\/div>\n\n\n\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Automatic visual detection of persons actively speaking on screen<\/p>\n","protected":false},"author":20,"featured_media":1390,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"inline_featured_image":false,"footnotes":""},"categories":[32,4],"tags":[],"acf":[],"_links":{"self":[{"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/posts\/1357"}],"collection":[{"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/users\/20"}],"replies":[{"embeddable":true,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/comments?post=1357"}],"version-history":[{"count":26,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/posts\/1357\/revisions"}],"predecessor-version":[{"id":1490,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/posts\/1357\/revisions\/1490"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/media\/1390"}],"wp:attachment":[{"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/media?parent=1357"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/categories?post=1357"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/sail.usc.edu:\/ccmi\/wp-json\/wp\/v2\/tags?post=1357"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}