Open Access System for Information Sharing

Graduate School of Artificial Intelligence (인공지능대학원) 2. Conference Papers

Conference

Cited 0 time in webofscience

Cited 0 time in scopus

Metadata Downloads

Full metadata record

Files in This Item:: There are no files associated with this item.

DC Field	Value	Language
dc.contributor.author	김동원	-
dc.contributor.author	김남엽	-
dc.contributor.author	곽수하	-
dc.date.accessioned	2024-03-07T00:23:57Z	-
dc.date.available	2024-03-07T00:23:57Z	-
dc.date.created	2024-03-06	-
dc.date.issued	2023-06	-
dc.identifier.uri	https://oasis.postech.ac.kr/handle/2014.oak/122809	-
dc.description.abstract	Cross-modal retrieval across image and text modalities is a challenging task due to its inherent ambiguity: An image often exhibits various situations, and a caption can be coupled with diverse images. Set-based embedding has been studied as a solution to this problem. It seeks to encode a sample into a set of different embedding vectors that capture different semantics of the sample. In this paper, we present a novel set-based embedding method, which is distinct from previous work in two aspects. First, we present a new similarity function called smooth-Chamfer similarity, which is designed to alleviate the side effects of existing similarity functions for set-based embedding. Second, we propose a novel set prediction module to produce a set of embedding vectors that effectively captures diverse semantics of input by the slot attention mechanism. Our method is evaluated on the COCO and Flickr30K datasets across different visual backbones, where it outperforms existing methods including ones that demand substantially larger computation at inference.	-
dc.language	English	-
dc.publisher	IEEE Computer Society	-
dc.relation.isPartOf	2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023	-
dc.relation.isPartOf	Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition	-
dc.title	Improving Cross-Modal Retrieval with Set of Diverse Embeddings	-
dc.type	Conference	-
dc.type.rims	CONF	-
dc.identifier.bibliographicCitation	2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, pp.23422 - 23431	-
dc.citation.conferenceDate	2023-06-18	-
dc.citation.conferencePlace	CA	-
dc.citation.endPage	23431	-
dc.citation.startPage	23422	-
dc.citation.title	2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023	-
dc.contributor.affiliatedAuthor	김동원	-
dc.contributor.affiliatedAuthor	김남엽	-
dc.contributor.affiliatedAuthor	곽수하	-
dc.description.journalClass	1	-
dc.description.journalClass	1	-

Show simple item record

qr_code

트윗하기

Communities & Collection

Graduate School of Artificial Intelligence (인공지능대학원)

Related Researcher

Researcher

곽수하KWAK, SU HA: Grad. School of AI

Read more

Open Access System for Information Sharing

Communities & Collection

Related Researcher

Views & Downloads

Browse