Logotipo do repositório

Interactive Content Retrieval in Egocentric Videos Based on Vague Semantic Queries

Carregando...
Imagem de Miniatura

Orientador

Coorientador

Pós-graduação

Curso de graduação

Título da Revista

ISSN da Revista

Título de Volume

Editor

MDPI

Tipo

Artigo

Direito de acesso

Acesso abertoAcesso Aberto

Resumo

Retrieving specific, often instantaneous, content from hours-long egocentric video footage based on hazily remembered details is challenging. Vision–language models (VLMs) have been employed to enable zero-shot textual-based content retrieval from videos. But, they fall short if the textual query contains ambiguous terms or users fail to specify their queries enough, leading to vague semantic queries. Such queries can refer to several different video moments, not all of which can be relevant, making pinpointing content harder. We investigate the requirements for an egocentric video content retrieval framework that helps users handle vague queries. First, we narrow down vague query formulation factors and limit them to ambiguity and incompleteness. Second, we propose a zero-shot, user-centered video content retrieval framework that leverages a VLM to provide video data and query representations that users can incrementally combine to refine queries. Third, we compare our proposed framework to a baseline video player and analyze user strategies for answering vague video content retrieval scenarios in an experimental study. We report that both frameworks perform similarly, users favor our proposed framework, and, as far as navigation strategies go, users value classic interactions when initiating their search and rely on the abstract semantic video representation to refine their resulting moments.

Descrição

Palavras-chave

Citação

Itens relacionados

Financiadores

Unidades

Tipo de item:Unidade,
Presidente Prudente, Faculdade de Ciências e Tecnologia - FCT
FCT
Campus: Presidente Prudente

Departamentos

Cursos de graduação

Programas de pós-graduação

Outras formas de acesso