Think Twice: a Two-Level Mismatch Discriminator for Generalized Referring Expression Grounding

Abstract

Generalized Referring Expression Grounding (GREG) introduces a novel task that encompasses both no-target and multi-target samples, distinguishing it from the traditional task. Current methods can be categorized into object-level and global-level discrimination based on how they handle no-target situations. However, global-level methods show inconsistency between mask and discrimination results, while object-level methods lack a global perspective to explore mismatches. To address this, we propose the Think Twice model, which features a unique Two-level Mismatch Discriminator that performs discrimination at both object-level and global-level. Object-level ensures that the model identifies high-scoring objects within the target samples, thereby maintaining consistency. Global-level incorporates the positional prompt of objects and leverages cross attention to capture more detailed global relationships to enhance discrimination accuracy. Additionally, our model proposes a Ground-Truth Mixed-Prompt Generation algorithm to enhance the prompt stability in the training stage. Experimental findings demonstrate the superior performance of our Think Twice model in detecting mismatches, achieving a new state-of-the-art on the gRefCOCO and Ref-ZOM datasets.

Publication
MM 26: Proceedings of the 34th ACM International Conference on Multimedia
Shangfei Wang
Shangfei Wang
Professor of Artificial Intelligence

My research interests include Pattern Recognition, Affective Computing, Probabilistic Graphical Models, Computation Intelligence.