DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

Current perceptual metrics primarily focus on low-level pixel and texture similarities, often failing to capture meaningful mid-level structural, semantic, or pose-based relationships between images. This limitation hinders the accurate assessment of visual similarity as understood by humans, particularly regarding layout and object content. To address this gap, the authors introduce DreamSim, a novel metric designed to align more closely with human perception by evaluating images holistically. The development of DreamSim leverages synthetic data generated through text-to-image models, allowing for the creation of controlled image variations labeled with consistent human similarity judgments. Despite being trained exclusively on synthetic data, the metric demonstrates strong generalization capabilities to real-world images. It effectively prioritizes foreground objects and semantic content while remaining sensitive to color and layout, outperforming both previous learned metrics and large vision models in tasks such as nearest-neighbor retrieval and feature inversion. This work is highly relevant to open data initiatives because it establishes a new benchmark for evaluating visual similarity that bridges the gap between automated computation and human intuition. By providing a robust, open-source metric that performs well across diverse datasets like ImageNet-R and COCO, DreamSim enables more reliable evaluation of computer vision systems. This facilitates better transparency and reproducibility in research, ensuring that models are assessed on dimensions that truly reflect semantic understanding rather than superficial pixel alignment.

Source: dreamsim-nights.github.io
Published on 2024-11-25