Helping computer vision and language models understand what they see

Vision and language models typically excel at identifying objects but often fail to understand spatial relationships and object attributes, such as understanding that a cup is sitting on a table. This limitation arises because standard training methods prioritize contrastive learning, which causes models to focus primarily on nouns while ignoring the semantic context of how items are arranged in a scene. Consequently, these systems struggle with complex concept understanding despite their proficiency in basic object recognition. To address this, researchers developed a technique using computer-generated synthetic data to fine-tune existing models. By creating a diverse dataset of photorealistic images with detailed annotations describing object positions and interactions, the team was able to teach models to recognize semantic concepts beyond simple identification. This approach leverages the cost-effectiveness, privacy, and scalability of synthetic data to expose models to a wider variety of scenarios than natural datasets typically provide, significantly enhancing their ability to interpret complex visual arrangements. This advancement is highly relevant to the open_data community because it demonstrates how synthetic data can effectively supplement real-world data to overcome specific model limitations without requiring additional labeled real images. The technique improves accuracy by up to ten percent while preventing the model from forgetting prior knowledge, offering a practical pathway for enhancing open-source AI systems. Such improvements benefit applications ranging from automated video captioning to healthcare and e-commerce, promoting more robust and semantically aware open artificial intelligence tools.

Source: news.mit.edu
Published on 2023-09-14