Improving Parquet Dedupe on Hugging Face Hub

Hugging Face is optimizing its storage architecture to handle the massive scale of datasets hosted on its Hub, with Parquet files consuming significant storage resources. The primary challenge lies in efficiently managing incremental updates to datasets, where current methods often require re-uploading entire snapshots. By implementing effective data deduplication, the team aims to store all dataset versions compactly, ensuring that space usage grows minimally even as datasets evolve, thereby reducing storage costs and improving user experience. Experiments reveal that while appending rows to Parquet files deduplicates effectively, modifying or deleting rows triggers widespread rewrites due to absolute file offsets in column headers and rigid row group structures. This results in significantly lower deduplication rates for updates, as minor changes force the rewriting of structural metadata. The findings highlight a critical tension between aggressive compression and efficient versioning, as disabling compression improves deduplication but drastically increases file sizes, creating a bottleneck for scalable data management. The proposed solution involves adopting content-defined chunking at the row group level, splitting groups based on data hashes rather than fixed counts. This approach allows the format to remain position-independent, enabling efficient deduplication even after row deletions or modifications without sacrificing compression. Implementing this requires minimal changes to Parquet writers and represents a significant step forward for open_data infrastructure, promoting more sustainable, efficient, and collaborative data storage standards within the Apache Arrow ecosystem.

Source: huggingface.co
Published on 2024-10-09