Learning the language of molecules to predict their properties

Researchers from MIT have introduced a unified machine-learning framework that significantly improves the efficiency of predicting molecular properties and generating new molecules. By learning a "molecular grammar"—a set of production rules dictating how atomic building blocks combine—the system achieves an underlying understanding of structural similarities. This allows it to accurately predict biological and mechanical properties while simultaneously designing viable molecular structures, overcoming the common limitation of requiring massive, expensive training datasets. This approach demonstrates remarkable data efficiency, outperforming existing deep-learning models even when trained on fewer than one hundred samples. Traditional methods often rely on costly pretraining on general datasets or require millions of labeled examples, leading to poor performance on specific domains. In contrast, this hierarchical grammar-based method avoids extensive pretraining and leverages reinforcement learning to quickly adapt to small, domain-specific datasets, achieving results comparable to methods trained on much larger corpora. This research is highly relevant to open_data as it challenges the prevailing assumption that high-performance AI requires massive, curated datasets. It illustrates how structured, interpretable knowledge representations can democratize access to advanced scientific modeling by lowering the barrier to entry regarding data volume. Furthermore, the potential for this grammar-based representation to apply to other graph-form data highlights the broader utility of open, efficient machine-learning frameworks in accelerating discovery across various scientific fields beyond just chemistry.

Source: sciencedaily.com
Published on 2023-07-09