LeoLM: Igniting German-Language LLM Research | LAION

LeoLM represents a significant milestone in open data by delivering the first comprehensive suite of German-language foundation models that challenge the English-centric dominance of current large language models. By continuing pretraining on Llama-2 with a high-quality, localized German corpus, the project demonstrates that existing models can be effectively adapted to new languages without suffering from catastrophic forgetting of their original capabilities. This approach not only preserves global knowledge but also significantly enhances proficiency in German, proving that multilingual expansion is technically viable and efficient. The initiative also addresses the lack of standardized evaluation for non-English models by introducing GermanBench, a collection of translated benchmarks that allow for rigorous and comparable performance assessment. This contribution provides the research community with essential tools to measure progress accurately, fostering a more transparent and data-driven ecosystem. By making both the models and the evaluation datasets openly available, LeoLM encourages reproducibility and sets a precedent for how language-specific advancements can be tracked and validated in open science. This release is crucial for open data because it reduces reliance on closed-source commercial entities by providing a powerful, accessible alternative for German AI applications. It empowers researchers and developers within the German-speaking community to build upon existing open weights, promoting innovation and collaboration. Ultimately, LeoLM serves as a proof-of-concept for language acquisition in pretrained models, accelerating the development of diverse, multilingual open-source AI that is equitable and inclusive.

Source: laion.ai
Published on 2023-09-30