AI4Bharat releases Hindi LLM ‘Airavata’ - ET Telecom
AI4Bharat has released Airavata, an open-source large language model tailored for Hindi, marking a significant milestone for Indian linguistic AI. Built upon SarvamAI’s OpenHathi and enhanced with diverse instruction-tuning datasets, the model demonstrates encouraging performance compared to other open-source alternatives. This release serves as a foundational step toward developing high-quality, open-source LLMs for all scheduled Indic languages, addressing the critical scarcity of local training data. The initiative emphasizes sustainability and accessibility by utilizing human-curated, license-friendly datasets rather than relying on distilled data from proprietary models. This approach avoids costly licensing restrictions, ensuring that downstream applications remain free to use. By sharing both the model and its instruction-tuning datasets, the project aims to empower broader research and development within the Indic LLM community, fostering a more inclusive and accessible ecosystem for multilingual AI. However, the team acknowledges inherent limitations, including risks of hallucination, bias, and struggles with complex topics or cultural subtleties. Despite these challenges, Airavata represents a crucial advancement in open_data efforts for non-English languages. It highlights the importance of transparent, community-driven data curation in building equitable AI systems that respect local contexts while remaining openly available for global innovation and adaptation.
Source: telecom.economictimes.indiatimes.comPublished on 2024-01-26
Related news
- Hugging Face and Google partner for open AI collaboration
- Google Cloud, Hugging Face join hands to accelerate Generative AI, ML Development
- Google Cloud, Hugging Face Partner on AI Development
- Google Cloud and Hugging Face ink AI infrastructure partnership
- Google Cloud partners with Hugging Face to attract AI developers