Why AI can’t do hiring

The central argument challenges the viability of using AI to automate hiring decisions, asserting that the primary obstacle is not technological capability but the severe lack of high-quality, proprietary data. While AI models have advanced significantly, effective recruitment requires training data that reflects true job performance and candidate suitability, which remains inaccessible in the public domain. Consequently, attempts to build AI matchers using only open-source information are fundamentally flawed, as the necessary signals to predict hiring success simply do not exist in publicly available formats. Publicly accessible data sources like LinkedIn profiles, GitHub repositories, and social graphs provide insufficient signal for accurate candidate assessment. Resume-based information often correlates weakly with actual engineering skill, while code repositories represent only a tiny fraction of the developer workforce, as most professionals do not maintain public activity. Furthermore, social networks fail to reliably map talent due to informal networking behaviors that do not strictly reflect professional competence. Without proprietary data on interview performance and long-term job outcomes, algorithms cannot distinguish between capable candidates and those who merely present a favorable public profile. This perspective is highly relevant to the open data community because it highlights the limitations of public datasets in solving complex, high-stakes problems like employment matching. It serves as a cautionary tale against the assumption that open data alone is sufficient for building robust AI systems, emphasizing that domain-specific, proprietary insights are often essential for meaningful application. Understanding these data gaps helps clarify where open data initiatives can realistically contribute to AI development and where they fall short, guiding more practical expectations for innovation in hiring technology.

Source: interviewing.io
Published on 2023-05-18