GitHub - FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card.

GitHub - FeSens/openTPU: An open-source AI accelerator, developed by AI: RTL, ISA, simulator, compiler and profiler in one repo. Runs Qwen3, LFM2.5 and Qwen3.5 on a Kintex-7 PCIe card.

OpenTPU demonstrates that AI agents can successfully design hardware accelerators capable of running their own inference, bridging the gap between software intent and physical implementation. By applying lessons from auto-arch-tournaments, the project validates that autonomous design processes can produce functional chip architectures that match simulator outputs bit-for-bit on real FPGA cards. This achievement highlights the viability of self-referential AI systems where the software creates the hardware infrastructure required to execute its own tasks, marking a significant step toward fully automated AI infrastructure development. The architecture emphasizes extreme transparency and educational value by housing the entire accelerator stack within a single, readable monorepo. This approach encompasses everything from the SystemVerilog hardware design and instruction set architecture to the compiler and host software, allowing developers to trace the complete data flow from high-level Python operations down to physical signal wires. This open-source nature serves as a comprehensive learning resource, demystifying the complex interactions between hardware constraints and software optimization that are often obscured in proprietary AI chip designs. This project is highly relevant to open data because it provides a fully auditable, end-to-end open-source stack for AI hardware, challenging the opacity of current commercial accelerator markets. By making the instruction set, simulator, and RTL code publicly accessible, it empowers the community to verify performance claims, fork designs, and understand the precise mechanics of AI inference efficiency. This level of visibility fosters trust and innovation in the open hardware ecosystem, ensuring that advancements in AI acceleration are not just theoretical benchmarks but empirically verifiable, community-driven achievements.

Source: github.com
Published on 2026-10-07