Research on the Application of Hardware Accelerators in Artificial Intelligence Systems
DOI:
https://doi.org/10.61173/fxcbk297Keywords:
Hardware Accelerators, Artificial Intelli-gence Systems, Applied ResearchAbstract
Traditional general-purpose processors are constrained by operational efficiency when handling computationally intensive tasks such as deep learning, convolutional neural networks, and recurrent neural networks. Hardware accelerators, as specialized computing architectures, significantly outperform general-purpose processors in parallel computing, data throughput, and operational efficiency. Therefore, hardware accelerators have become one of the key factors driving the advancement of AI. This article systematically studies the basic principles of hardware accelerators, their main types (GPU, FPGA, ASIC), as well as the design and implementation methods of hardware accelerators. By comparing the performance metrics, power consumption characteristics and application adaptability of different accelerator architectures, the advantages of hardware accelerators over other accelerators are demonstrated. The paper also investigates the advantages and challenges of hardware accelerators in accelerating deep learning inference, training, and edge computing. According to the research results, accelerators tailored for specific AI tasks can significantly reduce latency and improve energy efficiency, and have broad application prospects in scenarios such as 5G, autonomous driving, and intelligent manufacturing. This research provides a reference for the designers of artificial intelligence systems to select and optimize hardware acceleration solutions, and also offers a direction for the future innovation of accelerator architectures.
References
[1] Dally W J. High-performance hardware for machine learning[J]. Advances in Computers, 2020, 117: 1-70.
[2] Chen Y, Emer J, Sze V. Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks[J]. IEEE Journal of Solid-State Circuits, 2017, 52(1): 127-138.
[3] Zhu J, Zhang W, Liu S, et al. Energy-efficient hardware accelerators for machine learning in embedded systems[J]. Electronics, 2022, 11(6): 945.
[4] Li Z, Zhang C, Chen X, et al. Hardware acceleration for deep neural networks: A survey[J]. Wireless Personal Communications, 2024, 133: 2665-2691.
[5] Sze V, Chen Y H, Yang T J, et al. Efficient processing of deep neural networks: A tutorial and survey[J]. Proceedings of the IEEE, 2017, 105(12): 2295-2329.
[6] Jouppi N P, Young C, Patil N, et al. In-datacenter performance analysis of a tensor processing unit[C]//Proceedings of the 44th Annual International Symposium on Computer Architecture. ACM, 2017: 1-12.
[7] Chen T, Du Z, Sun N, et al. DianNao: A small-footprint high-throughput accelerator for ubiquitous machinelearning[C]//Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operating Systems. ACM, 2014: 269-284.
[8] Jouppi N P, Yoon D H, Kurian G, et al. TPUv4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings[C]//Proceedings of the 50th Annual International Symposium on Computer Architecture. ACM, 2023: 1-14.
[9] Chen Zhilu, et al. „A fast deep learning system using GPU.“ 2014 IEEE international symposium on circuits and systems (ISCAS). IEEE, 2014.
[10] G. Kim, M. Lee, J. Jeong and J. Kim, „Multi-GPU System Design with Memory Networks,“ 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, Cambridge, UK, 2014, pp. 484-495,
[11] Premalatha, S., et al. „FPGA Based AI Inference Accelerator for Low-End Embedded Systems.“ 2025 3rd International Conference on Artificial Intelligence and Machine Learning Applications Theme: Healthcare and Internet of Things (AIMLA). IEEE, 2025.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
