An Adaptive Structured Pruning Framework for Efficient Neural Network Inference on Resource-Constrained Edge Devices
编号:42访问权限:仅限参会人更新:2026-10-04 23:23:58浏览:7次Online
报告开始:2026年10月13日 15:15(Asia/Ho_Chi_Minh)
报告时间:15min
所在会场:[S5] Track 5: Emerging Trends of AI/ML [S5-6] Track 5: Emerging Trends of AI/ML
暂无文件
提示
无权点播视频
提示
没有权限查看文件
提示
文件转码中
摘要
Deep neural networks have been proven effective in computer vision tasks like image classification, object detection, and many other AI applications. Yet, the computational and memory requirements of these models make them hard to deploy on edge devices that have limited resources. Neural network pruning is an effective model compression technique that involves removing less important parameters or structural components of a trained model. We propose an adaptive structured pruning method that allows reduction of the complexity of a neural network without compromising too much on performance. In this method, first filters are ranked based on their contribution to the network and then low-importance ones are removed gradually. After that, a short fine-tuning period is done to get back some of the accuracy that was lost during pruning. The whole method is assessed through the model size, parameter reduction, FLOPs, inference latency, and classification accuracy metrics. Findings of this study revealed that a 50% structured pruning can significantly decrease the number of parameters and operations but still preserve an accuracy level of the original model. A comparison is provided between structured and unstructured pruning with the reasons of why structured pruning is more appropriate for actual edge hardware. It is concluded that pruning can be further combined with quantization and hardware-aware optimization methods through which more improvements can be made.
发表评论