Memory-centric accelerator design for Convolutional Neural Networks

Top Cited Papers

1 October 2013

conference paper
Published by Institute of Electrical and Electronics Engineers (IEEE)

No. 10636404,p. 13-19
https://doi.org/10.1109/iccd.2013.6657019

Abstract

In the near future, cameras will be used everywhere as flexible sensors for numerous applications. For mobility and privacy reasons, the required image processing should be local on embedded computer platforms with performance requirements and energy constraints. Dedicated acceleration of Convolutional Neural Networks (CNN) can achieve these targets with enough flexibility to perform multiple vision tasks. A challenging problem for the design of efficient accelerators is the limited amount of external memory bandwidth. We show that the effects of the memory bottleneck can be reduced by a flexible memory hierarchy that supports the complex data access patterns in CNN workload. The efficiency of the on-chip memories is maximized by our scheduler that uses tiling to optimize for data locality. Our design flow ensures that on-chip memory size is minimized, which reduces area and energy usage. The design flow is evaluated by a High Level Synthesis implementation on a Virtex 6 FPGA board. Compared to accelerators with standard scratchpad memories the FPGA resources can be reduced up to 13× while maintaining the same performance. Alternatively, when the same amount of FPGA resources is used our accelerators are up to 11× faster.

Keywords

This publication has 13 references indexed in Scilit:

Project Glass: An Extension of the Self
IEEE Pervasive Computing, 2013
Polyhedral-based data reuse optimization for configurable computing
Published by Association for Computing Machinery (ACM) ,2013
NeuFlow: A runtime reconfigurable dataflow processor for vision
Published by Institute of Electrical and Electronics Engineers (IEEE) ,2011
A dynamically configurable coprocessor for convolutional neural networks
Published by Association for Computing Machinery (ACM) ,2010
Refactoring for Data Locality
Computer, 2009
A practical automatic polyhedral parallelizer and locality optimizer
ACM SIGPLAN Notices, 2008
Memory-centric video processing
IEEE Transactions on Circuits and Systems for Video Technology, 2008
Traffic monitoring and accident detection at intersections
IEEE Transactions on Intelligent Transportation Systems, 2000
Gradient-based learning applied to document recognition
Proceedings of the IEEE, 1998
A data locality optimizing algorithm
Published by Association for Computing Machinery (ACM) ,1991