Real-Time Performance Factors of YOLO11 on Raspberry Pi 5 under a Constrained Resource Budget: Data Protocol and Key Results
DOI:
https://doi.org/10.71942/z5rz-tr11Keywords:
performance, computer vision, Raspberry Pi 5, Hailo-8L, iot, embedding systemAbstract
An object detector for small smart cameras and mobile robots must fit within the resources of an onboard computer that it shares with other tasks. On a Raspberry Pi 5 with a budget of two inference threads, we study how the frame rate, latency, accuracy and heating of the YOLO11n detector are affected by the following factors: model format and numerical precision, input resolution, pipeline implementation language, and the Hailo-8L neural processing unit. Particular attention is paid to the data protocol. It covers forming the COCO val2017 test subset and the calibration set, aligning ground-truth annotations with detection outputs, and evaluating accuracy on small objects. A fully integer INT8 model in TensorFlow Lite with a C++ pipeline runs 3.6 times faster than the baseline Python implementation with ONNX Runtime in FP32 (17.4 vs. 4.8 fps), at a cost of 4.3 pp in mAP50–95. A 320 × 320 input gives a further 4.1-fold speedup, but mAP_S drops to 3.7 %. This drop is explained by small objects being downscaled to sizes comparable to the stride of the finest feature map. The Hailo-8L delivers 65.6 fps at 640 × 640, keeps mAP_S at the FP32 level and does not load the CPU.
Downloads
Posted
License
Copyright (c) 2026 Arxiv Academy

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.