- https://doi.org/10.1109/inc465408.2025.11256254
Comparative Analysis of Torchvision Object Detection Models
- Mar 14, 2025
- Naga Madhurya Peram +3 more
Object detection sits at the core of contemporary computer vision, powering everything from security cameras to driverless cars. In this work, we put Torchvision’s latest detection models under the microscope, while Comet’s rocksolid experiment-tracking keeps score. Drawing on the finely annotated Penn-Fudan Pedestrian dataset, we guide the reader through every step-weight initialization, data wrangling, training routines, and evaluation protocols-examining each model’s knack for adaptation and generalization. A battery of trials teases apart the strengths and shortcomings of marquee architectures such as Mask RCNN, Fast RCNN, Faster RCNN, and RetinaNet. Key yardsticks-mean Average Precision, mean Average Recall, and the trusty F1 score-are logged and visualized in painstaking detail with Comet, revealing each network’s quirks and capabilities. The discussion then turns pragmatic, outlining best practices for deployment and lifecycle management and leaning on Comet’s artifact logging to streamline the hand-off from lab to production. Altogether, the study offers a clear, end-to-end framework that equips researchers and engineers to navigate the real-world maze of object-detection tasks with confidence.