Abstract

In recent years, multi-agent sensing technology has attracted considerable attention, particularly in the field of vehicle-to-vehicle (V2V) collaborative perception for autonomous driving. By sharing complementary information, connected autonomous vehicles (CAVs) can significantly improve scene understanding. However, this collaboration also presents challenges, such as balancing data transmission efficiency with the accuracy of environmental perception. To address these issues, we propose an autoencoder based on intermediate-level feature fusion. Specifically, we design a 3D convolutional autoencoder with a residual structure to maximize the compression of inter-vehicle bandwidth usage while retaining essential feature information. By encoding the spatial-channel dimensions of intermediate feature maps extracted from point clouds, network bandwidth is reduced by 3–10× under different configurations. Additionally, we propose a channel-wise autoencoder training strategy to reduce model parameters and enhance feature compression efficiency, while preserving collaborative perception performance. Finally, we conduct a comprehensive evaluation of perception accuracy and bandwidth consumption using a collaborative detection dataset collected via LiDAR. Extensive experimental results demonstrate that our method offers significant advantages in balancing object detection accuracy and bandwidth compression for CAVs.