If you found this work to be useful in your own research, please consider citing the following:
Multi-Modal Cooperative Autonomy Benchmark
with C-V2X Communication Characterization
CooperScene is the first real-world, multi-agent, multi-modal cooperative autonomy dataset with C-V2X communication characterization. It is organized into diverse real-world scenes—spanning urban intersection, freeway ramp, local street, and parking lot—that introduce different traffic and occlusion patterns. In urban intersection scnario, our data involves three connected autonomous vehicles and one instrumented infrastructure roadside unit, all equipped with multi-modal sensors and commercial C-V2X communication radios.
The dataset provides tightly aligned LiDAR, camera, GNSS/IMU, and C-V2X network data, enabling the community to jointly evaluate perception performance and the practical constraints under real-world vehicular networks. CooperScene supports 3D object detection and motion prediction benchmarks with synchronized pair-wise C-V2X throughput measurements.
High-quality interactive scenes from three vehicles and one infrastructure agent across diverse real-world urban intersection scenarios.
Real-world C-V2X network traces including latency, throughput, packet loss rate and jitter — bridging wireless networking and autonomous driving research.
LiDAR, camera, GNSS/IMU data streams per agent, well-calibrated and temporally synchronized with sub-millisecond accuracy.
GNSS-RTK positioning with spatial-temporal ICP alignment achieving an average RMSE of 0.2m across all agents.
PTP-based time synchronization and hardware triggering ensure globally consistent timestamps across all platforms and sensors.
Comprehensive benchmarks for detection and motion prediction evaluating both perception accuracy and communication efficiency.
| Model | Agents | mAP (C-V2X) | mAP (Unlimited) | Sharing Size (MB) |
Latency (ms) under C-V2X |
||||
|---|---|---|---|---|---|---|---|---|---|
| @0.3 | @0.5 | @0.7 | @0.3 | @0.5 | @0.7 | ||||
| V2VNet | V+I | 0.43 | 0.33 | 0.20 | 0.55 | 0.42 | 0.23 | 10.0 | 64,349 |
| V+V | 0.35 | 0.25 | 0.15 | 0.66 | 0.55 | 0.34 | 10.0 | 64,038 | |
| V+V+I | 0.36 | 0.27 | 0.15 | 0.70 | 0.60 | 0.37 | 20.0 | 76,879 | |
| V+2V | 0.32 | 0.23 | 0.13 | 0.79 | 0.74 | 0.47 | 20.0 | 72,773 | |
| V+2V+I | 0.32 | 0.23 | 0.13 | 0.81 | 0.75 | 0.50 | 30.0 | 107,432 | |
| V2X-ViT | V+I | 0.50 | 0.38 | 0.20 | 0.54 | 0.41 | 0.21 | 0.338 | 2,211 |
| V+V | 0.53 | 0.43 | 0.26 | 0.65 | 0.55 | 0.35 | 0.338 | 1,305 | |
| V+V+I | 0.56 | 0.45 | 0.26 | 0.70 | 0.59 | 0.36 | 0.667 | 9,215 | |
| V+2V | 0.60 | 0.53 | 0.33 | 0.79 | 0.73 | 0.49 | 0.667 | 7,022 | |
| V+2V+I | 0.60 | 0.52 | 0.33 | 0.80 | 0.74 | 0.50 | 1.014 | 9,976 | |
| V2VAM | V+I | 0.53 | 0.40 | 0.22 | 0.57 | 0.44 | 0.24 | 0.423 | 2,264 |
| V+V | 0.53 | 0.43 | 0.26 | 0.68 | 0.56 | 0.37 | 0.423 | 2,253 | |
| V+V+I | 0.58 | 0.47 | 0.27 | 0.74 | 0.63 | 0.39 | 0.846 | 10,301 | |
| V+2V | 0.61 | 0.54 | 0.34 | 0.82 | 0.76 | 0.53 | 0.846 | 8,117 | |
| V+2V+I | 0.62 | 0.54 | 0.34 | 0.84 | 0.78 | 0.53 | 1.269 | 11,302 | |
| CoBEVT | V+I | 0.48 | 0.39 | 0.20 | 0.52 | 0.43 | 0.21 | 0.500 | 2,707 |
| V+V | 0.50 | 0.41 | 0.26 | 0.64 | 0.55 | 0.37 | 0.500 | 2,393 | |
| V+V+I | 0.54 | 0.44 | 0.25 | 0.70 | 0.60 | 0.37 | 1.000 | 11,059 | |
| V+2V | 0.57 | 0.49 | 0.33 | 0.80 | 0.73 | 0.54 | 1.000 | 8,987 | |
| V+2V+I | 0.59 | 0.50 | 0.32 | 0.82 | 0.75 | 0.54 | 1.500 | 12,381 | |
| ERMVP | V+I | 0.61 | 0.46 | 0.24 | 0.65 | 0.51 | 0.26 | 0.346 | 1,730 |
| V+V | 0.61 | 0.49 | 0.30 | 0.76 | 0.63 | 0.42 | 0.346 | 1,730 | |
| V+V+I | 0.63 | 0.51 | 0.31 | 0.79 | 0.67 | 0.43 | 0.692 | 3,460 | |
| V+2V | 0.66 | 0.57 | 0.37 | 0.87 | 0.80 | 0.57 | 0.692 | 3,460 | |
| V+2V+I | 0.64 | 0.56 | 0.36 | 0.87 | 0.81 | 0.57 | 1.038 | 5,190 | |
| CoSDH | V+I | 0.57 | 0.43 | 0.23 | 0.58 | 0.44 | 0.24 | 0.0595 | 298 |
| V+V | 0.63 | 0.47 | 0.29 | 0.67 | 0.54 | 0.35 | 0.0595 | 298 | |
| V+V+I | 0.68 | 0.53 | 0.32 | 0.72 | 0.62 | 0.39 | 0.1190 | 595 | |
| V+2V | 0.74 | 0.61 | 0.40 | 0.79 | 0.72 | 0.51 | 0.1190 | 595 | |
| V+2V+I | 0.75 | 0.61 | 0.41 | 0.81 | 0.73 | 0.52 | 0.1785 | 893 | |
| Agent | Modality | mAP (BEV) | mAP (3D) | ||||
|---|---|---|---|---|---|---|---|
| @0.3 | @0.5 | @0.7 | @0.3 | @0.5 | @0.7 | ||
| V | LiDAR | 0.78 | 0.56 | 0.31 | 0.70 | 0.38 | 0.21 |
| LiDAR+Camera | 0.80 | 0.57 | 0.33 | 0.72 | 0.42 | 0.23 | |
| V+V | LiDAR | 0.81 | 0.65 | 0.39 | 0.75 | 0.54 | 0.22 |
| LiDAR+Camera | 0.80 | 0.65 | 0.41 | 0.74 | 0.54 | 0.25 | |
| V+2V | LiDAR | 0.90 | 0.84 | 0.56 | 0.87 | 0.73 | 0.30 |
| LiDAR+Camera | 0.90 | 0.83 | 0.61 | 0.87 | 0.74 | 0.35 | |
Multi-modal feature fusion (LiDAR+Camera) consistently improves performance, especially with more cooperative agents.
| Model | Agent | minADE@1s ↓ | minADE@3s ↓ | minADE@5s ↓ | minFDE@1s ↓ | minFDE@3s ↓ | minFDE@5s ↓ |
|---|---|---|---|---|---|---|---|
| V2VNet | V | 0.4895 | 2.1648 | 4.4538 | 0.9307 | 5.7344 | 10.6390 |
| V+V | 0.5424 | 1.5611 | 3.5591 | 0.7528 | 4.4820 | 9.1667 | |
| V+2V | 0.7610 | 1.3796 | 2.9096 | 1.0079 | 3.4762 | 7.3660 | |
| CMP | V | 0.3851 | 0.8421 | 1.4022 | 0.5786 | 1.6125 | 3.0712 |
| V+V | 0.4076 | 0.9330 | 1.6416 | 0.6143 | 1.8492 | 3.7389 | |
| V+2V | 0.3214 | 0.7252 | 1.2893 | 0.4723 | 1.4551 | 2.9716 |
Short-term predictions (1s) achieve strong accuracy, while long-horizon reasoning (5s) remains an open research challenge.
All dataset files are released under CC BY-NC-SA 4.0 (see DATA_LICENSE). The code is MIT licensed.
If you found this work to be useful in your own research, please consider citing the following:
@inproceedings{CooperScene,
title={CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization},
author={Bo Wu* and Ruoshen Mo* and Justin Yue and Yanyu Zhang and Janice Nguyen and Guoyuan Wu and Amit Roy-Chowdhury and Matthew J. Barth and Hang Qiu},
booktitle={European Conference on Computer Vision},
year={2026},
}