CooperScene

Multi-Modal Cooperative Autonomy Benchmark
with C-V2X Communication Characterization

University of California, Riverside
ECCV 2026, Oral Presentation

CooperScene

Multi-Modal Cooperative Autonomy Benchmark
with C-V2X Communication Characterization

Overview

CooperScene is the first real-world, multi-agent, multi-modal cooperative autonomy dataset with C-V2X communication characterization. It is organized into diverse real-world scenes—spanning urban intersection, freeway ramp, local street, and parking lot—that introduce different traffic and occlusion patterns. In urban intersection scnario, our data involves three connected autonomous vehicles and one instrumented infrastructure roadside unit, all equipped with multi-modal sensors and commercial C-V2X communication radios.

The dataset provides tightly aligned LiDAR, camera, GNSS/IMU, and C-V2X network data, enabling the community to jointly evaluate perception performance and the practical constraints under real-world vehicular networks. CooperScene supports 3D object detection and motion prediction benchmarks with synchronized pair-wise C-V2X throughput measurements.

59K Synchronized LiDAR Frames
53K Image Frames
344K 3D Bounding Box Labels
4 Cooperative Agents

Key Features

Multi-Agent Cooperative Scenes

High-quality interactive scenes from three vehicles and one infrastructure agent across diverse real-world urban intersection scenarios.

C-V2X Communication

Real-world C-V2X network traces including latency, throughput, packet loss rate and jitter — bridging wireless networking and autonomous driving research.

Multi-Modal Sensing

LiDAR, camera, GNSS/IMU data streams per agent, well-calibrated and temporally synchronized with sub-millisecond accuracy.

Centimeter-Level Localization

GNSS-RTK positioning with spatial-temporal ICP alignment achieving an average RMSE of 0.2m across all agents.

Precise Synchronization

PTP-based time synchronization and hardware triggering ensure globally consistent timestamps across all platforms and sensors.

Unified Benchmark

Comprehensive benchmarks for detection and motion prediction evaluating both perception accuracy and communication efficiency.

Benchmark

Model Agents mAP (C-V2X) mAP (Unlimited) Sharing
Size (MB)
Latency (ms)
under C-V2X
@0.3 @0.5 @0.7 @0.3 @0.5 @0.7
V2VNet V+I 0.43 0.33 0.20 0.55 0.42 0.23 10.0 64,349
V+V 0.35 0.25 0.15 0.66 0.55 0.34 10.0 64,038
V+V+I 0.36 0.27 0.15 0.70 0.60 0.37 20.0 76,879
V+2V 0.32 0.23 0.13 0.79 0.74 0.47 20.0 72,773
V+2V+I 0.32 0.23 0.13 0.81 0.75 0.50 30.0 107,432
V2X-ViT V+I 0.50 0.38 0.20 0.54 0.41 0.21 0.338 2,211
V+V 0.53 0.43 0.26 0.65 0.55 0.35 0.338 1,305
V+V+I 0.56 0.45 0.26 0.70 0.59 0.36 0.667 9,215
V+2V 0.60 0.53 0.33 0.79 0.73 0.49 0.667 7,022
V+2V+I 0.60 0.52 0.33 0.80 0.74 0.50 1.014 9,976
V2VAM V+I 0.53 0.40 0.22 0.57 0.44 0.24 0.423 2,264
V+V 0.53 0.43 0.26 0.68 0.56 0.37 0.423 2,253
V+V+I 0.58 0.47 0.27 0.74 0.63 0.39 0.846 10,301
V+2V 0.61 0.54 0.34 0.82 0.76 0.53 0.846 8,117
V+2V+I 0.62 0.54 0.34 0.84 0.78 0.53 1.269 11,302
CoBEVT V+I 0.48 0.39 0.20 0.52 0.43 0.21 0.500 2,707
V+V 0.50 0.41 0.26 0.64 0.55 0.37 0.500 2,393
V+V+I 0.54 0.44 0.25 0.70 0.60 0.37 1.000 11,059
V+2V 0.57 0.49 0.33 0.80 0.73 0.54 1.000 8,987
V+2V+I 0.59 0.50 0.32 0.82 0.75 0.54 1.500 12,381
ERMVP V+I 0.61 0.46 0.24 0.65 0.51 0.26 0.346 1,730
V+V 0.61 0.49 0.30 0.76 0.63 0.42 0.346 1,730
V+V+I 0.63 0.51 0.31 0.79 0.67 0.43 0.692 3,460
V+2V 0.66 0.57 0.37 0.87 0.80 0.57 0.692 3,460
V+2V+I 0.64 0.56 0.36 0.87 0.81 0.57 1.038 5,190
CoSDH V+I 0.57 0.43 0.23 0.58 0.44 0.24 0.0595 298
V+V 0.63 0.47 0.29 0.67 0.54 0.35 0.0595 298
V+V+I 0.68 0.53 0.32 0.72 0.62 0.39 0.1190 595
V+2V 0.74 0.61 0.40 0.79 0.72 0.51 0.1190 595
V+2V+I 0.75 0.61 0.41 0.81 0.73 0.52 0.1785 893
Agent Modality mAP (BEV) mAP (3D)
@0.3 @0.5 @0.7 @0.3 @0.5 @0.7
V LiDAR 0.78 0.56 0.31 0.70 0.38 0.21
LiDAR+Camera 0.80 0.57 0.33 0.72 0.42 0.23
V+V LiDAR 0.81 0.65 0.39 0.75 0.54 0.22
LiDAR+Camera 0.80 0.65 0.41 0.74 0.54 0.25
V+2V LiDAR 0.90 0.84 0.56 0.87 0.73 0.30
LiDAR+Camera 0.90 0.83 0.61 0.87 0.74 0.35

Multi-modal feature fusion (LiDAR+Camera) consistently improves performance, especially with more cooperative agents.

Model Agent minADE@1s ↓ minADE@3s ↓ minADE@5s ↓ minFDE@1s ↓ minFDE@3s ↓ minFDE@5s ↓
V2VNet V 0.4895 2.1648 4.4538 0.9307 5.7344 10.6390
V+V 0.5424 1.5611 3.5591 0.7528 4.4820 9.1667
V+2V 0.7610 1.3796 2.9096 1.0079 3.4762 7.3660
CMP V 0.3851 0.8421 1.4022 0.5786 1.6125 3.0712
V+V 0.4076 0.9330 1.6416 0.6143 1.8492 3.7389
V+2V 0.3214 0.7252 1.2893 0.4723 1.4551 2.9716

Short-term predictions (1s) achieve strong accuracy, while long-horizon reasoning (5s) remains an open research challenge.

Download

All dataset files are released under CC BY-NC-SA 4.0 (see DATA_LICENSE). The code is MIT licensed.

Mini Set

Download

~2.9 GB

Full Dataset

Download

~94.3 GB

Visualization

Download

~90.9 GB

Citation

If you found this work to be useful in your own research, please consider citing the following:

@inproceedings{CooperScene,
  title={CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization},
  author={Bo Wu* and Ruoshen Mo* and Justin Yue and Yanyu Zhang and Janice Nguyen and Guoyuan Wu and Amit Roy-Chowdhury and Matthew J. Barth and Hang Qiu},
  booktitle={European Conference on Computer Vision},
  year={2026},
}