CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios
While modern autonomous driving systems excel at perceptiontasks such as object detection and trajectory prediction, they lackthe high-level causal reasoning required to interpret traffic accidents.In particular, determining responsibility, such as identifying who is atfault and which traffic rule was violated, remains largely unexploredin current benchmarks. To this end, we introduce CAViAR (CausalAccident Video and Incident Analysis Repository), a human-annotateddashcam benchmark comprising 2,249 real-world accident videos collectedfrom CarCrashDataset (CCD) and Nexar. Each video is annotatedwith structured labels spanning environmental conditions, accidenttype, causal explanation, apparent At-Fault Agent, affected agent, andapparent rule-violation category. We benchmark state-of-the-art visionlanguagemodels (VLMs), including Cosmos-Reason2, Qwen3-VL, andInternVL3. Once class imbalance is accounted for with majority/randombaselines and balanced metrics, perceptual competence is uneven–lighting

