NEC Laboratories America

Media Analytics | Abhishek Aich

PROJECTS

Abhishek Aich

Abhishek Aich

Senior Researcher

Media Analytics

Linkedin Logo

About

Abhishek Aich is a Senior Researcher in the Media Analytics Department at NEC Laboratories America. He received his Bachelor of Technology (B.Tech.) in Electronics and Communications Engineering at Biju Patnaik University of Technology, Odisha, India, Master of Science (M.S.) in Electronics and Communication Engineering at National Institute of Technology, Tiruchirappalli, India and his Ph.D. in 2023 in Electrical and Computer Engineering from the University of California, Riverside, USA. At UC Riverside, he worked with Prof. Amit K. Roy-Chowdhury on topics in computer vision and deep learning.

At NEC Labs, Dr. Aich’s work focuses on vision-language frameworks, perception problems for autonomous driving, dynamic multi-task architectures, efficient transformers, and generative AI. His recent contributions include research on progressive token length scaling for universal segmentation (presented at ICLR 2025, and efficiency–accuracy trade-offs in DETR-style models (CVPRW 2024). He has also helped organize key workshops at ICCV, AAAI, and WACV, and was part of the winning team in the U.S. DOT Intersection Safety Challenge 2025, earning top recognition in Stage 1B.

Projects

Autonomous Driving

Overview: While autonomous cars are rapidly becoming a reality, it remains a challenge to scalably deploy them across geographies and conditions. Our full-stack autonomy solutions include perception, prediction, planning, simulation and DevOps that leverage the latest advances in generative AI, neural rendering, large language models, diffusion models and transformers.


Dynamic Multi-Task Architectures

Overview: Multi-task learning commonly encounters competition for resources among tasks when model capacity is limited. We develop neural architectures that allow control over the relative importance of tasks and total compute cost during inference time.


Open Vocabulary Perception

Overview: We develop open vocabulary perception methods that combine the power of vision and language to provide rich descriptions of objects in scenes, including their attributes, behaviors, relations and interactions.


Publications

Teaching AI to Edit Driving Scenes It Has Never Seen with HorizonWeaver

Our HorizonWeaver software edits driving scenes with instruction-guided AI, adding traffic, changing weather, and generalizing to unseen roads, all while preserving the safety-critical details that keep autonomous vehicle testing honest.

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centric or artistic data, struggle with dense, safety-critical driving layouts. We propose

NEC Labs America Attends CVPR 2026 in Denver, CO June 3-7, 2026

NEC Labs America headed to Denver for CVPR 2026, one of the most prestigious gatherings in computer vision, machine learning, and pattern recognition. The IEEE/CVF Conference on Computer Vision and Pattern Recognition brought innovators from around the world to share breakthroughs.

Which Should We Test Next? Performance Gap Discovery for Driving VLMs

Driving vision-language models (VLMs) must accurately understand scenes across diverse conditions defined by Operational Design Domains (ODDs), yet verificationremains sparse: many slices are missing, making empirical failure rates unreliable. We propose SLICESCORER, a deterministic scoring rule for

Driving Video Retrieval for Complex Queries with Structured Grounding

Video retrieval at scale is central to data curation and safety validation in autonomous driving, where users want to find not only scenes but also dynamic events such as cut-ins and hard braking. Existing vision-language and keyword-based retrieval methods often miss these events because the relevant

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes

Ensuring safety in autonomous driving requires scalable generation of realistic, controllable driving scenes beyond what real-world testing provides. Yet existing instruction guided image editors, trained on object-centric or artistic data, struggle with dense, safety-critical driving layouts. We propose

Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation

Vision transformer-based models bring significant improvements for image segmentation tasks. Although these architectures offer powerful capabilities irrespective of specific segmentation tasks, their use of computational resources can be taxing on deployed devices. One way to overcome this challenge

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality available for such analysis (i.e. no LiDAR, GPS, etc.), existing

NeurIPS 2025 in San Diego from Nov 30th to Dec 5th, 2025

NEC Laboratories America is heading to San Diego for NeurIPS 2025, where our researchers will present cutting-edge work spanning optimization, AI systems, language modeling, and trustworthy machine learning. multi-agent coordination, scalable training, efficient inference, and techniques for detecting

Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations

Obtaining high-quality fine-grained annotations for traffic signs is critical for accurate and safe decision-making in autonomous driving. Widely used datasets, such as Mapillary, often provide only coarse-grained labels without distinguishing semantically important types such as stop signs or speed

Abhishek Aich is Organizing the Anomaly Detection with Foundation Models Workshop, held in conjunction with ICCV 2025

We are proud to share that our Abhishek Aich is serving as one of the organizers of the Anomaly Detection with Foundation Models Workshop, held in conjunction with the International Conference on Computer Vision, October 20, 2025, 08:55 AM – 12:15 PM HST in Room 314 at theHawaii Convention Center,

Sparsh Garg Presents Mapillary Vistas Validation for Fine-Grained Traffic Signs at DataCV 2025

Our Sparsh Garg, a Senior Associate Researcher in the Media Analytics Department, will present “Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations” at the Data Computer Vision (DataCV) 2025 workshop as part of ICCV 2025 in Honolulu,

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality available for such analysis (i.e., no LiDAR, GPS, etc.), existing

Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation

A powerful architecture for universal segmentation relies on transformers that encode multi-scale image features and decode object queries into mask predictions. With efficiency being a high priority for scaling such models, we observed that the state-of-the-art method Mask2Former uses >50% of its compute

NEC Labs America Attends the 39th Annual AAAI Conference on Artificial Intelligence #AAAI25

Our NEC Lab America team attended the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-25), in Philadelphia, Pennsylvania at the Pennsylvania Convention Center from February 25 to March 4, 2025. The purpose of the AAAI conference series was to promote research in AI and foster scientific

Improving the Efficiency-Accuracy Trade-off of DETR-Style Models in Practice

This report aims to provide a comprehensive view on the inference efficiency of DETR-style detection models. We provide the effect of the basic efficiency techniques and identify the factors that are easily applicable yet effectively improve the efficiency-accuracy trade-off. Specifically, we explore

Efficient Transformer Encoders for Mask2Former-style Models

Vision transformer based models bring significant improvements for image segmentation tasks. Although these architectures offer powerful capabilities irrespective of specific segmentation tasks, their use of computational resources can be taxing on deployed devices. One way to overcome this challenge

Efficient Controllable Multi-Task Architectures

We aim to train a multi-task model such that users can adjust the desired compute budget and relative importance of task performances after deployment, without retraining. This enables optimizing performance for dynamically varying user needs, without heavy computational overhead to train and save models