NEC Laboratories America

Projects | eigen Video Understanding

eigen Video Understanding

Aiming to achieve near-human video understanding, in this project we analyze several gigabytes of spatiotemporal data to perform action recognition, multi-person tracking, object permanence and video reasoning. eigen has built a scalable video-understanding platform for long-form video reasoning that scales to new environments and camera angles without any re-training.

eigen also provides a system platform. Running both on the cloud (AWS) and on-prem, it can scale up to thousands of streams into it for cloud-based AI processing. Our AI video algorithms provide efficient streaming and inference. eigen has a web frontend and support for iOS/Android playback. eigen has been tested in several retail POCs serving 200+ streams; its behavioral analytics have also been evaluated through various NEC customers. Using mixed precision and TensorRT, eigen is extremely efficient and incurs very low cloud costs.

Publications

  • Hopper: Multi-hop Transformer for Spatiotemporal Reasoning. In ICLR 2021.
  • 15 Keypoints is All You Need. Michael Snower, Asim Kadav, Farley Lai, Hans Peter Graf. In CVPR 2020. PDFRanked #1 in PoseTrack
  • Tripping Through Time: Efficient Localization of Activities in Videos. (Spotlight) Meera Hahn, Asim Kadav, James M. Rehg, Hans Peter Graf. In CVPR Workship on Language and Vision, 2019. Also, appears in BMVC’20. PDF
  • Visual Entailment: A Novel Task for Fine-Grained Image Understanding. Ning Xie, Farley Lai, Derek Doran, Asim Kadav. In NeurIPS Workshop on Visually-Grounded Interaction and Language, 2018 (ViGIL’18). PDFDatasetleaderboard
  • Teaching Syntax using Adverserial Distraction. Juho Kim, Christopher Malon, and Asim Kadav. In FEVER-EMNLP, 2018. PDF
  • Attend and Interact: Higher-Order Object Interactions for Video Understanding. Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf. In CVPR, 2018. PDFVideo

Collaborators: Hans Peter Graf, Farley Lai, Asim Kadav

Eigen Collaborators: Asim Kadav (Lead), Farley Lai, Deep Patel, Rowena Chen, Anupriya Prasad, Shayan Bhatti, Shrey Shah, Ramakrishnan Sundareswaran, Likitha Lakshminarayan, Aagam Shah, Maxwell Springer, Aaron Rosati

eigen Video Understanding

Video Publications

Understanding Road Layout from Videos as a Whole

In this paper, we address the problem of inferring the layout of complex road scenes from video sequences. To this end, we formulate it as a top-view road attributes prediction problem and our goal is to predict these attributes for each frame both accurately and consistently. In contrast to prior work,

S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data Generation

We propose a sequential variational autoencoder to learn disentangled representations of sequential data (e.g., videos and audios) under self-supervision. Specifically, we exploit the benefits of some readily accessible supervision signals from input data itself or some off-the-shelf functional models