NEC Laboratories America

Edge Intelligence | Projects

Intelligence is often most valuable when it runs close to the physical world it is interpreting.

A decision about a vehicle should happen in the vehicle, a decision about a production line on the line, and a decision about who is at a door at the door. Sending everything to a distant data center adds latency, consumes bandwidth, and can create privacy concerns that operators are unwilling to accept.

But edge devices have limited resources. A camera, phone, vehicle module, or industrial gateway has only a fraction of the compute and memory that today’s most capable AI models require. Running the entire model locally can therefore sacrifice accuracy, while sending all the data to the cloud gives up many of the benefits of edge computing.

The natural solution is to split the inference between the device and the cloud. But that creates a second problem: data must move between them, and today’s compression technologies were designed for people, not AI. Traditional video codecs optimize what looks good to the human eye. In doing so, they may discard subtle visual information that matters to a detector while preserving details the model does not need. The result is a double penalty: accuracy is lost once by shrinking the model to fit the device, and again by compressing its inputs for the wrong consumer. At scale, both penalties become expensive. They apply to every device and every stream, so even small improvements in accuracy, bandwidth, or compute efficiency multiply across thousands or millions of endpoints.

We are building toward an edge-cloud inference fabric that jointly optimizes models, compute, and communication.

Models can split dynamically between device and cloud based on the application, available resources, and operating conditions. At the same time, the data crossing the network is compressed for the machine that will consume it, preserving the information that matters for inference rather than the information that matters primarily to human perception. The goal is to deliver server-class intelligence with edge-class latency, bandwidth, and cost, while bringing more capable AI to resource-constrained devices without requiring every device to become a data center.

Two structural barriers stand in the way of efficient edge-cloud intelligence: today’s models were not designed to be split, and today’s compression standards were not designed for machines. First, modern multitask models are optimized for accuracy, not distributed execution. Their internal representations are rarely organized around a natural split point where a small amount of data preserves most of the information needed for downstream tasks. Even when a useful split exists, the best place to divide the model depends on conditions that change at runtime like available compute, bandwidth, workload, and application context. A model partitioned once at deployment can therefore be poorly matched to the conditions under which it actually operates. Second, the compression infrastructure already deployed in the world was built for human perception. Standards such as JPEG, H.264, and H.265 are deeply embedded in real systems, but their operations are not naturally differentiable. That makes it difficult to directly optimize them using the downstream machine-vision objective. Replacing standard codecs with learned alternatives can simplify the optimization problem, but at the cost of compatibility with the enormous base of existing cameras, encoders, networks, and hardware.

Our approach addresses both constraints directly rather than designing around them. For model splitting, we make partitioning dynamic and context-aware. Chimera builds splittable multitask models for device-edge collaboration, allowing execution to adapt to the context in which the model is running. FactionFormer extends this approach to vision transformers, enabling model components to collaborate across the edge according to changing conditions. DyCo dynamically contextualizes models so that smaller models running on constrained devices can recover accuracy they would otherwise lose. For compression, we optimize standard codecs rather than replace them. Differentiable JPEG addresses the challenge of optimizing through JPEG’s non-differentiable operations. Deep Video Codec Control and its vision-oriented variant steer standard video encoders toward representations that preserve the information downstream models need, while remaining compatible with existing codec infrastructure. Our broader work on deep vision under standard codecs characterizes how these compression decisions affect model performance in the first place. Deep learning-based real-time rate and quality control extends the same principle to live streaming, continuously adapting compression as wireless bandwidth and network conditions change. Our earlier work on coprocessors provides the hardware foundation for this broader effort, tracing the problem back to efficiently executing convolutional networks on resource-constrained systems.

Our AI systems make both sides of edge-cloud inference adaptive to the machine rather than fixed around the infrastructure: models can change how and where they execute, while standard compression can change what information it preserves. The result is an edge-cloud pipeline that uses limited compute and bandwidth more efficiently without giving up the accuracy of larger models or compatibility with deployed systems.

Read Related Publications

A Perspective on Deep Vision Performance with Standard Image and Video Codecs

Resource-constrained hardware such as edge devices or cell phones often rely on cloud servers to provide the required computational resources for inference in deep vision models. However transferring image and video data from an edge or mobile device to a cloud server requires coding to deal with network

Deep Video Codec Control for Vision Models

Standardized lossy video coding is at the core of almost all real-world video processing pipelines. Rate control is used to enable standard codecs to adapt to different network bandwidth conditions or storage constraints. However standard video codecs (e.g. H.264) and their rate control modules aim to

Deep Learning-Based Real-Time Quality Control of Standard Video Compression for Live Streaming

Ensuring high-quality video content for wireless users has become increasingly vital. Nevertheless, maintaining a consistent level of video quality faces challenges due to the fluctuating encoded bitrate, primarily caused by dynamic video content, especially in live streaming scenarios. Video compression

Deep Learning-Based Real-Time Rate Control for Live Streaming on Wireless Networks

Providing wireless users with high-quality video content has become increasingly important. However, ensuring consistent video quality poses challenges due to variable encodedbitrate caused by dynamic video content and fluctuating channel bitrate caused by wireless fading effects. Suboptimal selection

Differentiable JPEG: The Devil is in The Details

JPEG remains one of the most widespread lossy image coding methods. However, the non-differentiable nature of JPEG restricts the application in deep learning pipelines. Several differentiable approximations of JPEG have recently been proposed to address this issue. This paper conducts a comprehensive

Deep Video Codec Control

Deep Video Codec Control Lossy video compression is commonly used when transmitting and storing video data. Unified video codecs (e.g., H.264 or H.265) remain the emph(Unknown sysvar: (de facto)) standard, despite the availability of advanced (neural) compression approaches. Transmitting videos in the

Retrospective : A Dynamically Configurable Coprocessor For Convolutional Neural Networks

In 2008, parallel computing posed significant challenges due to the complexities of parallel programming and the bottlenecks associated with efficient parallel execution. Inspired by the remarkable scalability achieved by networking and storage systems in handling extensive packet traffic and persistent

FactionFormer: Context-Driven Collaborative Vision Transformer Models for Edge Intelligence

Edge Intelligence has received attention in the recent times for its potential towards improving responsiveness, reducing the cost of data transmission, enhancing security and privacy, and enabling autonomous decisions by edge devices. However, edge devices lack the power and compute resources necessary

DyCo: Dynamic, Contextualized AI Models

Devices with limited computing resources use smaller AI models to achieve low-latency inferencing. However, model accuracy is typically much lower than the accuracy of a bigger model that is trained and deployed in places where the computing resources are relatively abundant. We describe DyCo, a novel

Chimera: Context-Aware Splittable Deep Multitasking Models for Edge Intelligence

Design of multitasking deep learning models has mostly focused on improving the accuracy of the constituent tasks, but the challenges of efficiently deploying such models in a device-edge collaborative setup (that is common in 5G deployments) has not been investigated. Towards this end, in this paper,