The natural solution is to split the inference between the device and the cloud. But that creates a second problem: data must move between them, and today’s compression technologies were designed for people, not AI. Traditional video codecs optimize what looks good to the human eye. In doing so, they may discard subtle visual information that matters to a detector while preserving details the model does not need. The result is a double penalty: accuracy is lost once by shrinking the model to fit the device, and again by compressing its inputs for the wrong consumer. At scale, both penalties become expensive. They apply to every device and every stream, so even small improvements in accuracy, bandwidth, or compute efficiency multiply across thousands or millions of endpoints.
We are building toward an edge-cloud inference fabric that jointly optimizes models, compute, and communication.
Models can split dynamically between device and cloud based on the application, available resources, and operating conditions. At the same time, the data crossing the network is compressed for the machine that will consume it, preserving the information that matters for inference rather than the information that matters primarily to human perception. The goal is to deliver server-class intelligence with edge-class latency, bandwidth, and cost, while bringing more capable AI to resource-constrained devices without requiring every device to become a data center.
Two structural barriers stand in the way of efficient edge-cloud intelligence: today’s models were not designed to be split, and today’s compression standards were not designed for machines. First, modern multitask models are optimized for accuracy, not distributed execution. Their internal representations are rarely organized around a natural split point where a small amount of data preserves most of the information needed for downstream tasks. Even when a useful split exists, the best place to divide the model depends on conditions that change at runtime like available compute, bandwidth, workload, and application context. A model partitioned once at deployment can therefore be poorly matched to the conditions under which it actually operates. Second, the compression infrastructure already deployed in the world was built for human perception. Standards such as JPEG, H.264, and H.265 are deeply embedded in real systems, but their operations are not naturally differentiable. That makes it difficult to directly optimize them using the downstream machine-vision objective. Replacing standard codecs with learned alternatives can simplify the optimization problem, but at the cost of compatibility with the enormous base of existing cameras, encoders, networks, and hardware.
Our approach addresses both constraints directly rather than designing around them. For model splitting, we make partitioning dynamic and context-aware. Chimera builds splittable multitask models for device-edge collaboration, allowing execution to adapt to the context in which the model is running. FactionFormer extends this approach to vision transformers, enabling model components to collaborate across the edge according to changing conditions. DyCo dynamically contextualizes models so that smaller models running on constrained devices can recover accuracy they would otherwise lose. For compression, we optimize standard codecs rather than replace them. Differentiable JPEG addresses the challenge of optimizing through JPEG’s non-differentiable operations. Deep Video Codec Control and its vision-oriented variant steer standard video encoders toward representations that preserve the information downstream models need, while remaining compatible with existing codec infrastructure. Our broader work on deep vision under standard codecs characterizes how these compression decisions affect model performance in the first place. Deep learning-based real-time rate and quality control extends the same principle to live streaming, continuously adapting compression as wireless bandwidth and network conditions change. Our earlier work on coprocessors provides the hardware foundation for this broader effort, tracing the problem back to efficiently executing convolutional networks on resource-constrained systems.
Our AI systems make both sides of edge-cloud inference adaptive to the machine rather than fixed around the infrastructure: models can change how and where they execute, while standard compression can change what information it preserves. The result is an edge-cloud pipeline that uses limited compute and bandwidth more efficiently without giving up the accuracy of larger models or compatibility with deployed systems.