NEC Laboratories America

From Intent to Executable Code on the Computing Continuum | Projects

The physical world operates on deadlines that software cannot negotiate.

A production line moves at a fixed speed. A battleship cannot wait for a response from a distant data center. A traffic system either detects a conflict in time to act or it does not. Software is increasingly embedded in these operations, and it must respond within the time limits the physical process allows.

Generative AI has made writing that software dramatically easier. But it has done little to solve the harder question that comes next: where should the code run?

Modern applications rarely execute in one place. They span a computing continuum from sensors and embedded devices, through on-premise edge infrastructure and 5G networks, to hyperscale cloud. Each layer offers a different combination of latency, capacity, availability, and cost. Today, developers manage these tradeoffs largely by hand. They decide how to divide an application, where to place each component, and how to adapt that placement as workloads, infrastructure, and prices change. For latency-critical environments such as factories, ports, vessels, and transportation systems, this can become one of the largest barriers to deploying new applications. It consumes the resource organizations have least of: engineers with the distributed-systems expertise to make these decisions correctly.

As code generation becomes increasingly commoditized, the source of advantage shifts. The challenge is no longer just generating the code, but deciding where and how that code should run. We are building toward a world where an application is described once, in natural language, and generated in a form that can be distributed across the computing continuum. A runtime then continuously decides where each component should execute, adapting to changes in workload and infrastructure while meeting the application’s latency requirements at the lowest practical cost. The goal is to move from intent to running code without requiring developers to manually engineer everything in between.

We are focused on overcoming several barriers. First, generated code is not designed for distributed execution. An LLM typically produces a monolithic program with no representation of where different components should run or how placement affects latency and cost. Distributing that application therefore requires developers to restructure it after generation. What is missing is a way to generate code whose architecture reflects the infrastructure on which it will execute. Second, placement is a continuous decision, not a one-time optimization. Edge resources can provide low latency when lightly loaded but degrade quickly as demand rises. The best placement therefore changes with the workload, network, infrastructure, and cost. Third, distributed applications introduce their own orchestration challenges. Microservice pipelines often need to scale up and scale out together, even though today’s orchestration systems tend to treat these as separate decisions. And in many pipelines, communication between services becomes the primary performance bottleneck.

Our approach is to treat code generation, placement, and execution as a single problem. DiCE and DiCE-M generate applications as distributed programs from the outset and execute them across edge and cloud resources, rather than generating a monolith and partitioning it afterward. Our work on LLM-based distributed code generation extends this approach toward cost-efficient cloud execution, while latency-driven execution focuses on applications where meeting a physical-world deadline is the binding constraint. Once an application is running, placement must be adapted. CLAP makes these decisions dynamically and cost-consciously, while our more recent work on agentic microservice placement allows the placement strategy to reason for changing workloads and infrastructure. ECO-LLM approaches the problem from the other direction, using LLMs to solve and customize systems-optimization problems as conditions change. Around this are the mechanisms required to make distributed pipelines practical in production. Our work on latency-aware resource allocation, content-aware autoscaling, and simultaneous scale-up and scale-out improves how orchestration platforms respond to changing demand. The DataX family (DataXe, DataXc, and DataX Allocator) simplifies the development of distributed streaming applications, reduces the cost of communication between microservices, and dynamically manages resources at the edge.

Our AI systems move beyond AI that simply writes code toward AI that can generate, place, and continuously operate distributed applications across the computing continuum, and adapting execution as workloads and infrastructure change while meeting the application’s latency and cost requirements.

Read Related Publications

Agentic Placement of Microservices on the Computing Continuum

Deploying microservices across the computing continuum (edge–cloud) requires placement decisions that adapt to workload variation and heterogeneous infrastructure, yet existing solutions often rely on static policies or opaque heuristics. We present Bellona a system for reliable and auditable Large

Latency-driven Execution of LLM-generated Application Code on the Computing Continuum

Latency-critical applications demand quick responses. Ideally, detailed insights are preferable for the best decision making and response actions. However, in situations when detailed insights cannot be provided quickly, even basic information goes a long way in tackling the situation effectively. For

LLM-based Distributed Code Generation and Cost-Efficient Execution in the Cloud

The advancement of Generative Artificial Intelligence (AI), particularly Large Language Models (LLMs), is reshaping the software industry by automating code generation. Many LLM-driven distributed processing systems rely on serial code generation constrained by predefined libraries, limiting flexibility

DiCE-M: Distributed Code Generation and Execution for Marine Applications – An Edge-Cloud Approach

Edge computing has emerged as a transformative technology that reduces application latency, improves cost efficiency, enhances security, and enables large-scale deployment of applications across various domains. In environmental monitoring, systems such as MegaSense[49], use low-cost sensors to gather

DiCE: Distributed Code generation and Execution

Generative artificial intelligence (GenAI), specifically, Large Language Models (LLMs), have shown tremendous potential in automating several tasks and improving human productivity. Recent works have shown them to be quite useful in writing and summarizing text (articles, blogs, poems, stories, songs,

ECO-LLM: LLM-based Edge Cloud Optimization

AI/ML techniques have been used to solve systems problems, but their applicability to customize solutions on-the-fly has been limited. Traditionally, any customization required manually changing the AI/ML model or modifying the code, configuration parameters, application settings, etc. This incurs too

CLAP: Cost and Latency-Aware Placement of Microservices on the Computing Continuum

For microservices-based real-time stream processing applications, computing at the edge delivers fast responses for low workloads, but as workload increases, the response time starts to slow down due to limited compute capacity. Abundant compute capacity in the cloud delivers fast responses even for

Improving Real-time Data Streams Performance on Autonomous Surface Vehicles using DataX

In the evolving Artificial Intelligence (AI) era, the need for real-time algorithm processing in marine edge environments has become a crucial challenge. Data acquisition, analysis, and processing in complex marine situations require sophisticated and highly efficient platforms. This study optimizes

LARA: Latency-Aware Resource Allocator for Stream Processing Applications

One of the key metrics of interest for stream processing applications is “latency”, which indicates the total time it takes for the application to process and generate insights from streaming input data. For mission-critical video analytics applications like surveillance and monitoring, it is of paramount

Scale Up while Scaling Out Microservices in Video Analytics Pipelines

Modern video analytics applications comprise multiple microservices chained together as pipelines and executed on container orchestration platforms like Kubernetes. Kubernetes automatically handles the scaling of these microservices for efficient application execution. There are two popular choices for

Content-aware auto-scaling of stream processing applications on container orchestration platforms

Modern applications are designed as an interacting set of microservices, and these applications are typically deployed on container orchestration platforms like Kubernetes. Several attractive features in Kubernetes make it a popular choice for deploying applications, and automatic scaling is one such

DataX Allocator: Dynamic resource management for stream analytics at the Edge

Serverless edge computing aims to deploy and manage applications so that developers are unaware of challenges associated with dynamic management, sharing, and maintenance of the edge infrastructure. However, this is a non-trivial task because the resource usage by various edge applications varies based

DataXc: Flexible and efficient communication in microservices-based stream analytics pipelines

A big challenge in changing a monolithic application into a performant microservices-based application is the design of efficient mechanisms for microservices to communicate with each other. Prior proposals range from custom point-to-point communication among microservices using protocols like gRPC to

DataXe: A System for Application Self-optimization in Serverless Edge Computing Environments

A key barrier to building performant, remotely managed and self-optimizing multi-sensor, distributed stream processing edge applications is high programming complexity. We recently proposed DataX [1], a novel platform that improves programmer productivity by enabling easy exchange, transformations, and