Modern applications rarely execute in one place. They span a computing continuum from sensors and embedded devices, through on-premise edge infrastructure and 5G networks, to hyperscale cloud. Each layer offers a different combination of latency, capacity, availability, and cost. Today, developers manage these tradeoffs largely by hand. They decide how to divide an application, where to place each component, and how to adapt that placement as workloads, infrastructure, and prices change. For latency-critical environments such as factories, ports, vessels, and transportation systems, this can become one of the largest barriers to deploying new applications. It consumes the resource organizations have least of: engineers with the distributed-systems expertise to make these decisions correctly.
As code generation becomes increasingly commoditized, the source of advantage shifts. The challenge is no longer just generating the code, but deciding where and how that code should run. We are building toward a world where an application is described once, in natural language, and generated in a form that can be distributed across the computing continuum. A runtime then continuously decides where each component should execute, adapting to changes in workload and infrastructure while meeting the application’s latency requirements at the lowest practical cost. The goal is to move from intent to running code without requiring developers to manually engineer everything in between.
We are focused on overcoming several barriers. First, generated code is not designed for distributed execution. An LLM typically produces a monolithic program with no representation of where different components should run or how placement affects latency and cost. Distributing that application therefore requires developers to restructure it after generation. What is missing is a way to generate code whose architecture reflects the infrastructure on which it will execute. Second, placement is a continuous decision, not a one-time optimization. Edge resources can provide low latency when lightly loaded but degrade quickly as demand rises. The best placement therefore changes with the workload, network, infrastructure, and cost. Third, distributed applications introduce their own orchestration challenges. Microservice pipelines often need to scale up and scale out together, even though today’s orchestration systems tend to treat these as separate decisions. And in many pipelines, communication between services becomes the primary performance bottleneck.
Our approach is to treat code generation, placement, and execution as a single problem. DiCE and DiCE-M generate applications as distributed programs from the outset and execute them across edge and cloud resources, rather than generating a monolith and partitioning it afterward. Our work on LLM-based distributed code generation extends this approach toward cost-efficient cloud execution, while latency-driven execution focuses on applications where meeting a physical-world deadline is the binding constraint. Once an application is running, placement must be adapted. CLAP makes these decisions dynamically and cost-consciously, while our more recent work on agentic microservice placement allows the placement strategy to reason for changing workloads and infrastructure. ECO-LLM approaches the problem from the other direction, using LLMs to solve and customize systems-optimization problems as conditions change. Around this are the mechanisms required to make distributed pipelines practical in production. Our work on latency-aware resource allocation, content-aware autoscaling, and simultaneous scale-up and scale-out improves how orchestration platforms respond to changing demand. The DataX family (DataXe, DataXc, and DataX Allocator) simplifies the development of distributed streaming applications, reduces the cost of communication between microservices, and dynamically manages resources at the edge.
Our AI systems move beyond AI that simply writes code toward AI that can generate, place, and continuously operate distributed applications across the computing continuum, and adapting execution as workloads and infrastructure change while meeting the application’s latency and cost requirements.