In the rapidly evolving landscape of artificial intelligence, the development of agentic large language models (LLMs) marks a significant milestone. Our project, Agentic LLMs for AI Orchestration, is at the forefront of this innovation. We have engineered an advanced LLM capable of solving complex workflows by integrating computer vision, logic, and compute modules. By leveraging natural language task specifications, our LLM dynamically generates plans to execute tasks using available tools, translating these plans into Python programs that can be seamlessly synthesized and deployed. This adaptability is underpinned by our LLM’s ability to quickly assimilate new tools based on their documentation and code, a feature that sets it apart in the AI orchestration domain.

Agentic LLMs for AI Orchestration Blog Post Graphic

Vijay Kumar B G NEC Labs America“Integrating various modules such as computer vision and logic within the LLM framework allows for a more holistic approach to solving complex tasks. The ability to generate and deploy Python programs based on natural language specifications is a game-changer in the field of AI orchestration,” said Vijay Kumar B G, a Media Analytics department team member.

Project Overview
Our agentic LLM is designed to understand and execute complex workflows by employing a combination of cutting-edge techniques. The natural language task specification is the starting point, guiding the LLM to generate a coherent plan. This plan is then represented as a Python program, which is synthesized to deploy the necessary tools programmatically. The adaptability of our planner is one of its most significant strengths, allowing it to incorporate new tools efficiently by interpreting available documentation and code. This capability ensures that our LLM remains versatile and scalable, ready to tackle a wide range of tasks.

Our training methodology is innovative and efficient. By utilizing reinforced self-training with weak supervision, our LLM can learn and adapt quickly, with the added option of incorporating human feedback to refine its performance further. This approach not only enhances training efficiency but also ensures that our LLM outperforms its competitors on benchmark visual reasoning tasks despite using fewer parameters.

Samuel Schulter NEC Labs AmericaMedia Analytics department team member Samuel Schulter adds, “Our focus on reinforced self-training with weak supervision has been pivotal. It enables the LLM to adapt quickly and efficiently with minimal human intervention. This not only streamlines the workflow but also sets a new benchmark in AI performance.”

Our project is supported by four foundational research papers, each contributing critical insights and advancements to developing our agentic LLM.

  1. Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement: This paper explores visual program synthesis to leverage the reasoning capabilities of large language models for compositional computer vision tasks. The emphasis on training LLMs to write better visual programs highlights the potential for significant improvements in AI orchestration.
  1. Exploring Question Decomposition for Zero-Shot VQA: This paper investigates a question decomposition strategy for visual question answering (VQA) to address the limitations of treating VQA as a single-step task. The findings underscore the importance of nuanced question-answering strategies, integral to our LLM’s ability to handle complex workflows.
  1. Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images: This research focuses on finetuning large vision-language models on specialized datasets, particularly in scenarios where data is scarce. The self-training approach on unlabeled images is particularly relevant to our project, as it informs our methodology for training LLMs under constrained conditions.
  1. Single-Stream Multi-level Alignment for Vision-Language Pretraining: This paper highlights the effectiveness of self-supervised vision-language pretraining and discusses the limitations of dual-stream architectures and the benefits of fine-grained alignment. These insights are crucial for optimizing our LLM’s performance by integrating visual and language components.

Conclusion
The development of Agentic LLMs for AI Orchestration represents a significant advancement in artificial intelligence. By seamlessly integrating computer vision, logic, and compute modules, our LLM is poised to revolutionize the way complex workflows are managed and executed. Supported by robust research and driven by innovative training methodologies, our agentic LLM sets a new standard in AI orchestration, offering unparalleled performance and adaptability.

Read Our News Posts

Haven for Hope

A New Hope: AI Research is Conquering Today’s Computer Vision Plateau

The age of computer vision is upon us, and it’s transforming the way we live, work and interact with the world. Nearly every industry has found a use case to grow revenue, reduce cost, or create exceptional experiences using computer vision. From self-driving cars to retail automation, surgeons to farmers, computer vision is everywhere, providing critical insights that are driving progress in virtually every aspect of our daily lives. Today, images and videos are annotated to train artificial intelligence (AI) models to recognize specific objects, but there is still so much more to be done when it comes to understanding what those objects are doing in real-time.
Our Time Series Data Research Drives Space Systems Innovation

NEC Labs America’s Time Series Data Research Drives Space Systems Innovation

With decreasing hardware costs and increasing demand for autonomic management, many of today’s physical systems are equipped with an extensive network of sensors, generating a considerable amount of time series data daily. A highly valuable source of information, time series data is used by businesses and governments to measure and analyze change over time in complex systems. Organizations must consolidate, integrate and organize a vast amount of time series data from multiple sources to generate insights and business value.
Next-Generation Computing Finally Sees Light Blog Post

Next-Generation Computing Finally Sees Light

Moore's law is dead, as we have squeezed all the innovation out of silicon. Fiber optics is the solution to meet the computing needs of tomorrow. Today, we can already use the light traveling inside fiber optic cables as sensors that measure vibrations, sound, temperature, light, and pressure changes. We're now developing the means to take this to the next level with photonic computing at the speed of light to provide faster reaction time, reduce energy consumption and improve battery range
Using AI To Safely Put The First Woman On The Moon

Using AI To Safely Put The First Woman On The Moon

We are helping to safely bring the first woman astronaut to the moon as part of NASA - National Aeronautics and Space Administration's Artemis Project with our System Invariant Analysis Technology (SIAT). With Lockheed Martin Space's T-Tauri AI platform, our SIAT analytics engine takes the data from the 150,000 sensors and creates a model incorporating over 22 billion data relationships. The AI model is then analyzed to find any irregularities which could lead to a possible malfunction of any of the spacecraft's systems.

Our AI Research Contributing to NASA’s Artemis Space Program

By 2024, the spacecraft "Orion" developed by Lockheed Martin will bring humans to the moon in NASA's Artemis program. The system invariant analysis technology, one of NEC's Artificial Intelligence technologies, will perform checks to ensure that the spacecraft is tested and operating properly during the production phase.

NEC Provides AI-Based Traffic Monitoring System with Fiber-Optic Sensing Technology for NEXCO CENTRAL

NEC Corporation has deployed an AI-based traffic monitoring system to Central Nippon Expressway Company Limited (NEXCO CENTRAL). The system uses fiber-optic sensing and AI technologies to visualize traffic conditions, such as the location, speed, and direction of travel, from vibrations produced by vehicle movement.