TacitFlow: Learning Workflow Representations for Tacit-Knowledge-Grounded Machine Learning Engineering Agent
Publication Date: 8/13/2026
Event: The 1st International Workshop on AI Data Scientist (KDD 2026 Workshop)
Reference: pp. 1-17, 2-26
Authors: Yushan Jiang, University of Connecticut; Wenchao Yu, NEC Laboratories America, Inc.; Minyoung Choe, KAIST; Dongjin Song, University of Connecticut; Jingchao Ni, University of Houston; Wei Cheng, NEC Laboratories America, Inc.; Haifeng Chen, NEC Laboratories America, Inc.
Abstract: Large language model (LLM) agents have shown promise for automating machine learning engineering (MLE), but their iterative improvement remains weakly grounded. In particular, most existing methods lack structured representations of how expert pipeline components compose across tasks. Instead, their solution development is often driven by prior trials and flat textual context from external retrieval sources, with limited explicit support for reasoning over step dependencies and reusable pipeline compositions. To bridge this gap, we introduce TacitFlow, a knowledge-grounded MLE agent built on structured representations of expert workflows. We distill expert solutions across image, tabular, and text domains into heterogeneous workflow graphs of canonicalized pipeline steps and execution dependencies. From this graph, we learn contrastive node embeddings that capture multi-modal step semantics from technical descriptions and code snippets, and train a task-conditioned ranking model to capture workflow performance preferences from competition leaderboard rankings. During runtime, these representations support a retrieve, edit, and rerank refinement loop, where structured workflow context guides explicit candidate edits and the task-conditioned ranking model prioritizes promising workflows before execution. This design avoids exhaustively running every candidate while making agent decisions traceable to individual workflow steps. Experiments across three modalities show that our workflow-structured design yields promising performance on a diverse subset of competitions from MLE-Bench, while case studies illustrate transparent, workflow-grounded agent behavior on end-to-end MLE tasks.
Publication Link: https://usail-hkust.github.io/aidatasci/assets/papers/tacitflow.pdf


Leave a Reply
Want to join the discussion?Feel free to contribute!