While Large Language Models (LLMs) excel at automated reasoning, their ability to interpret specialized Cyber-Physical Systems (CPSs) remains largely unestablished. Industrial Control Systems (ICSs) are especially challenging, as understanding their dynamics typically relies heavily on proprietary documentation or domain-specific algorithms. This paper presents a systematic empirical evaluation of general-purpose LLMs for ICS process comprehension from PLC register time series with limited semantic labels. We introduce a modular benchmarking framework that transforms register logs into structured prompts and evaluates model outputs across five tasks, ranging from register behavior classification and domain inference to device inventory, operating bounds, and candidate causal invariant extraction. Across four real and simulated datasets and six LLMs, we find that frontier models are reliable on closed-set and coarse-grained tasks, reflecting familiarity with well-studied testbeds, but degrade sharply on open-set device inventory and fine-grained relationship extraction. These results suggest that LLMs can support industrial process analysis when coupled with structured validation, but they are not yet sufficient for autonomous process comprehension in realistic ICS settings.
Towards LLM-based Process Comprehension for Industrial Control Systems
Furri, Geremia;Donadel, Denis;Merro, Massimo
2026-01-01
Abstract
While Large Language Models (LLMs) excel at automated reasoning, their ability to interpret specialized Cyber-Physical Systems (CPSs) remains largely unestablished. Industrial Control Systems (ICSs) are especially challenging, as understanding their dynamics typically relies heavily on proprietary documentation or domain-specific algorithms. This paper presents a systematic empirical evaluation of general-purpose LLMs for ICS process comprehension from PLC register time series with limited semantic labels. We introduce a modular benchmarking framework that transforms register logs into structured prompts and evaluates model outputs across five tasks, ranging from register behavior classification and domain inference to device inventory, operating bounds, and candidate causal invariant extraction. Across four real and simulated datasets and six LLMs, we find that frontier models are reliable on closed-set and coarse-grained tasks, reflecting familiarity with well-studied testbeds, but degrade sharply on open-set device inventory and fine-grained relationship extraction. These results suggest that LLMs can support industrial process analysis when coupled with structured validation, but they are not yet sufficient for autonomous process comprehension in realistic ICS settings.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



