The full upstream README, mirrored here for reference. Install config, tool schemas, adoption signals, and an original overview live on the Industrial AIOps listing page.
English · 中文
Ask an AI agent why the line stopped — and get an answer that cites its evidence.
A vendor-neutral, read-first data tap for the factory floor. It speaks 14 field protocols, correlates what it reads across them, and hands your agent an evidence-cited verdict instead of a guess. Every call is audited, and no reading ever phones home.
Prefer a container? The published image is cosign-signed and runs non-root. It speaks MCP over stdio, so keep stdin open and mount a volume for the audit store:
For a hardened or air-gapped deployment (read-only rootfs, cap_drop: ALL, no-new-privileges,
optional on-box LLM) use deploy/margo/compose.yaml and
deploy/airgap/. The analysis engine needs no GPU and no model API — it is
deterministic; an LLM is optional and only phrases the verdict.
| Reads | OPC-UA (+ Historical Access, tag auto-discovery) · Modbus TCP/RTU · S7comm · Mitsubishi MC · Omron FINS · MTConnect · MQTT/Sparkplug B · EtherNet/IP · EtherCAT · PROFINET · SECS/GEM · HART-IP · BACnet/IP · IO-Link — plus read-only REST layers for BAS supervisors (Metasys / Niagara) and Ignition Gateway |
| Figures out | downtime root cause (the flagship copilot), alarm floods (ISA-18.2), broken dataflows, data trustworthiness, OEE, asset inventory, legacy PLC program explainer (ST/AWL/L5X) |
| Governs | audit · budget · risk-tier · undo — on every call, through one engine, from both MCP and CLI |
| Stays yours | no telemetry, no phone-home. Six tools can send data off-box by design (stream_publish, stream_publish_event, uns_publish, historian_push, mqtt_publish, rca_narrate) — IAIOPS_NO_EGRESS=1 withholds all six for an air-gapped posture |
Nine per-industry editions ship in this package — fab · factory · process · building · water ·
warehouse · clinical · renewables · plcnext — each adding its own read-only advisory checks.
Substation / utility telecontrol (IEC-104 · DNP3 · IEC-61850) ships separately as
iaiops-energy.
Four commands. Only one of them touches a device, and it prints what it will send before it sends anything.
onboard status answers the smaller question you have first: which of the six
steps is this site on, and what is the one command that advances it? The six
were always there and nothing stated the order. It is derived from your store and
config.yaml every time, so there is no onboarding state to go stale — edit
config.yaml by hand and the answer stays true.
readiness reads your config and local store and answers one question: which
scenarios can this site run today, and what does each gap need? Every gap comes
with the command that closes it, ranked by how much it unlocks. No agent, no
cloud, no account, and nothing on the wire.
Then the path, in the order that matters — survey what is there, take a bounded sample, and only then explain it:
| contacts a device? | ||
|---|---|---|
| Survey | iaiops scan plan → iaiops scan run | preview sends nothing; the run itemises every packet class it sent |
| Configure | iaiops onboard draft → you merge it into config.yaml | no — it reads the stored scan, and writes nothing |
| Tap | iaiops collect run line1 --duration 7d | yes — and it reports what it saw and what it missed |
| Declare | iaiops tags export → a person fills in role → iaiops tags apply --by <you> | no — the role column comes out empty on purpose |
| Explain | iaiops oee measure --since … --until … · iaiops investigate open · iaiops diag rca | no — all over collected history |
See the whole thing run against a real device in about two minutes, including
a genuine mid-run outage, with ./demo/oee-line/run_demo.sh — no hardware, no
configuration, nothing written outside a temporary directory.
demo/oee-line/ explains what each step is for and what the
numbers do and do not claim.
OT is exactly where you want an agent on a tight leash. The read paths are the product; the few write paths are OT-dangerous, off by default, and gated by MOC discipline — dry-run, one-shot approval, undo capture, hash-chained audit.
The analysis layers cannot reach a language model. That is a guard, not a slogan:
tests/test_brain_is_llm_free.py scans eight packages — brain, discovery, runtime,
readiness, collect, knowledge, retain, connectors — for any import that could reach one,
and an empty result is the guarantee. A model is used in exactly two places, and neither is
load-bearing: rca_narrate rephrases a verdict that was already computed and already cited, and an
agent front-end decides which tool to call. Remove both and the numbers are the same numbers.
That guard is static — it proves nothing can call a model. For a validation team the sentence they are asked to accept is the executed one, so it is executed:
A pinned in-repo dataset goes through availability, production counts, the Six Big Losses,
ISA-18.2 alarm load, control charts, the conservative baseline and the RCA copilot. Each result is
canonically encoded and digested; the suite runs twice in this process and once in each of two
fresh interpreters started at different PYTHONHASHSEED values — the arm that catches a set or
dict iteration order reaching a result, which a single run never can. The socket API raises
throughout, so a computation that reached for a device or a hostname fails here instead of quietly
working on a machine that happened to be online. Afterwards the run is asked what it pulled in:
a model library that was already loaded (an MCP server holds iaiops.core.llm for the opt-in
narration tool) is recorded, not judged — only what the suite itself imported can condemn it.
The record separates result (identical every run — the part to sign) from context (when and
where this run happened). Two good runs are not byte-identical records, and someone will diff
them, so the halves are named rather than mixed.
This is the form the claim has to take to be usable: not "our model is accurate", which is not
evidence in a GxP context, but a test case someone can write into an IQ/OQ protocol — remove the
model, block the network, re-run the standard dataset, compare the hash — execute, and sign.
verify_determinism is the same check from the MCP side; iaiops verify suite lists what it
covers without running it.
Short version: verified against real protocol libraries, containers and in-process servers — not yet against real plant gear. We grade evidence rather than saying "tested", because a real container round-trip and a synthetic fixture are not the same claim.
| Rung | What it means | Status |
|---|---|---|
| Real libraries / containers / in-process servers | OPC-UA (incl. cert Sign/SignAndEncrypt + A&C), Modbus-RTU over a socat PTY + pymodbus, BACnet/IP via bacpypes3 on a two-IP subnet, MTConnect against the Institute's own cppagent, IoTDB / TDengine live write→read, HART codec vs hart-protocol, PLCnext route via asyncua | ✅ |
| Mock-verified (protocol logic exercised, no real device) | Omron FINS, IO-Link, BAS (Metasys / Niagara), Ignition Gateway, EtherNet/IP PCCC, Sparkplug B, S7 / MC / SECS-GEM | ⚠️ |
| Real gear | physical RS-485 devices, EtherCAT slaves, live HART gateways, live HVAC / BAS / Ignition, real PLCs | zero, for every protocol |
Per-protocol evidence — including what each test does not cover — is in
docs/VERIFICATION-RECORD.md, one row per protocol, naming the test
behind each claim. Every 待核实 is hardware-gated, not forgotten — each one names the equipment that would settle it.
我们在找现场测试伙伴。 软件里能验证的我们都验证了(真实 in-process 服务器、真实协议库、Docker 容器 loopback)——剩下的 待核实 清单只有真设备能回答:物理 Modbus-RTU(RS-485)、EtherCAT 从站、HART 网关、在线 BACnet 楼宇设备、在线 Metasys/Niagara BAS 控制器、在线 Ignition 网关、国产 PLC(汇川/信捷)、真机 PLCnext、真实变电站 RTU/IED、欧姆龙 FINS 真机、IO-Link 主站。如果你是 OT 工程师、系统集成商或工厂团队,手上有任何这类设备:装上 iaiops,对你的设备跑一遍 iaiops doctor,把结果告诉我们。经你验证的设备会署名写进支持矩阵;现场反馈的问题我们优先分诊;功能可以通过 GitHub Issues/Discussions 直接共创。
We're looking for field-testing partners. Everything software-verifiable has been verified; what's left on the honest 待核实 list only real equipment can answer — physical Modbus-RTU (RS-485), EtherCAT slaves, HART gateways, live BACnet HVAC, live Metasys/Niagara BAS controllers, live Ignition gateway, domestic PLCs (Inovance/Xinje), live PLCnext, substation RTUs/IEDs, live Omron FINS PLCs, IO-Link masters. If you're an OT engineer, integrator, or factory team with access to any of these: install iaiops, run iaiops doctor against your gear, and tell us what happened. Verified-equipment reports get credited in the support matrix, field-reported issues get fast triage, and features are co-designed in the open via GitHub Issues/Discussions.
👉 参与入口 | Start here: open an issue with the protocol and device model in the title, or email zhouwei008@gmail.com. Either reaches a person, and a report gets answered against the current release.
| Protocol | Tool | Operation | R/W | risk_tier | Returns (key fields) |
|---|---|---|---|---|---|
| OPC-UA | opcua_server_info | server status | R | low | state, product_name, namespaces |
| OPC-UA | opcua_browse | browse node tree | R | low | [{node_id, browse_name, depth}] |
| OPC-UA | opcua_read_node | read one node | R | low | value, datatype, source_timestamp, good |
| OPC-UA | opcua_read_many | batch read | R | low | [{node_id, value, ...}] |
| OPC-UA | opcua_subscribe_sample | bounded sample | R | low | {collected, samples[]} |
| OPC-UA | opcua_read_alarms | alarm surfacing | R | low | {active_alarms[], active_count} |
| OPC-UA | opcua_read_history | Historical Access (HDA) | R | low | {supported, count, values[]} |
| OPC-UA | opcua_diagnose_connection | connection triage | R | low | {verdict, checks[]} |
| OPC-UA | opcua_discover_tags | tag auto-discovery → semantic asset model | R | low | {tag_count, assets[], naming_report} |
| OPC-UA | opcua_health_summary | threshold classify (was health_summary¹) | R | low | {overall, counts, offenders[]} |
| OPC-UA | opcua_anomaly_scan | stddev outliers (was anomaly_scan¹) | R | low | {mean, stddev, outliers[]} |
| Modbus | modbus_read_holding | FC03 | R | low | {raw_registers, decoded[]} |
| Modbus | modbus_read_input | FC04 | R | low | {raw_registers, decoded[]} |
| Modbus | modbus_read_coils | FC01 | R | low | {bits[]} |
| Modbus | modbus_read_discrete | FC02 | R | low | {bits[]} |
| Modbus | modbus_detect_byte_order | byte/word-order auto-detect | R | low | {best_order, candidates[]} |
| Modbus | modbus_list_templates | vendor register templates | R | low | {templates[]} |
| Modbus | modbus_apply_template | decode block via template | R | low | {values:{name: engineering_value}} |
| Modbus | modbus_health_summary | threshold classify | R | low | {overall, counts, offenders[]} |
| S7comm | s7_cpu_info | CPU id + run/stop | R | low | {cpu_status, cpu_info} |
| S7comm | s7_read_area | read DB/M/I/Q | R | low | {items:[{address, value}]} |
| S7comm | s7_read_db | read data block | R | low | {items:[{address, value}]} |
| S7comm | s7_read_many | batch addresses | R | low | {items:[{address, value}]} |
| S7comm | s7_write_db | write data block | W | high/MOC | {before, written, _undo_id} |
| Mitsubishi MC | mc_cpu_status | CPU type | R | low | {cpu_type, cpu_code} |
| Mitsubishi MC | mc_read_words | word devices | R | low | {words[]} |
| Mitsubishi MC | mc_read_bits | bit devices | R | low | {bits[]} |
| Mitsubishi MC | mc_read_many | random read | R | low | {words[], dwords[]} |
| Mitsubishi MC | mc_write_words | write words | W | high/MOC | {before, written, _undo_id} |
| Omron FINS | fins_cpu_info | controller data read (0501) | R | low | {controller_model, controller_version} |
| Omron FINS | fins_cpu_status | controller status (0601) | R | low | {run_mode, status} |
| Omron FINS | fins_read_words | memory-area word read (DM/CIO/W/H/A/EM) | R | low | {words[]} |
| Omron FINS | fins_read_bits | memory-area bit read | R | low | {bits[]} |
| Omron FINS | fins_read_many | batch reads | R | low | {items[]} |
| Omron FINS | fins_write_words | memory-area write | W | high/MOC | {before, written, _undo_id} |
| MTConnect | mtconnect_probe | device model | R | low | {devices:[{components:[{data_items}]}]} |
| MTConnect | mtconnect_current | latest values | R | low | {observations[]} |
| MTConnect | mtconnect_sample | bounded stream | R | low | {observations[]} |
| MTConnect | mtconnect_assets | assets | R | low | {assets[]} |
| MTConnect | mtconnect_oee_snapshot | OEE inputs | R | low | {availability, execution, verdict} |
| MQTT/Sparkplug | mqtt_read_topic | bounded read | R | low | {messages:[{topic, payload}]} |
| MQTT/Sparkplug | sparkplug_subscribe_sample | bounded SpB sample (full decode) | R | low | {samples:[{sparkplug, payload:{metrics[]}}], seq_gap_count} |
| MQTT/Sparkplug | sparkplug_decode_payload | decode raw SpB payload | R | low | {metrics:[{name, alias, datatype, value, is_historical}]} |
| MQTT/Sparkplug | sparkplug_node_list | node discovery + state | R | low | {nodes:[{group_id, edge_node_id, online, devices}], primary_hosts[]} |
| MQTT/Sparkplug | uns_browse | topic-tree browse | R | low | {topics[], tree{}} |
| MQTT/Sparkplug | uns_topic_audit | UNS naming + sprawl governance | R | low | {verdict, sprawl_findings, findings{casing_collisions[], scattered_leaves[], …}} |
| MQTT/Sparkplug | uns_schema_drift | Sparkplug schema-drift (baseline vs current) | R | low | {verdict (none/additive/breaking), node_changes[]} |
| MQTT/Sparkplug | uns_live_audit | live UNS audit (bounded broker sample) | R | low | {verdict, findings{}} |
| MQTT/Sparkplug | sparkplug_live_schema | live NBIRTH schema snapshot | R | low | {nodes[], metrics[]} |
| MQTT/Sparkplug | uns_live_drift | live drift vs stored baseline | R | low | {verdict, node_changes[]} |
| MQTT/Sparkplug | mqtt_publish | publish/command | W | high/MOC | {published_bytes, applied} |
| EtherNet/IP | eip_controller_info | Logix controller id | R | low | {controller:{vendor, product_name, revision, serial}} |
| EtherNet/IP | eip_list_tags | tag discovery | R | low | {tag_count, tags:[{name, data_type, structure}]} |
| EtherNet/IP | eip_read_tag | read one tag/array | R | low | {tag, value, type, good} |
| EtherNet/IP | eip_read_many | batch read | R | low | {items:[{tag, value, type}]} |
| EtherNet/IP | eip_write_tag | write tag | W | high/MOC | {before, written, _undo_id} |
| Diagnostics | diagnose_dataflow | localize no-data | R | low | {verdict, diagnosis, hops[]} |
| Diagnostics | alarm_bad_actors | ISA-18.2 flood | R | low | {flood_verdict, top_offenders[]} |
| Diagnostics | tag_health | offender ranking | R | low | {overall, offenders[]} |
| Diagnostics | historian_health | gap/flatline | R | low | {verdict, gaps[]} |
| Diagnostics | subscription_health | sequenced-feed loss/reorder/overload | R | low | {verdict, missed_count, overloaded_channels[]} |
| Diagnostics | downtime_root_cause | AI downtime RCA copilot (cited, advisory) | R | low | {verdict, primary_cause, hypotheses:[{cause, confidence, evidence[]}]} |
| Diagnostics | downtime_root_cause_live | RCA copilot that gathers its own live evidence | R | low | {…downtime_root_cause…, collected_evidence} |
| Diagnostics | learn_cause_weights | learn per-site RCA cause weights from labeled incidents | R | low | {cause_weights{}, rationale} |
| Diagnostics | data_quality_scorecard | fleet data-trust rollup | R | low | {fleet_score, fleet_status, issue_breakdown, worst_tags[], endpoints[]} |
| Diagnostics | data_quality_fleet_rollup | cross-endpoint fleet view | R | low | {fleet_score, endpoints[]} |
| Diagnostics | heartbeat_health | heartbeat/watchdog liveness | R | low | {alive, distinct_transitions, longest_stall_s, reason} |
| Alarm (ISA-18.2) | alarm_flood_analysis | flood episodes / chattering / stale / summary | R | low | {episodes[], chattering[], stale[], summary{}} |
| Alarm (ISA-18.2) | alarm_rationalization_worksheet | CSV-exportable rationalization rows | R | low | {rows[], csv_path?} |
| Baseline | baseline_learn | conservative change-log baseline (refuses thin history) | R | low | {band{p1,p99,median,mad} | insufficient_data} |
| Baseline | baseline_check | silent-by-default violation check | R | low | {status, violations[] (cited)} |
| Baseline | baseline_record_change | record operator change (restarts learning) | R | low | {recorded, change_point} |
| Baseline | baseline_status | no_baseline / learning / ok / violation | R | low | {status, window} |
| Historian | historian_query | read history back out of sqlite/TDengine/IoTDB | R | low | {rows[], truncated} |
| Historian | historian_coverage | per-tag row counts + first/last ts | R | low | {tags:[{tag, rows, first, last}]} |
| PLC program | plc_program_outline | structure of exported ST/AWL/L5X program | R | low | {blocks[], call_graph, timers[]} |
| PLC program | plc_program_xref | symbol/address cross-reference (cited lines) | R | low | {sites:[{kind, source_file, line, quote}]} |
| PLC program | plc_program_section | one named block's source (≤200 lines) | R | low | {text, source_file} |
| Export | export_data | export local store → CSV/SQLite/Parquet | R | low | {path, row_count, preview[]} |
| Analytics | oee_compute | OEE = A×P×Q | R | low | {availability, performance, quality, oee, oee_pct} |
| Analytics | downtime_events | stoppage detect + categorize | R | low | {event_count, total_downtime_s, by_category, events[]} |
| Analytics | oee_multidim | OEE machine×part×shift | R | low | {matrix[], worst_performers[], mean_oee} |
| Analytics | asset_inventory | active fingerprint | R | low | {assets:[{protocol, vendor, model, firmware, reachable}]} |
| Analytics | cross_protocol_asset_model | merge discovered tags into one asset model | R | low | {assets[], tag_count} |
| Analytics | adopt_alias_map / diff_alias_map | tag alias-map adopt/diff | R | low | {aliases{}, changes[]} |
| Analytics | monitor_changes | bounded change-of-value | R | low | {change_count, changes:[{value, previous, wall_clock}]} |
| EtherCAT | ethercat_master_state | master/WKC + slave count | R | low | {master_state, expected_working_counter, slaves_found, slaves_expected} |
| EtherCAT | ethercat_slaves | bus scan | R | low | {slave_count, slaves:[{index, name, vendor_id, product_code, state}]} |
| EtherCAT | ethercat_slave_info | slave detail | R | low | {sync_managers[], fmmus[], object_dictionary[], input_bytes} |
| EtherCAT | ethercat_read_sdo | CoE SDO upload | R | low | {index, byte_length, hex, as_uint} |
| EtherCAT | ethercat_read_pdo | input PDO snapshot | R | low | {working_counter, input_hex, input_byte_length} |
| EtherCAT | ethercat_write_sdo | CoE SDO download | W | high/MOC | {before, written, applied} |
| EtherCAT | ethercat_set_state | AL-state transition | W | high/MOC | {before, requested, reached, applied} |
| PROFINET | profinet_discover | DCP IdentifyAll (segment-wide) | R | low | {station_count, stations:[{name_of_station, mac, ip, vendor_id, device_roles[]}]} |
| PROFINET | profinet_identify_station | identify by name-of-station | R | low | {found, name_of_station, mac, ip, device_family} |
| PROFINET | profinet_station_params | targeted DCP Get (by MAC) | R | low | {found, name_of_station, ip, netmask, gateway} |
| PROFINET | profinet_asset_inventory | DCP asset register | R | low | {asset_count, io_controller_count, assets[]} |
| PROFINET | profinet_dcp_set | DCP Set (station name / IP suite) | W | high/MOC | {before, applied, _undo_id} |
| SECS/GEM | secsgem_equipment_status | GEM link + identity (S1F1/F2) | R | low | {communication_state, are_you_there} |
| SECS/GEM | secsgem_list_status_variables | SVID namelist (S1F11/F12) | R | low | {count, status_variables[]} |
| SECS/GEM | secsgem_read_status_variables | SVID values (S1F3/F4) | R | low | {svids, values[]} |
| SECS/GEM | secsgem_list_equipment_constants | ECID namelist (S2F29/F30) | R | low | {count, equipment_constants[]} |
| SECS/GEM | secsgem_read_equipment_constants | ECID values (S2F13/F14) | R | low | {ecids, values[]} |
| SECS/GEM | secsgem_list_alarms | alarm list (S5F5/F6) | R | low | {count, alarms[]} |
| SECS/GEM | secsgem_list_process_programs | PPID directory (S7F19/F20) | R | low | {count, process_programs[]} |
| BACnet (building) | bacnet_discover | Who-Is device discovery | R | low | {device_count, devices:[{device_id, address}]} |
| BACnet (building) | bacnet_object_list | a device's objects | R | low | {object_count, objects:[{object_type, instance}]} |
| BACnet (building) | bacnet_read_property | one object property | R | low | {object_type, instance, property, value} |
| BACnet (building) | bacnet_read_points | all present-values (HVAC snapshot) | R | low | {point_count, points:[{object_type, instance, present_value}]} |
| BACnet (building) | bacnet_cov_subscribe | bounded COV capture (always unsubscribes) | R | low | {notifications[], terminated_reason} |
| BACnet (building) | bacnet_read_trend_log | TrendLog readRange (bounded) | R | low | {records:[{timestamp, value}]} |
| BACnet (building) | bacnet_write_property | present-value write (priority) | W | high/MOC | {before, written, _undo_id} |
| HART-IP (process) | hart_device_identity | cmd 0 identity | R | low | {manufacturer, device_type, revision} |
| HART-IP (process) | hart_primary_variable | cmd 1 PV | R | low | {value, unit} |
| HART-IP (process) | hart_dynamic_variables | cmd 3 PV/SV/TV/QV + loop current | R | low | {variables[], loop_current} |
| HART-IP (process) | hart_burst_sample | bounded burst-variable sampling | R | low | {samples[]} |
| IO-Link | iolink_master_info | master identity | R | low | {vendor, product, serial} |
| IO-Link | iolink_ports | ≤32-port sweep (mode/status/device id) | R | low | {ports[]} |
| IO-Link | iolink_device_info | per-port device identity | R | low | {vendor_id, device_id, product_name} |
| IO-Link | iolink_read_pdin | process-data-in (raw hex + bytes) | R | low | {hex, bytes[]} |
| IO-Link | iolink_read_isdu | ISDU acyclic parameter read | R | low | {index, subindex, value} |
| IO-Link | iolink_scan | master + all connected devices | R | low | {master{}, devices[]} |
| BAS (Metasys/Niagara) | bas_point_list | supervisory point directory | R | low | {point_count, points:[{id, name, type}]} |
| BAS (Metasys/Niagara) | bas_point_read | read one supervisory point | R | low | {point, value, unit, status} |
| BAS (Metasys/Niagara) | bas_alarm_list | active controller alarms | R | low | {alarm_count, alarms:[{id, priority, state}]} |
| BAS (Metasys/Niagara) | bas_trend_read | trend/history samples (bounded) | R | low | {records:[{timestamp, value}]} |
| BAS (Metasys/Niagara) | bas_command | supervisory command (default-OFF; life-safety object denylist refuses fire/smoke/egress/pressurization before any I/O) | W | high/MOC | {before, written, _undo_id} |
| Ignition | ignition_gateway_status | Gateway + module health | R | low | {state, version, modules:[{name, state}]} |
| Ignition | ignition_tag_browse | tag-tree browse | R | low | {tags[], tree{}} |
| Ignition | ignition_tag_read | current tag values | R | low | {values:[{path, value, quality, timestamp}]} |
| Ignition | ignition_alarm_status | active alarms | R | low | {alarm_count, alarms:[{path, priority, state}]} |
| Ignition | ignition_tag_history | tag-history query (bounded) | R | low | {rows:[{path, timestamp, value}]} |
| 信创 / compliance | compliance_mapping | 《工控网络安全防护指南》↔ iaiops | R | low | {pillars[], status_summary, controls:[{pillar, status, gap}]} |
| 信创 / compliance | compliance_frameworks | 等保 2.0 + IEC 62443 FR1–6 crosswalk | R | low | {controls:[{crosswalk}]} |
| 信创 / compliance | compliance_dengbao_levels | 等保 二级 baseline vs 三级 增量 | R | low | {pillars:[{l2, l3_delta, status}]} |
| 信创 / compliance | compliance_report | deliverable compliance report (md/html) | R | low | {markdown | out_path} |
| 信创 / compliance | compliance_evidence_bundle | audit-evidence zip (hash-chain verified) | R | low | {bundle_path, manifest} |
| 信创 / historian | historian_push | push telemetry to sqlite/TDengine/IoTDB | R(→historian) | low | {sink, received, written, skipped_non_numeric} |
| Self | protocols_supported | capability map | R | low | {protocols[], diagnostics[], analytics[]} |
(The energy protocols — IEC-104 / DNP3 / IEC-61850 — moved to iaiops-energy in 0.8.0; their tool matrix lives in that repo.)
196 governed tools = 183 read + 10 MOC-gated device writes + historian_push (a write, to a historian rather than to a device: [WRITE][risk=low]) + the 2 deprecated aliases below. The device writes are (s7_write_db, mc_write_words, fins_write_words, mqtt_publish, eip_write_tag, ethercat_write_sdo, ethercat_set_state, profinet_dcp_set, bacnet_write_property, bas_command). The read side now includes two vendor-REST read-only layers above the field protocols — a BAS controller layer (Metasys/Niagara, building edition) and an Ignition Gateway MES/SCADA layer (factory edition). ¹ The 2 deprecated aliases are the two deprecated brain aliases health_summary / anomaly_scan, renamed to opcua_health_summary / opcua_anomaly_scan in 0.10.0 — the deprecated aliases are still registered and will be removed in a future release (target: 1.0.0). Read-only per-edition tools load ONLY under their edition (see per-edition tool modules below), so a bare protocol / single-edition surface is smaller than this line-wide total. The table above is representative, not exhaustive; run protocols_supported() (or iaiops protocols) for the live map.
opc.tcp:// via asyncua (sync facade). Security: anonymous + username/password, plus application-certificate message security (Sign / SignAndEncrypt) — set client_cert + client_key (+ optional server_cert) and the client opens a signed/encrypted secure channel (no cert ⇒ the anonymous / username path is unchanged). Validated end-to-end against an in-process asyncua server (tests/test_opcua_security.py) for Basic256Sha256 in both Sign and SignAndEncrypt modes: server_cert pinning and client-side server-cert auto-discovery are exercised, and the test asserts the negotiated policy URI + message-security mode on the live encrypted channel (plus a negative test that anonymous is refused by a secure-only server).endpoint_url, username (password encrypted), security_mode, security_policy; for cert security client_cert / client_key / optional server_cert (PEM or DER paths; aliases certfile / keyfile).opcua_alarm_events — bounded event subscription + ConditionRefresh, events carry the server's own timestamps (verified against an in-process asyncua server; third-party A&C servers 待核实). Untimed fallback: opcua_read_alarms browses alarm-like boolean nodes.待核实): cert-security interop with third-party / vendor servers (KEPServerEX / Prosys / Siemens / real PLCs), the other policies (Aes128Sha256RsaOaep / Aes256Sha256RsaPss / Basic128Rsa15 / Basic256), strict server-side certificate-trust enforcement, and cert-based user identity (X509 identity token, distinct from channel security).pymodbus (+ pyserial). Read function codes FC01 (coils), FC02 (discrete), FC03 (holding), FC04 (input). Write FCs (FC05/06/15/16) = not implemented (read-only).host, port (502), unit_id. RTU — transport: rtu, serial_port (e.g. /dev/ttyUSB0), baudrate, unit_id. Registers are untyped 16-bit words → decode hint (uint16/int16/uint32/int32/float32/raw); modbus_detect_byte_order auto-detects the byte/word order (AB/BA · ABCD/DCBA/BADC/CDAB) from hint values — pure logic, no extra device load.modbus_list_templates / modbus_apply_template): named register maps decoding a block into engineering values — energy meters (Eastron SDM630, Schneider PM5xxx, Carlo Gavazzi EM24), PV inverters (Huawei SUN2000, Growatt), Phoenix PLCnext process data, and water-industry templates (E+H Promag, Hach SC controller, generic dosing pump). Each template carries an explicit 待核实 caveat — no invented "verified" addresses.待核实.pyS7 (pure-Python, ISO-on-TCP / RFC1006 — no native libsnap7). S7-300/400/1200/1500 and compatible clones. Memory areas DB / M (merker) / I / Q. No protocol auth (CPU gates via "Permit access with PUT/GET").host, port (102), rack, slot (0/1 for 1200/1500; 0/2 common for 300/400).s7_write_db = high risk_tier, MOC, dry-run default, captures BEFORE value + undo.pymcprotocol — MC 3E frame (binary) only. 1E / 4E frames = not supported. PLC types Q / L / QnA / iQ-R / iQ-L. Devices: D/W/R (word), M/X/Y/B (bit).host, port (5007 default; set to the module's open MC port), plctype.mc_write_words = high/MOC/dry-run default, captures BEFORE + undo.iaiops[fins] extra pins nothing): 10-byte FINS header framing, FINS/UDP (default port 9600) and FINS/TCP (node-address handshake per Omron W342), SID matching, bounded response parsing, end-code table per W227/W342. Commands: 0101 memory-area read (words/bits over DM/CIO/W/H/A/EM), 0102 write, 0501 controller data read, 0601 controller status.host, port (9600), transport (udp default / tcp), FINS network/node/unit addressing.fins_write_words = high/MOC/dry-run default, captures BEFORE + undo; CLI double-confirm on --apply.tests/test_fins.py); live Omron PLC behaviour and banked-EM access stay 待核实.flavor: — iotcore (ifm IoT-Core POST envelope, default) and rest (plain-REST GET, Balluff/Turck-style). Reads: master identity, bounded ≤32-port sweep, per-port device identity, process-data-in (raw hex + bytes), ISDU acyclic parameter read. NO write tools. Bounded/size-capped HTTP (256 KiB response cap), schema-checked JSON with teaching errors. Reuses the MTConnect HTTP pin (iaiops[iolink] → requests).host/URL, flavor, timeout_s. protocol: iolink.tests/test_iolink.py); live master datapoint paths stay 待核实.transport: tcp, length-delimited framing) via an in-tree transport; the HART command codec is verified vs hart-protocol. Tools: hart_device_identity (cmd 0), hart_primary_variable (cmd 1), hart_dynamic_variables (cmd 3, PV/SV/TV/QV + loop current), hart_burst_sample (bounded sampling of burst-published variables). No write / device-specific commands exposed (OT-dangerous on live instruments).host (HART-IP server/gateway), port (5094), transport (udp default / tcp).待核实.requests + xml.etree), namespace-agnostic (parses MTConnect 1.x Devices/Streams/Assets schemas). Endpoints: /probe, /current, /sample, /assets. Read-only by specification. XML parsing is hardened (DTD/entity declarations rejected — XXE/billion-laughs defense).agent_url (e.g. http://host:5000).interval=); only bounded count= samples.paho-mqtt — MQTT 3.1.1 & 5. Sparkplug B topic convention spBv1.0/{group}/{type}/{edge}/[device] (NBIRTH/DBIRTH/NDATA/DDATA/NDEATH/DDEATH/STATE). TLS + username/password supported.sparkplug_b.proto generated module (depends only on protobuf). Per metric you get name, alias (resolved to its name via the BIRTH model), datatype (Int8…Int64/UInt…/Float/Double/Boolean/String/DateTime/Text/UUID/DataSet/Bytes/File/Template/PropertySet…), value, timestamp, and the is_historical / is_null flags. A birth/death + seq model tracks node/device online state (NBIRTH/DBIRTH ↔ NDEATH/DDEATH), builds the alias→name map from BIRTH, applies NDATA/DDATA by alias, and flags seq gaps / out-of-order. Primary-host awareness: STATE/<host_id> topics surface in sparkplug_node_list. sparkplug_decode_payload decodes a single raw payload (base64/hex) offline.host/broker, port (1883 / 8883 TLS), topic, use_tls, username (password encrypted).mqtt_publish = high/MOC/dry-run default. A transient publish has no automatic inverse (delivered is delivered); a retained one overwrites durable broker state, so it captures the prior retained payload and records an inverse.pycomm3 (pure-Python — no native deps). Tag-based, symbolic access: read/write tags by name (Conveyor.Speed, Array[3], Program:Main.X) and discover the controller's tag list at runtime (eip_list_tags, the headline feature). eip_controller_info reads the controller identity.host, slot (0 for CompactLogix; the CPU slot for a ControlLogix chassis), port (44818). protocol: ethernetip (alias eip).eip_write_tag = high risk_tier, MOC, dry-run default, captures BEFORE value + undo.Supported: a real EtherCAT master via pysoem (the Python binding for the SOEM C stack). CoE SDO read (ethercat_read_sdo, acyclic mailbox upload) + SDO write (ethercat_write_sdo, download), input PDO read (ethercat_read_pdo, one bounded cyclic snapshot), bus scan / slave enumeration (ethercat_slaves, ethercat_slave_info — identity, SM/FMMU mapping, object-dictionary summary), master/working-counter state (ethercat_master_state), and AL-state transitions INIT↔PREOP↔SAFEOP↔OP (ethercat_set_state).
HARD REQUIREMENTS (no way around them): Linux, root or CAP_NET_RAW, a dedicated NIC cabled to the bus, and real EtherCAT slave hardware. pysoem is an OPTIONAL extra: pip install iaiops[ethercat] — the base package installs and imports without it, and every EtherCAT tool then degrades to a teaching error (never crashes, never imports pysoem at module load).
NOT supported: no software simulator exists (unlike OPC-UA / Modbus) — EtherCAT is hardware-only and not testable in mock-only CI; macOS is unsupported. EoE / FoE / SoE mailbox protocols and full PDO-mapping decode/expansion = roadmap.
Connection params: nic (the dedicated interface name, e.g. eth1; alias interface), optional expected_slaves (a sanity check vs the bus scan). protocol: ethercat.
Operations matrix:
| Tool | Op | R/W | risk | Capture/notes |
|---|---|---|---|---|
ethercat_master_state | master + WKC state, slave count | R | low | expected vs found |
ethercat_slaves | bus scan / enumerate | R | low | index/vendor/product/rev/addr/AL-state |
ethercat_slave_info | one-slave detail | R | low | SM/FMMU + OD summary |
ethercat_read_sdo | CoE SDO upload | R | low | hex + uint interpretation |
ethercat_read_pdo | input PDO snapshot | R | low | single cycle, never loops |
ethercat_write_sdo | CoE SDO download | W | high/MOC | before-value (SDO read-back) + undo |
ethercat_set_state | AL-state transition | W | high/MOC | before-state + undo; can start/stop motion |
Write/state safety: ethercat_write_sdo (hex little-endian bytes) and ethercat_set_state are high risk_tier, MOC, dry-run by default, capture the BEFORE value/state for undo, and need a CLI double-confirm. Changing EtherCAT state can START or STOP machine motion — treat with extreme care. 未经授权勿对生产控制系统写入.
pnio-dcp — profinet_discover (DCP IdentifyAll: one broadcast surfaces every station on the segment — name-of-station, MAC, IP, vendor/device id, role — closer to passive discovery than a per-device fingerprint), profinet_identify_station (by name-of-station), profinet_station_params (targeted DCP Get by MAC → name + IP suite), and profinet_asset_inventory (a register with IO-controller vs IO-device role decoding).profinet_dcp_set re-addresses one station (name-of-station and/or IP suite, by MAC) — high risk_tier, MOC, dry-run default, captures the BEFORE addressing + undo descriptor. Re-addressing a live station can disrupt its IO connection.CAP_NET_RAW) on the NIC on the PROFINET subnet. pnio-dcp is an OPTIONAL extra: pip install iaiops[profinet] — the base package installs/imports without it, and every tool then degrades to a teaching error.host — THIS machine's IP on the PROFINET subnet (the DCP broadcast goes out on it). protocol: profinet.pnio-dcp DCP — not verified against live PROFINET devices yet.iaiops-energyThe energy vertical — IEC 60870-5-104 / DNP3 / IEC 61850 MMS read-only monitoring for substation RTUs/IEDs — moved to its own package in 0.8.0: iaiops-energy (pip install iaiops-energy), built on iaiops.core (shared governance / brain / runtime). Its protocol reference, support matrix, and validation status live in that repo.
The building vertical adds BACnet/IP (ASHRAE 135) — the dominant building-automation protocol for HVAC, lighting, metering, and facility plant. Install with pip install iaiops[building] and expose with IAIOPS_MCP=building (bundle: bacnet + modbus + opcua + iolink).
BAC0 over bacpypes3): bacnet_discover (Who-Is device discovery), bacnet_object_list (a device's objects), bacnet_read_property (one object property), bacnet_read_points (present-value of all analog/binary/multistate points — the HVAC snapshot), bacnet_cov_subscribe (bounded change-of-value capture — capped by count AND wall-clock, always unsubscribes), bacnet_read_trend_log (TrendLog buffered records via one bounded readRange). Config: host = THIS machine's BACnet/IP interface (ip or ip/mask) / port (47808).bacnet_write_property (present-value at a BACnet priority 1..16, or relinquish) = high risk_tier, MOC, dry-run default, BEFORE-value read-back + undo. Overriding a live building-control point can move real HVAC/plant.tests/test_bacnet_live.py); COV / trend-log / writes on live HVAC gear stay 待核实.IAIOPS_MCP=water (or iaiops-mcp-water, pip install iaiops[water]) exposes modbus + opcua + hart + the brain — the protocol set waterworks / wastewater plants actually run. Adds water-domain tag semantics (溶解氧 DO / ORP / 余氯 chlorine / 氨氮 ammonia / TSS/MLSS / 跨膜压差 TMP / UV / 加药 dosing / 曝气 aeration) and water-industry Modbus templates (E+H Promag, Hach SC controller, generic dosing pump — all with explicit 待核实 caveats).
IAIOPS_MCP=warehouse (or iaiops-mcp-warehouse, pip install iaiops[warehouse]) exposes eip + profinet + modbus + opcua + sparkplug + the brain — conveyor & sorter drives over EtherNet/IP (Rockwell) and Profinet (Siemens), VFD / energy meters over Modbus (conveyor_vfd / agv_battery templates), WMS/WCS gateways over OPC-UA, and AMR/IoT telemetry over MQTT-Sparkplug. Edition tools (read-only, advisory): line_bottleneck (Theory-of-Constraints throughput bottleneck across stations) + sortation_health. PdM (pdm_forecast), downtime_triage and OEE are reused as-is.
IAIOPS_MCP=clinical (or iaiops-mcp-clinical, pip install iaiops[clinical]) exposes bacnet + modbus + opcua + the brain — hospital facilities as a distinct patient-safety vertical over the building brain. Edition tools (read-only, advisory): isolation_room_check (负压/正压 isolation-room pressurization), medical_gas_check (medical-gas alarm-panel safety), or_environment_check (OR temperature / humidity / pressure envelope). BACnet BMS + Modbus gas-alarm panels + OPC-UA plant SCADA.
IAIOPS_MCP=pharma (or iaiops-mcp-pharma, pip install iaiops[pharma]) exposes bacnet + modbus + hart + opcua + the brain. No new protocol — that is the point: no field protocol is specific to pharma. Cleanrooms run BACnet, purified-water systems run Modbus and HART, filling and lyophilization run S7, DCS and bioreactors run OPC-UA, and all of it was already here. What pharma needed was semantics: the water edition's indicators are municipal (DO, ORP, chlorine, turbidity) and the clinical edition grades one room's pressure, where Annex 1 inspects the cascade.
Edition tools (read-only, advisory): cleanroom_pressure_cascade (EU GMP Annex 1, door by door — adjacency is declared, never inferred from a room list), cleanroom_particle_check, pharma_water_check (USP <645> stage-1 procedure: the non-temperature-compensated reading, the measured temperature rounded down to the tabulated step, and exceeding stage 1 reported as proceed to Stage 2 rather than as a failure).
No compendial limit tables are shipped. The particle limits, the stage-1 conductivity table and the TOC limit belong to the site's qualified specification at its compendial revision. A transcription nobody in this repository can verify would end up deciding whether a batch environment passed — and the error that hurts is the flattering one, since a limit set too loose reads as "in specification". Limits are passed in and cited back; anything not declared is reported no_limit / not_graded and named, never counted as passing. Known gaps are listed in the edition's skill: no PI historian connector, S7 without hardware verification, no GxP (Annex 11 / Part 11) crosswalk yet, and LIMS / QMS deliberately out of scope — they run REST and databases, not field protocols.
IAIOPS_MCP=renewables (or iaiops-mcp-renewables, pip install iaiops[renewables]) exposes modbus + opcua + sparkplug + the brain — PV inverters (SUN2000 / Growatt templates) + wind-turbine controllers over Modbus, OPC-UA plant SCADA, and MQTT-Sparkplug telemetry. Edition tool (read-only, advisory): pv_performance (PV string performance vs expectation). Device-level monitoring + PdM via baseline / RCA.
IAIOPS_MCP=plcnext (or iaiops-mcp-plcnext, pip install iaiops[plcnext]) exposes opcua + modbus + the brain — the Phoenix Contact PLCnext virtualized PLC reached over its built-in OPC-UA server (opc.tcp 4840, Arp.Plc.Eclr address space) + Modbus-TCP process-data server; no new connector. Route-verified in-process; live PLCnext hardware reads stay 待核实 (see How far it's actually been verified).
For 自主可控 / 信创 deployments — see docs/CHINA.md for the full guide.
pip install --no-index --find-links ./wheelhouse "iaiops[...]"; secrets stay local (encrypted store), no cloud KMS.historian_push, CLI iaiops historian push): write collected telemetry to TDengine (iaiops[tdengine]) or Apache IoTDB (iaiops[iotdb]) — domestic, controllable; we don't build our own store or bind InfluxDB. Data egress to the operator's own historian, not a control write.compliance_mapping, CLI iaiops compliance): an honest 《工控系统网络安全防护指南》 ↔ iaiops self-assessment across 分区隔离 / 可审计 / 双向认证 / 最小权限 / 数据保护 / 自主可控, with per-control status (addressed / partial / 待核实) and the named gap.oee_compute — OEE = Availability × Performance × Quality from production inputs (planned time, run time, ideal cycle, total/good counts). Each factor is reported raw + clamped to [0,1]; a capped performance >1.0 flags an optimistic ideal cycle.downtime_events — auto-detects running→stopped transitions in a {timestamp, state} series and produces stoppage events with durations, categorized (changeover / material / mechanical / quality / break / unknown, by keyword heuristics or a {state: category} override).oee_multidim — aggregates OEE across machine × part × shift (or any dimensions) from labelled records → the matrix + worst performers.mtconnect_oee_snapshot surfaces the live MTConnect availability/execution inputs that feed these.asset_inventory — for each configured (or named) endpoint, actively connects with our own protocol client and reads its identity call (S7 s7_cpu_info, EtherNet/IP eip_controller_info, OPC-UA server build info, Modbus Device Identification FC43/0x2B, Mitsubishi CPU type, MTConnect device model), aggregating vendor / model / firmware / serial / reachable / last_seen into an asset register.opcua_read_history — reads stored historical values for a node over a [start,end] ISO-8601 window via the server's HistoryRead service (asyncua read_raw_history), bounded by max_points (≤2000). Returns {supported:false, note} gracefully when the server does not historize the node (no crash). Read-only.monitor_changes — bounded deadband report: polls a point and returns only the value CHANGES (with timestamps), not every sample. Works over OPC-UA / Modbus / S7 / Mitsubishi MC / EtherNet-IP. Never an infinite loop — hard-capped by both duration_s (≤120) and max_changes (≤500). Read-only.baseline_learn/check/record_change/status, CLI iaiops baseline …) — a change-log baseline, explicitly NOT black-box anomaly detection: robust p1/p99 + median/MAD band over the local history, refuses thin history (<100 samples or <24h) with an explicit insufficient_data verdict, restarts at recorded operator changes, and is silent by default — a violation needs >3×MAD beyond the band AND ≥3 consecutive samples, and every violation cites its baseline window and offending samples.historian_query / historian_coverage, CLI iaiops historian query|coverage) — query history back out of the sqlite/TDengine/IoTDB sinks; an optional per-site historian: config block lets the RCA copilot pull the 2h pre-incident window as one more cited evidence class (strictly additive — without the config, RCA output is byte-identical, test-proven).plc_program_outline/xref/section, CLI iaiops program …) — structural extraction over exported program files (Siemens SCL/ST .scl/.st, AWL/STL .awl, Rockwell Studio 5000 .L5X — never a live PLC upload); every element carries source_file + line (rung for L5X) so the explaining agent must cite real locations. XXE-hardened, ≤5 MB, extension allowlist.alarm_flood_analysis / alarm_rationalization_worksheet, CLI iaiops diag alarm-flood|alarm-worksheet) — flood episodes (≥10 alarms/10 min), chattering, stale/standing (>24h), percent-time-in-flood vs target, and a CSV-exportable rationalization worksheet; over injected events or a live OPC-UA active-condition scan.iaiops export csv|sqlite|parquet (from the local SQLite sink; Parquet via iaiops[export]) / MCP export_data; iaiops metrics serve --port 9184 exposes Prometheus /metrics (latest tag values + counters, binds 127.0.0.1 by default) — Grafana recipe in docs/GRAFANA.md.iaiops compliance report (等保 2.0 L2/L3 status + IEC 62443 FR1–6 crosswalk + honest gap list, md/html) and iaiops compliance evidence (audit-evidence zip with hash-chain verification + manifest); MCP compliance_report / compliance_evidence_bundle. Onboarding aids, 非认证.downtime_triage) — composes alarm cascade + RCA verdict + PdM precursors into one triage and cross-checks whether the first-out alarm agrees with the diagnosed cause; advisory, cite-first (builds on the earlier alarm_cascade first-out reconstruction and pdm_forecast time-to-threshold early-warning).plc_program_visibility) — a risk/maintainability read over an exported ST/AWL/L5X program (size, block count, xref density, undocumented sections), never a live upload — pairs with the plc_program_outline/xref/section explainer.EDITION_MODULES in mcp_server/profiles.py) — a named edition can carry its own @mcp.tool group that loads only when that edition is selected — never for a bare protocol key and never in the always-on brain, so edition-specific tools stay off other surfaces and don't inflate the base. Every edition tool is read-only, cite-first, advisory:
line_bottleneck (Theory-of-Constraints throughput bottleneck) + sortation_healthisolation_room_check (负压隔离病房 pressurization) + medical_gas_check + or_environment_checkeconomizer_check (AHU economizer FDD) + zone_comfortcontrol_loop_health (PID oscillation/offset/saturation) + heat_exchanger_foulingspc_check (SPC control-chart rules) + defect_paretochangeover_analysis (SMED)disinfection_ct + water_quality_compliancepv_performance (PV string performance)skills/iaiops) plus ten per-edition skills (iaiops-fab / iaiops-factory / iaiops-process / iaiops-building / iaiops-water / iaiops-warehouse / iaiops-clinical / iaiops-pharma / iaiops-renewables / iaiops-plcnext) that route an agent to the right MCP server and document the tool surface.iaiops is designed to ride on a hardened, centrally-managed edge host as a portable, governed edge application — not to own the host or the fleet manager. It maps naturally onto the Margo edge-interoperability roles: the host/device is the immutable edge OS, a compliant orchestrator places workloads by desired-state, and iaiops is the OT-domain application — read-first tap + cross-protocol RCA, exposed as governed MCP tools, with an optional on-box LLM brain for a fully air-gapped diagnostic path (data never leaves the plant).
Honest status: iaiops is a natural Margo edge application but is NOT Margo-compliant yet — a container image + application description + a published conformance-toolkit result are roadmap
⏳(see docs/MARGO-ALIGNMENT.md anddocs/ROADMAP.md). No material claims Margo-compliant until that test result exists.
A container + application-description skeleton lives in deploy/margo/
(hardened Dockerfile · compose · 待核实-marked app descriptor); per-host distribution overlays
that reuse it live under deploy/ (one folder per candidate edge host).
Protocol client libraries are optional extras — install only the 1–2 protocols a site actually runs (every protocol library is imported lazily; the base package installs and imports without any of them, and a call to a not-installed protocol returns a teaching error pointing at the right extra):
Protocol extras: opcua · modbus · s7 · mc · fins (stdlib — pins nothing) · eip · mtconnect · sparkplug · secsgem · ethercat · profinet · bacnet · hart · iolink · bas (BAS supervisory REST — reuses the mtconnect HTTP pin) · ignition (Ignition Gateway read layer — reuses the mtconnect HTTP pin) · plus tdengine · iotdb · influxdb (historian sinks) · nats (stream egress) · ollama (on-box LLM narration) · export (Parquet) · all (every pip-installable connector).
Adapter belt (
docs/ADAPTERS.md): iaiops is a small neutral core (ingress → normalize/govern/RCA → egress) with pluggable, lazily-imported adapters — bind no store/bus/host/model, install only what a site runs. The RCA core is deterministic + cited, not a black box (docs/RCA.md); footprint is small by design (docs/FOOTPRINT.md).
Edition bundles (match the same-named IAIOPS_MCP profiles — install the protocols a vertical runs):
fab (secsgem + opcua + s7 + modbus) · factory (the discrete-manufacturing set: opcua + modbus + s7 + mc + fins + eip + mtconnect + sparkplug + ethercat + profinet + iolink + ignition) · process (opcua + modbus + hart) · building (bacnet + modbus + opcua + iolink + bas) · water (modbus + opcua + hart) · warehouse (仓储/物料搬运: eip + profinet + modbus + opcua + sparkplug) · clinical (医疗设施: bacnet + modbus + opcua) · renewables (光伏/风电: modbus + opcua + sparkplug — PV inverters (SUN2000/Growatt) + wind turbines + plant SCADA; device-level monitoring + PdM via baseline/RCA) · plcnext (opcua + modbus). The grid/substation energy bundle (IEC-104/DNP3/61850) ships in iaiops-energy.
Secrets (per-endpoint passwords, MQTT credentials) are never stored in plaintext — they live in ~/.iaiops/secrets.enc (Fernet + scrypt). Export IAIOPS_MASTER_PASSWORD so the MCP server/CLI can unlock non-interactively:
~/.iaiops/config.yaml (one block per protocol)iaiops init walkthrough (per protocol)(MQTT prompts add TLS/topic/username; MTConnect prompts for agent_url; EtherCAT prompts for the nic + expected_slaves and warns about the Linux/root/NIC/optional-extra requirement; OPC-UA/MQTT prompt for a hidden password stored encrypted.)
asyncua demo server (the test suite runs a real in-process one).pymodbus server simulator.:102.mosquitto broker (+ a Sparkplug edge for SpB topics).tests/test_fins.py) or a spare CP/CJ PLC.tests/test_iolink.py, both JSON dialects) or any ifm/Balluff/Turck master on the bench.CAP_NET_RAW, on a dedicated NIC wired to real slaves (e.g. a Beckhoff EK1100 coupler + EL terminals). iaiops doctor reports a clear "needs Linux/root/NIC/pysoem" status off the bus rather than failing.Every other command needs an endpoint you already configured. scan answers the
question that comes first. It has no full-port mode, no raw sockets, no
half-open SYNs, and no write path of any kind; what it may touch is a fixed
industrial port allowlist, and how fast is capped by a ceiling the caller cannot
raise.
scan plan puts nothing on the wire. It prints every host and port that
would be touched, every class of packet that would be sent, the worst-case
duration, and the explicit list of what this tool never does — so you can run it
against a network before you have permission to scan it, and hand the output to
whoever grants that permission. scan run shows the same preview and asks once
before it sends anything (--yes to skip).
Postures run from passive (reads the local ARP cache, emits nothing at all) to
legacy-safe (reachability only, one host at a time, five connects a second —
for 1990s controllers where even a well-formed identify request is a risk).
standard and deep refuse to run without a recorded sign-off.
The HTML report is self-contained: no fonts, scripts, styles or images from anywhere, and no network request when opened. Its first section is what the scan touched — per-class emission counts, including requests that failed — followed by the list of things it never does. The device table comes after that.
scan answers what is out there. These answer what can I do with it, and what
is the number. Each step is a real command; nothing here is a roadmap item.
What this installation can run today, and for each thing it cannot, the specific input that is missing — ranked by how much supplying it would unlock. It touches no device, so you can run it against a site you have not been authorised to probe, which is the site that most needs the answer. It reports gaps; it never fills one in (§9.4/D16 — a guessed production counter yields a plausible-looking OEE, which is worse than an error).
A bounded assessment run — capped at 14 days, and the operator must state the end. There is deliberately no run-forever mode: a resident process on an OT network needs change management, a laptop running for a week does not (D21). Every run records the windows it could not see, so a gap is never silently readable as a stoppage.
To get an OEE out of it, three tags have to be declared — which value means running, which register counts parts, and (for Quality) which counts good ones:
running_when is declared, never inferred. On the 0=stopped 1=idle 2=running 3=fault status word most PLCs expose, "any non-zero means running"
counts three states of four as production.
Scope the measurement with --since / --until, together. Without them the
period is everything the store holds for that endpoint — fine for one assessment
run, wrong the moment there are two: the idle weeks between a March run and an
August one become one enormous blind span and neither run can be measured.
The window is charged for in full. The parts of it that were never sampled — before the first sample and after the last — count as blind exactly like a gap in the middle, so narrowing the question to the minutes that happen to have data cannot raise your coverage. Ask about a shift you observed two hours of, and the answer is 25% coverage and a refusal, not "100% of what we saw".
Availability measured over the time the collector could see, with blind windows excluded rather than counted as downtime; plus Performance, Quality and the Six Big Losses. Each factor is reported only when its inputs were declared — a partial OEE that names what is missing beats a whole one with a guess inside it.
--report writes one self-contained HTML file — no fonts, scripts, styles or
images from anywhere, and no network request when opened, so it works on an
air-gapped laptop and survives being forwarded as an attachment. Its first
section is what the measurement could see (coverage, blind time, sample cadence),
before the number, and it carries the row a sales deck usually leaves out: what
had to be declared to produce each figure, and what is still missing.
The label is a by-product of work already being done: the audit trail already recorded that someone ran a write four minutes after the line stopped, so the case shows it. Confirmation is one choice from a fixed vocabulary, never free text, and a dismissal is a label too. Whether an answer counts as independent is derived from whether the tool had suggested it — the answerer cannot claim it.
readiness answers which scenarios this site can run. This answers the next
question down: if something stopped tomorrow, how far could we actually get?
Eight evidence steps — define the incident, collect the evidence, normalize and check it, compress and rank, correlate the timeline, test the hypotheses, check against known mechanisms, conclude and close. For each one it cannot walk, it says whether that is something you have not supplied (with the command that would) or something this product cannot express at all. Those two send a person to very different places.
The same eight steps over a real past window, persisted so it can be re-read and advanced later. No device is contacted — the window is already over, and its evidence is whatever was collected at the time.
Any of the three writes the forwardable version — one self-contained HTML file that opens on an air-gapped laptop in a plant office:
Unlike oee measure --report, which refuses to write a file for a refused
measurement, this one writes for a blocked investigation on purpose. An OEE
report is a number, so a file existing at all claims one was measured. This
report's content is how far this got and what each step still needs — which
makes the blocked case the one most worth handing over, and for a site nobody has
instrumented yet it is the whole deliverable. What it will not do is let a
blocked investigation look finished: the headline is always the walk (2 / 8),
never a conclusion, and no step's own words appear above it.
Everything on the path existed; nothing joined it. A site could scan forty devices and then retype all forty by hand, and nothing anywhere said which of the six commands came next.
draft writes nothing into config.yaml — you merge it, exactly as with
tags apply. What it emits is constrained on purpose:
plctype, whether an
OPC-UA server advertises an unsecured endpoint or will need credentials.tags: comes out empty. A scan finds devices; it establishes nothing about
what their data means. That step is the next one, below.status is derived from the store and config.yaml every time. There is no
onboarding state file, so a hand-edited config or a restored backup still gets a
true answer rather than a remembered one — and a step that is genuinely done
stays done even if you did the steps out of order.
For the point-list step it names the command for your protocol: opcua browse, eip tags, mtconnect probe, iolink ports, mqtt browse, bacnet objects, ethercat slaves, hart dynamic, or modbus templates. Where there
genuinely is nothing to ask — an S7 CPU exposes no symbol table on the wire, and
MELSEC and Omron memory carry none either — it says so in that protocol's own
terms and tells you where the addresses do come from, rather than one sentence
covering everything that is not OPC-UA.
The one thing this product refuses to infer, and until now the only way to supply
it was hand-editing role: in a config file — which is exactly what stops working
at a hundred rows.
The role column comes out empty even next to a tag called
GoodPartsCounter. A name is not a declaration; plenty of plants have one that
counts something else, and a wrong production counter yields a plausible OEE —
worse than an error (D16).
apply emits the patch rather than writing it. config.yaml stays the single
source of truth: oee measure reads roles off the config tag objects, so a
parallel store would let readiness call the mapping met while oee measure
still could not run. A run_state with no running_when is refused here, for
the same reason MonitorTag refuses it — "anything non-zero" counts idle and
fault as production.
Or tick it through in a page instead of a spreadsheet:
HLD §13.9's App front end, delivered as a file rather than a served app. A
localhost server inside an OT box has to answer which address it binds and who
authenticates — and since every declaration here requires --by, a page with no
identity cannot record who confirmed a tag, which is the one thing this step
exists to capture. So the page collects, and the author is supplied at apply.
The page re-implements no refusal. run_state needing running_when, a ref
having to be monitored, a role claimed twice — reproducing any of those in
JavaScript is how they drift from the ones that actually gate the config, and a
page that says "looks fine" while apply refuses is worse than no page. It ships
script (it is a form) but makes no network request: it works with the cable
out.
There is deliberately no MCP tool for this. An agent filling in the role column is precisely the guess D16 exists to forbid.
Two of the steps need something a person has to state:
The second axis of root-cause analysis. With time alone, an upstream stoppage produces a string of equally-confident downstream false causes — on a line, downstream co-occurrence is guaranteed whatever the cause. That guarantee is why this is declared and not inferred (D25). Without it the timeline still runs; it degrades to a single asset and says so.
A fault-mechanism library, shaped by ISO 14224: the failure mode (what you saw), the mechanism (what to go and check) and the cause (what to fix) stay separate, because they answer different questions. Entries attach to the seven taxonomy causes; they never add new ones — past roughly forty codes, two operators stop picking the same one.
It may exclude and never confirm. A mechanism that cannot apply to this equipment rules the candidate out, which is the strong move a ranker cannot make:
And a cause the library has never heard of reports nothing known — never "no objection". A knowledge base that knows nothing about something has not cleared it.
A control program is a controlled document, and the usual way an undocumented change to one gets noticed is that somebody remembers. Record the version you consider approved, then ask a later export whether anything moved:
The snapshot stores the file's SHA-256 plus a per-block structural fingerprint — name/kind/language, declared variables, calls, branch conditions, timers — and deliberately excludes line numbers, comments and block order, so adding one comment at the top of a file does not report the whole program as changed. What is stored on disk is block names, hashes and counts; never a declaration, a source line or a comment, so the baseline store is not a second copy of your program.
Three verdicts, and each word is load-bearing:
| Verdict | Means |
|---|---|
identical | The same SHA-256. Nothing else earns the word. |
logic_changed | The extracted structure differs — reported per block, naming which of variables / calls / branches / timers_counters moved. |
changed_outside_extracted_structure | The bytes differ and every block fingerprint matched. |
That third one is the honest one. It is usually comments or formatting — but these parsers extract structure, they do not parse a grammar, so a real change inside a construct they do not model looks identical from here. Calling it "documentation only" would be the comfortable reading of evidence that does not support it, so it is not called that: line and comment counts are reported beside it and the verdict still says look. A drift report is a reason to read the diff, never a clearance.
iaiops program history lists what is tracked; iaiops program compare <name> <before> <after>
diffs two stored snapshots. Deleting history is iaiops program forget and is CLI-only — an
agent should not be one call away from removing change-control evidence. Nothing is pruned
automatically. The name (not the path) is the identity, because the export directory changes every
time somebody opens the engineering station and the program does not; absent --name the file stem
is used and the output says so. No device is touched at any point — this reads a file a person
exported.
--apply)s7_read_db:
s7_write_db (dry-run):
mtconnect_oee_snapshot:
eip_read_tag:
eip_write_tag (dry-run):
ethercat_read_sdo (CoE SDO upload):
ethercat_set_state (dry-run; can start/stop motion):
sparkplug_decode_payload (full SpB metric decode):
oee_compute:
asset_inventory (active fingerprint):
diagnose_dataflow(endpoint="line1", ref="ns=2;i=5", freshness_threshold_s=30):
alarm_bad_actors(events=[…]):
tag_health(tags=[…]):
downtime_root_cause correlates whatever evidence you can hand over — alarm
events, tag samples, a diagnose_dataflow verdict, a machine-state series —
around an incident window and returns an evidence-cited, advisory verdict.
Read-first: it proposes a human-approved, MOC-gated, undoable action and executes
nothing. Anti-hallucination by design — it cites only signals actually present in
the input, weights them by temporal proximity to onset (a cause precedes its
effect), and downgrades to insufficient_evidence (with a recommended_next_data
list) rather than guessing when evidence is thin.
downtime_root_cause(window={"start":"2026-06-28T10:00:00Z","asset":"line1"}, alarms=[{"source":"M1_DRIVE","timestamp":"2026-06-28T09:59:52Z","message":"motor overload trip"}], tags=[{"ref":"DRV1.Torque","samples":[10,11,99,99],"alarm_high":80}], dataflow={"verdict":"healthy"}):
The same copilot is on the CLI: iaiops diag rca --input bundle.json where the
bundle is {window, alarms?, tags?, dataflow?, state_series?}.
Let it gather its own evidence. downtime_root_cause_live (CLI iaiops diag rca-live) takes just an endpoint + window + the refs to look at, then pulls the
evidence itself — a cross-protocol diagnose_dataflow probe, a short sampled
series per ref (so flatline / bad-quality / anomaly surface via tag_health),
and active OPC-UA conditions — before running the same advisory, read-only copilot.
The gathered bundle is echoed back under collected_evidence (no hidden inputs):
Two more pure-analysis layers — fully testable without live gear, and they feed the RCA copilot.
data_quality_scorecard (CLI iaiops diag dataquality) — a fleet data-TRUST rollup: scores each tag 0-100 on whether its data can be believed — staleness, dead heartbeat (first-class), bad-quality, flatline, gaps, anomaly — then rolls up per endpoint and across the fleet with an issue breakdown and ranked worst offenders. Distinct from process health: it asks "can I trust this number," not "is this number alarming." heartbeat_health (CLI iaiops diag heartbeat) is the standalone watchdog-liveness check (a flatlined heartbeat = dead upstream even when comms look fine).uns_topic_audit (CLI iaiops mqtt uns-audit) — governs a UNS topic tree: naming conformance (allowed roots / min depth) + topic sprawl (casing collisions of the same logical name, leaf metrics scattered under many parents, depth outliers, duplicates) → a clean/minor/sprawling verdict. uns_schema_drift (CLI iaiops mqtt uns-drift) — compares two Sparkplug NBIRTH-style snapshots and classifies the change none / additive / breaking (a metric removed or its datatype changed). Positions the UNS as a governable neutral data source, not just a broker.Menu — expose only the protocols a site runs. A fab usually runs 1–2 protocols;
exposing all 14 floods the model with tools it can't use. Set IAIOPS_MCP to a
comma-list of protocols and/or a named profile. There is no default (since
0.10.0): a bare iaiops-mcp prints the selection menu (profiles, protocol keys,
tool counts) to stderr and exits 2 instead of silently exposing 100+ tools. The
cross-protocol brain (OEE / downtime / diagnostics / asset / analysis) is included
by default with every selection.
Named entry-point sugar. For the common single-protocol / single-edition case there is a pre-scoped console script per protocol and per named profile — no env var to set. Each is a thin shim over the same server:
Multi-process sites — 1 brain MCP + N protocol MCPs. Running several protocol
servers side by side (e.g. iaiops-mcp-opcua + iaiops-mcp-modbus) would duplicate
the ~30 brain tools in every server. Instead run one dedicated iaiops-mcp-brain
and set IAIOPS_MCP_NO_BRAIN=1 on the protocol servers to strip the brain from
them — the protocols_supported discovery tool stays exposed everywhere:
Write authorisation is not the tap's job. iaiops does not withhold write
tools behind a server switch. Whether a write is allowed is the caller's
decision — the agent's judgement or account/permission management — and the tap's
job is to make that write accurate, efficient, and un-bypassably audited.
Every call, read or write, on either front-end (MCP tool and iaiops CLI),
goes through @governed_tool and leaves a row in ~/.iaiops/audit.db. Writes are
additionally HIGH risk_tier and MOC-gated (dry-run + double confirmation + undo
capture + a recorded approver). protocols_supported reports this posture so the
model is told the rules rather than left to infer them.
Since 0.20.3 that promise is held by contract tests over the real tool
surface rather than by synthetic stand-ins: every one of the ten high-risk writes
is driven end to end and must be denied without an approver with the connector
never reached — "it raised" only proves an exception, not that nothing reached the
device. Two things those tests exposed on the way in: a call that failed was
audited as ok (tools return the canonical {error, hint} envelope rather than
raising, so the governance wrapper saw an ordinary return), which also told the
pattern circuit breaker "success" on every failure; and the runaway guard, blind to
a caller retrying a denial forever, let 500 identical denied writes through a
ceiling of 10. Both are fixed and pinned.
Sealed sites — make the data-shipping tools cease to exist.
IAIOPS_NO_EGRESS=1 removes every tool whose job is to transmit local or plant
data to a destination the caller names, at registration time so a weak or
prompt-injected model cannot call what it cannot see:
Withheld: stream_publish, stream_publish_event (NATS message bus),
uns_publish (MQTT broker / Unified Namespace),
historian_push (external TSDB), mqtt_publish (broker), rca_narrate (POSTs
the RCA verdict — plant tags, values and citations — to a caller-supplied model
base_url). This is a data-exfiltration / airgap axis, not read/write
authorisation: historian_push is risk_level="low" — it changes no plant state
— yet it ships telemetry off-box, so this switch withholds it. It gates MCP tools
only, and is not a firewall (reads still open sockets to plant devices).
protocols_supported reports each posture independently.
Scope, stated plainly — this is not a firewall:
iaiops audit forward (SIEM) is a CLI path no
registry gate can reach; block it at the host if the box must be sealed.export_data and
compliance_evidence_bundle stay exposed. The bytes never leave the box;
getting them off it afterwards is a host-level concern.iaiops-mcp server only (including its
per-protocol / per-edition entry-point shims). iaiops-energy-mcp is a
separate server in a separate package and does not honour them yet — it
mirrors in the base brain/compliance tools, so IAIOPS_NO_EGRESS=1 there
still leaves historian_push, rca_narrate, stream_publish and
stream_publish_event exposed on iaiops-energy 0.1.6 and earlier.
Fixed in iaiops-energy 0.1.7, which pins iaiops>=0.17 for exactly
this reason. Said out loud because a switch believed to be on is worse than
one known to be absent.Named profiles: all · brain · fab · factory · process · building ·
plcnext · water · renewables · warehouse · clinical. In an MCP client (e.g. Claude Desktop) set IAIOPS_MCP per
server entry — or point the entry straight at the matching iaiops-mcp-<name>
script — one entry per site/line, each a lean single- or dual-protocol server.
historian_push writes to a historian rather than to a device. The 10 write/command tools (s7_write_db, mc_write_words, fins_write_words, mqtt_publish, eip_write_tag, ethercat_write_sdo, ethercat_set_state, profinet_dcp_set, bacnet_write_property, bas_command) are OT-dangerous: governed at high risk_tier, off by default (dry-run), require a double-confirm in the CLI, and a recorded approver (one-shot iaiops approve tokens; with no risk_tiers configured, high/critical operations default to the dual tier) — MOC discipline. All ten declare an undo (no exemptions since 0.20.3); a successful write captures the BEFORE value/state and registers an inverse descriptor. The inverse honestly reports "none" where none exists — a transient (retain=False) mqtt_publish cannot be unsent, and ethercat_set_state's +ERR/NONE/BOOT are not cleanly re-requestable AL-states. An undo that over-promises is worse than none, because someone will replay it onto live equipment. ethercat_set_state can START or STOP machine motion. 未经授权勿对生产控制系统写入.iaiops CLI, runs through @governed_tool and leaves a row in ~/.iaiops/audit.db. Writes are additionally high risk_tier, MOC-gated, and undo-captured (see above). High/critical calls fail closed when the audit DB cannot be written.IAIOPS_NO_EGRESS=1 withholds the 6 tools that ship data off-box (stream_publish, stream_publish_event, uns_publish, historian_push, mqtt_publish, rca_narrate), fail-closed, for airgap/sealed-box deployments. This is a data-exfiltration axis, not authorisation — historian_push is low-risk (it changes nothing) yet pushes telemetry to an external TSDB, so this switch withholds it. Which tools count is derived from @governed_tool(egress=True) metadata and guarded by an AST scan in CI, so the next egress tool cannot silently escape the gate.~/.iaiops/audit.db, SHA-256 hash-chained rows + iaiops audit verify; audit fails closed for high/critical writes), token/call budget + runaway breaker, risk-tier gate (policy engine fails closed on a broken rules.yaml), undo recording. The MCP server refuses to start if any registered tool lacks the governance marker.Four items that used to sit here had in fact shipped — including two this same
README already listed as verified, three sections above. Listing built work as
future work is the same defect as claiming unbuilt work, so they are gone:
EtherNet/IP PCCC (PLC-5/SLC-500) and Micro800, passive asset discovery
(iaiops scan --posture passive, ARP-cache only, emits nothing),
OPC-UA certificate security, and MTConnect streaming long-poll.
What is genuinely open:
pysoem extra).Missing a protocol, device, or feature? 缺功能提 issue/PR 欢迎留言 — open a GitHub issue or PR.
MIT © wei