- NVIDIA raised Jetson prices on 22 July, hardest on the memory-rich modules a 7B model needs. Any edge rollout costed before then is priced wrong.
- Rockwell named a 9B model for its HMI panels. Its OptixPanels ship with 1 to 4 GB of RAM and a 2.3 TOPS NPU. The arithmetic does not close.
- Measured: an 8B model on the mainstream edge board runs at 9.9 tokens a second with 7.66 seconds to first token. Interactive work needs something smaller.
Your automation vendor has offered to run a language model in the cabinet. Nothing leaves the plant, no cloud account, no argument with IT about what is permitted out through the DMZ. The proposal probably carries a hardware price set before July. Before the question of whether it earns its place, there is a smaller one that settles it, answered in gigabytes and watts.
Rockwell announced on 13 November 2025 that it would bring NVIDIA's Nemotron-Nano-9B-v2 into FactoryTalk Design Studio, naming "HMI panels" and "air-gapped environments" among the deployment targets. No general availability has been announced since.
Now read Rockwell's own OptixPanel technical data. The Compact carries 1 GB of RAM, the Standard 4 GB, both passively cooled and rated 0 to 50 °C, with the i.MX 8M Plus NPU rated up to 2.3 TOPS.
NVIDIA's technical report says the model was compressed so that inference at 128k tokens fits on a single A10G, 22 GiB at bfloat16, using 19.66 GiB in practice. The community four-bit build is a 6.53 GB weight file. That does not go into four gigabytes. Whatever "HMI panels" means there, it is not the panels in the current catalogue, and the announcement does not say what it does mean.
The vertical industrial model was supposed to be the defensible thing: fine-tune on maintenance logs, batch records and alarm histories until the model knows a compressor the way a general model never could. Two forces are eroding that. General-purpose models keep absorbing the reasoning a vertical model was meant to supply, and they improve on a release cadence no industrial vendor can match. And the data required to fine-tune is going the other way, into stacks whose commercial logic is to hold it rather than release it for training. A fine-tune needs volume and variety of representative data, and the consolidation of the last eighteen months makes that harder to assemble.
So the model in the cabinet will be a general one, squeezed until it fits.
What the cabinet will hold
An independent benchmark published on 8 June 2026 ran heavily quantised models on a Jetson Orin Nano Super, the 8 GB board at the mainstream of edge hardware. At one-bit quantisation, where an 8B model compresses to 1.1 GB, it managed 9.9 tokens per second in the 15 W mode and took 7.66 seconds to return a first token. At 25 W it reached 14 tokens per second and 5.5 seconds, drawing a measured 7.73 W and peaking at 70.4 °C. A 1.7B model gave 34.7 tokens per second with 1.5 seconds to first token.
Then the part a cabinet makes decisive. An academic study published on 7 June 2026 ran one four-bit 1.5B model across four devices under sustained load. The laptop GPU held 131.7 tokens per second flat. The phone settled 41.5 per cent below its peak within a few iterations. The dedicated low-power accelerator held 6.914 tokens per second with a coefficient of variation of 0.04 per cent and no throttling at all. Anything with a boost clock hands its performance back when the load does not stop, and a control cabinet is sustained load with no fan by definition.
Read across those and the working ceiling for a passively cooled cabinet this year is roughly a 2B model at conversational speed, or an 8B model at ten tokens a second with a fan and a real power budget. That is inference across benchmarks on different hardware rather than a published figure, and it will move. Siemens' SIMATIC IPC BX-59A, its NVIDIA-accelerated box PC, takes 8 to 64 GB and draws 130 to 190 W at base. That is the honest shape of a machine holding a useful model: an industrial PC with a fan, not a controller and not a panel.
What it costs now
On 22 July 2026 NVIDIA raised Jetson prices across the range. The Orin NX 16 GB went from $599 to $999, the AGX Orin 32 GB doubled to $1,799, and the 64 GB went from $1,599 to $2,999. The memory-rich modules, the only ones that hold a 7B model, rose hardest.
The board side reports the same pressure. VersaLogic's supply brief through August has suppliers telling customers to plan for DRAM increases of 10 to 20 per cent per month to year end, lead times of 8 to 52 weeks, quote validity cut to 72 hours, and memory bought only against firm purchase orders. Its own hedge is that constraints are "likely to persist into 2027."
There is also no published case of a small local model doing plant work at a quality an operator signed off on. Not fault triage, not alarm summarisation. The vendors have announcements, the academics have benchmarks, and the gap between them is where the capital goes.
Three things before the next request goes up. Decide per task whether the answer has to arrive while someone is standing at the machine, because that call sets the model size and the model size sets the bill of materials. Stop speccing a node per line: consolidate the memory-heavy boxes into a handful per plant, with thin gateways wherever a microcontroller will do. And re-price anything approved before July, because at forty-week lead times the compute is now the critical path, and a 72-hour quote does not survive a six-week approval cycle.
