Redefining Technology

Manufacturing (Automotive)AI Implementation & Best Practices

Reinforcement learning in automotive plants: what a learned policy is allowed to control

Reinforcement learning is a method for training a policy that chooses actions in sequence and learns from the consequences they produce, rather than from labelled examples. In an automotive plant it fits a narrow set of repeated, measurable, bounded decisions — and how much authority that policy holds is a governance question long before a modelling one.

Automotive body shop with a sequencing decision point between the buffer bank and the paint line, control screens showing proposed and executed actions
Manufacturing (Automotive) · AI Implementation & Best Practices

Key takeaways

  1. Reinforcement learning fits sequential plant decisions whose consequence arrives later and whose outcome the plant already records — paint-shop colour sequencing, buffer release between shops, AMR dispatch, utilities setpoints. Perception and prediction problems are not RL problems, and treating them as such is the most common category error in plant AI.
  2. Maturity here is control authority, not model quality. The ladder runs Rule-bound, Simulated, Shadowed, Bounded authority, Governed autonomy — and the public automotive record stops at Simulated and Shadowed. No OEM has published a plant-floor policy holding unattended write authority.
  3. The safety envelope is always enforced outside the policy. Protective functions stay in the certified hardware and logic they already live in; the policy proposes a value inside a window the safety system would hold even if the policy output were random. A constraint expressed only in a reward term is not a constraint.
  4. The honest KPI is counterfactual, never a training curve. Divergence rate against the incumbent rule plus off-policy evaluation on the logged decisions — with an interval, not a point estimate — is what tells you whether the policy would have been better on your line.
  5. Reward design is where automotive RL fails, and it fails as a budget boundary: a single-term reward buys throughput by spending quality, maintenance life, energy or due-date performance that lands in another department's ledger. Write the constraint terms and the gaming modes before the first training run.

Abbreviations used on this page

RL
Reinforcement learning
MDP
Markov decision process (the state–action–reward formalism RL assumes)
OPE
Off-policy evaluation (scoring a policy on decisions it did not take)
MPC
Model predictive control
PLC
Programmable logic controller
MES
Manufacturing execution system
AMR
Autonomous mobile robot (lineside material supply)
OEE
Overall equipment effectiveness
FTT
First-time-through rate
EOL
End-of-line (test)
PL d
Performance level d — a machinery safety-function rating under ISO 13849
IATF
International Automotive Task Force (IATF 16949)

Free · 8 questions · ~3 minutes

Score the control authority of your decisions

Eight questions, one at a time, about three minutes. Answer them and we build your personalised control-authority report — your stage on the ladder, your score on each of the four dimensions, and the specific constraint standing between you and the next stage — and send it to your inbox. Your result doubles as the scoping note for a first instrumented decision.

0 of 8 answered

Question 1 of 8Decision & environment

Take the sequential decision you would most like a policy to make. What do you have logged about it today?

Off-policy evaluation is only possible where the state, the action actually taken and the realised outcome were recorded together against one identifier.

How the score maps to a stage
  • 0–5 — Stage 1, Rule-bound. Every sequential decision on the line is made by a fixed rule or a person, and the decision itself is not recorded — so no policy can be evaluated and no baseline exists to beat.
  • 6–11 — Stage 2, Simulated. A model of the decision exists — a replay of the log, a discrete-event model of the line or a robot simulation — and policies train in it, but nothing has ever run against the plant.
  • 12–16 — Stage 3, Shadowed. The policy runs on live plant state and proposes an action at every decision point while the incumbent still acts, and both the divergence and its counterfactual value are logged.
  • 17–21 — Stage 4, Bounded authority. The policy writes its decision into the MES or the controller inside an envelope enforced outside the model, with a named approver, a drilled one-switch revert to the incumbent rule, and a change record in the quality system.
  • 22–24 — Stage 5, Governed autonomy. Several policies hold bounded authority under a versioned control-policy register, retraining and redeployment ride the plant's change gateway, and people set the reward and the bounds rather than the individual actions.

What reinforcement learning is in an automotive plant — and what it is not

A definition, the loop as it actually exists on a line, and the thesis of this page: on a car line, reinforcement learning is a control-authority question long before it is a modelling question.

Reinforcement learning is a way of training a decision-maker — a policy — that repeatedly observes the state of a system, chooses an action, and learns from the consequences that follow, including consequences that arrive many steps later. It is not trained on labelled examples of the right answer, because for most sequential problems nobody knows the right answer; it is trained on outcomes. In an automotive plant that description fits a specific and fairly small family of decisions: which body to pull next out of the paint shop's selectivity bank, when to release a buffer between shops, which tugger to send to which lineside call, what to do with a booth setpoint for the next quarter of an hour. It fits none of the plant's perception problems and few of its prediction problems, and the single most common category error in industrial AI is using it where a supervised model or a well-posed optimisation would do the job with a hundredth of the risk.

Three neighbours are worth separating from it precisely, because plant conversations blur them constantly. Supervised learning maps an input to a label — this weld is porous, this claim is this failure mode — and its consequence is immediate and observable, which is why inspection and prediction are supervised problems. Model predictive control chooses actions over a horizon too, but it does so from an explicit written model of the dynamics and constraints, and where such a model exists it is almost always the better engineering choice: cheaper to validate, easier to explain to a safety engineer, and behaves predictably outside its training data because it has none. Reinforcement learning earns its place in the gap between the two — decisions whose consequence is delayed and whose dynamics nobody can write down, but which the plant repeats often enough, and logs well enough, to learn from experience.

Reinforcement learning is learning what to do — how to map situations to actions — so as to maximize a numerical reward signal.

That definition contains the whole problem of putting it in a plant. "Maximize a numerical reward signal" is a promise to pursue whatever is written down with more persistence, and less common sense, than any rule or any person. A shift leader told to reduce colour changes will not starve the assembly line of a due model to do it; a policy will, unless due-date deviation is a term in its reward with a sign and a weight. And "map situations to actions" is a promise to act, which is what separates this page from every other AI topic in this series. A vision model that is wrong produces a false reject and an operator's irritation. A policy that is wrong produces an action, in a plant, on a machine — which is why the rest of this page is largely about what stands between the policy's output and the equipment, and why the ladder that organises it measures authority rather than accuracy.

Decision value released along the control-authority ladder

The curve is not linear, and its flat section is longer than most programmes plan for. Value stays close to zero through Rule-bound and Simulated — where simulated gains are claims about a model — and only begins to release at Shadowed, the first stage that produces evidence about your own line. The steep segment is Bounded authority, where an action is actually written; Governed autonomy adds breadth rather than height, because the second and third policies reuse the envelope, the evidence pattern and the register built for the first.

Decision value released by stage

  • Stage 1 · Rule-bound — 34% of operators. Every sequential decision on the line is made by a fixed rule or a person, and the decision itself is not recorded — so no policy can be evaluated and no baseline exists to beat.
  • Stage 2 · Simulated — 31% of operators. A model of the decision exists — a replay of the log, a discrete-event model of the line or a robot simulation — and policies train in it, but nothing has ever run against the plant.
  • Stage 3 · Shadowed — 22% of operators. The policy runs on live plant state and proposes an action at every decision point while the incumbent still acts, and both the divergence and its counterfactual value are logged.
  • Stage 4 · Bounded authority — 10% of operators. The policy writes its decision into the MES or the controller inside an envelope enforced outside the model, with a named approver, a drilled one-switch revert to the incumbent rule, and a change record in the quality system.
  • Stage 5 · Governed autonomy — 3% of operators. Several policies hold bounded authority under a versioned control-policy register, retraining and redeployment ride the plant's change gateway, and people set the reward and the bounds rather than the individual actions.

Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the published record of bounded industrial RL deployments.

The reinforcement-learning loop as it actually exists on a line

Three lanes, and only the middle one touches the plant. The training lane is entirely offline: it learns from the decision log and is evaluated counterfactually. The live lane runs left to right at the real decision cadence, and every action passes an action filter before any write occurs. The governance lane owns the two artefacts that make the whole thing approvable — the reward and envelope specification, and the change record — while the safety-rated stop sits outside the loop altogether and is unchanged by any of it.

  • Data & feeds
  • AI / model
  • Where value leaks
  • System-of-record action
  • Human in the loop

The process, in words

  • The training lane never touches the plant. It reads the decision log — state, action taken, alternatives available, realised outcome — replays it through an environment that has been validated against real production, trains a policy, and scores that policy counterfactually. A policy is promoted out of this lane only on evidence, never on a training curve.
  • The live lane runs at the real decision cadence. Plant state feeds both the incumbent rule and the learned policy; the policy proposes one action; the action filter checks it against a written envelope and either passes it to the MES or controller write, or rejects it and lets the incumbent's action stand. Every write carries the policy version, and every realised outcome goes back into the log.
  • The governance lane contains nothing that is learned. The reward and envelope specification is written and signed before training begins, and it constrains both the training lane and the filter. Every write produces a change-controlled record — policy version, evidence, approver, revalidation date.
  • The safety-rated functions sit outside the loop by design. Protective stops, interlocks, guarding and robot speed-and-separation monitoring are unchanged by the presence of a policy, act on the equipment regardless of what any software proposed, and are validated independently. If a proposed architecture requires the policy to be trusted for safety, the architecture is wrong.
Step-by-step insights
The decision log — the artefact that makes everything else possible
The record has four parts and all four are load-bearing: the state at the moment of the decision, the action that was actually taken, the alternatives that were genuinely available, and the outcome that followed. Most plants have the fourth and none of the first three. The alternatives matter more than teams expect, because off-policy evaluation is a comparison against the options that existed, and a log without them can only tell you what happened, never what could have. A stable decision identifier joining all four is what turns a stream of events into a dataset — and it is the same discipline that the unit-level data backbone imposes elsewhere in the plant, which is why programmes with a serialised production record reach usable logs months faster.
The replay environment — validate the model before you trust the policy
The environment is the second-most-abused artefact in industrial RL, after the reward. The correct first use of it is not training but backtesting itself: run last quarter's real decisions through it and check that it reproduces the realised KPIs within a tolerance you publish. Nominal-flow models are the usual failure — no breakdowns, no rework re-entry, no manual overrides, no changeover — and a policy trained in one becomes fluent in states the plant never occupies. Adding disruption almost always reduces the simulated gain and increases its truthfulness, which is an uncomfortable conversation and the right one.
Off-policy evaluation — the only honest pre-deployment number
Off-policy evaluation estimates how a new policy would have performed on decisions taken by a different policy — your incumbent rule. It is the industrial equivalent of a holdout, and it comes with an interval that widens where the log is thin or where the new policy diverges hardest. That widening is the useful part: it tells you exactly which regions of the decision space you have no evidence about, which is the list of situations to watch during shadow. Treat a point estimate with no interval as a marketing artefact rather than an evaluation, whether it comes from a vendor or from your own team.
The incumbent rule — the baseline that never leaves
The incumbent stays live for the whole life of the policy, not just during shadow. It is the fallback the filter reverts to on a rejected action, the destination of the one-switch revert, and the comparison the regret number is computed against. Programmes that decommission the rule at promotion lose all three at once, and discover on the first bad week that they have nothing to fall back to and no way to say how bad the week actually was. Keeping a rule alive costs almost nothing; recreating one under pressure costs a quarter.
The action filter — a constraint in the reward is not a constraint
This is the single most important architectural claim on the page. Constraints expressed as reward penalties are preferences: a policy that finds enough upside will pay the penalty, and a policy whose inputs have gone stale will violate them without knowing. The filter is deterministic code, outside the model, in the same service that performs the write, and it checks range, rate of change, hysteresis and suppression conditions before anything reaches the plant. It is also the artefact that makes the deployment reviewable, because a quality or safety engineer can read it, test it and sign it without understanding the policy at all.
Safety-rated functions — the layer no policy may enter
Machinery-safety practice puts protective functions in certified hardware and logic, rated for the required performance level, validated independently of whatever application software exists. A learned policy changes none of that: it may choose a setpoint inside a validated window, but it may not be the thing that stops a machine, opens a guard or relaxes speed and separation limits. The practical test is simple and it belongs in every design review — if the policy stopped working, or started returning random values, would anyone be hurt? If the answer is anything other than an immediate no, the boundary has been drawn in the wrong place.

The control-authority ladder: five stages

For each stage: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving actually costs.

The ladder below measures how much authority a learned policy holds over a real decision, from Rule-bound — where the decision is not even recorded — to Governed autonomy, where several policies write inside stated bounds under a versioned register. It deliberately does not measure model quality, because model quality is not what makes a policy safe, useful or approvable in an automotive plant; a mediocre policy inside a well-designed envelope is a manageable process change, and an excellent policy with unbounded write access is an incident waiting for its Tuesday. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own MES and controls estate this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.

One thing to notice as you read. The stages get harder in the opposite order to the one most programmes assume. Stages 1 and 2 are engineering problems with known solutions and can be bought with effort. Stages 3, 4 and 5 are evidence and governance problems, and they are gated by people who will not be persuaded by a training curve — process owners, quality engineers, safety engineers, and eventually a customer auditor. That is why a plant with a mediocre model and an excellent envelope will ship, and a plant with a state-of-the-art policy and no written bounds will not.

Select a stage

Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.

Stage 1

Rule-bound

34% of operators sit here

Every sequential decision on the line is made by a fixed rule or a person, and the decision itself is not recorded — so no policy can be evaluated and no baseline exists to beat.

Rule-bound is the correct default and the honest starting point for almost every automotive plant, and it deserves respect before critique. The rules are usually good: decades of industrial engineering are compressed into a paint shop's block-building heuristic, a conveyor's release logic and a fleet manager's priority queue, and they have been tuned against real breakdowns by people who watched the line. A learned policy that beats them is doing something genuinely difficult, not something obvious. The problem at stage 1 is not that the rules are bad. It is that the plant keeps no record of them operating, so there is no way for anyone — including the people who wrote them — to ask whether a different choice would have been better.

The tell is the decision log, and it is a five-minute diagnostic. Ask for one week of one specific decision: which body the paint sequencer pulled next out of the selectivity bank, which tugger the fleet manager assigned to which lineside call, when the shift leader released the buffer between body shop and paint. Ask for it with the state at the moment of the decision and the alternatives that were physically available. At stage 1 the realised outcome is in the MES — colour changes per shift, starvation minutes, kilowatt-hours — and the decision is nowhere. The plant knows what happened and not what was chosen, which is precisely the wrong half of the pair for any technique that learns from consequences.

This stage is cheap to leave for one decision and expensive to leave for a plant, and the difference is worth being disciplined about. Instrumenting a single decision point is a controls-and-MES task measured in weeks: write the state, the action, the alternatives and a stable decision identifier to a log every time the rule fires. The common mistake is to treat that logging as the first phase of an AI project and therefore to fund it, staff it and justify it as one. It is process instrumentation, it belongs to industrial engineering as much as to data science, and it pays for itself in ordinary dispatch analysis — how often does the rule take the option the shift leader would have overridden? — long before any policy exists.

In practice

The sequencer nobody could question

A paint shop's block-building rule decides which body leaves the selectivity bank next, several thousand times a week. The shift's colour-change count and purge consumption are in the MES and reviewed every morning. Which bodies were sitting in the bank at 09:14, and which one the rule chose, is recorded nowhere. When a vendor proposes a learned sequencer, the only feasibility question that matters — how much better could anything have done on last quarter's real bank states? — cannot be answered at all. The decision reverts to whoever presents most persuasively, which is exactly how plants end up buying a policy they can neither evaluate nor revert.

What it looks like

  • Sequencing, release and dispatch decisions live in PLC logic, an MES rule set or a shift leader's judgement
  • The decision is not logged — only the shift's aggregate KPIs survive
  • Reinforcement learning appears in strategy decks and vendor demonstrations, never against plant state
  • Nobody can say what the incumbent rule chose last Tuesday, or what else was available to choose

Diagnostic signals you can check this week

  • Ask for one week of one decision with its state and available alternatives; if the answer is a KPI report, you are here
  • Ask for the incumbent rule in writing. Undocumented PLC logic is common, and it is the baseline any policy has to beat
  • Ask anyone to state the decision's frequency — decisions per shift is the unit RL economics are actually computed in
  • Ask whether any AI in the plant today chooses an action, or only classifies and predicts. Almost always the latter

Anti-pattern · Buying a policy before you own the log

Vendors will offer a sequencer, a dispatcher or an energy optimiser pre-trained on somebody else's plant, and at stage 1 there is no way to test the claim. Without your own decision log you cannot evaluate the policy before it is deployed, cannot detect the week it stops working, and cannot revert with evidence rather than with a shrug. The log is the asset that makes every later vendor conversation an evaluation instead of an argument — and it is yours whether or not you ever train anything. Build it first, and build it for the decision you care about rather than for the plant.

What holds you here

The decision is not recorded with the state it was taken in, so no policy can be evaluated and there is no measurable incumbent to beat.

Highest-leverage next move

Pick one repeated sequential decision and log state, action taken, alternatives available and realised outcome against a stable decision identifier.

Cost of leaving

Effort
6–12 weeks
Team
One controls engineer and one MES engineer, part-time, plus the process owner for the decision
Risk
Low — logging is additive and nothing in production changes
To next stage
3–6 months

If this is you, the next step is

A four-week scope: pick the decision, define the state and action record, and start the log.

Instrument one decision

Stage 2

Simulated

31% of operators sit here

A model of the decision exists — a replay of the log, a discrete-event model of the line or a robot simulation — and policies train in it, but nothing has ever run against the plant.

Stage 2 is where automotive reinforcement learning looks most convincing and is least informative. There is a training curve that climbs, a simulated improvement quoted as a double-digit percentage, and a demonstration that runs faster than real time in front of a steering group. Every part of that is real work and none of it is evidence about the plant: a reward improvement inside an environment is a statement about the environment. The gap between the two is the sim-to-real gap, and at stage 2 it is almost never quantified — not because teams are dishonest, but because quantifying it requires the decision log that stage 1 was supposed to produce.

The discipline that separates a useful stage 2 from a decorative one is environment validation, and it runs backwards from what people expect. Before you evaluate a policy in the environment, evaluate the environment itself: replay last quarter's real decisions through it and check that it reproduces the realised KPIs — colour changes per shift, starvation and blocking minutes, kilowatt-hours per unit, due-date deviation — within a stated tolerance, and publish the tolerance. An environment that cannot reproduce what the incumbent rule actually achieved has no standing to tell you what a new policy would achieve. This single check retires more bad RL proposals than any amount of model review.

There is a second, quieter trap: the environment usually omits exactly the things that make a real automotive line hard. Breakdowns and micro-stops, rework bodies re-entering the sequence out of order, the shift leader's manual overrides, the Thursday changeover, the supplier delivery that arrives on the wrong pallet, the colour that is temporarily unavailable because a tank is being cleaned. A model of the line on a good day trains a policy for a plant that does not exist, and the policy's advantage is often precisely a fluency in states the plant never occupies. Adding disruptions to the environment usually reduces the simulated gain and increases its usefulness, which is an uncomfortable trade to present and the right one to make.

In practice

The gain that did not survive contact

A sequencing policy showed a large simulated reduction in colour changes against the plant's block-building rule, sustained across a long training run. When the same policy was replayed against the previous quarter's actual bank contents, most of the advantage disappeared: the gain had been earned in states where four bodies of the same colour were simultaneously available in the bank, which the simulator generated freely and the upstream release rule makes rare. The residual advantage on real states was real, small, and worth pursuing — but it was a different number, attached to a different argument, and only the replay could tell them apart.

What it looks like

  • A replay or simulation environment exists, and a policy trains in it to a better reward than the incumbent rule
  • The environment's fidelity against realised production has never been measured as a number
  • Results are reported as reward curves and simulated percentage gains
  • No live plant system reads the policy's output

Diagnostic signals you can check this week

  • Ask what the sim-to-real gap is, as a number, on the KPI the business case rests on
  • Check whether the environment contains breakdowns, rework re-entry and manual overrides, or only nominal flow
  • Check that the incumbent rule is implemented inside the same environment — without it there is no comparison, only a score
  • Ask whether the simulated gain carries a confidence interval or is quoted as a single figure

Anti-pattern · Optimising the policy instead of the environment

When simulated gains fail to reproduce, the reflex is to train harder — a bigger network, a longer run, a different algorithm, more reward shaping. It almost never helps, because the fidelity of the environment sets the ceiling on everything trained inside it, and a policy that exploits a modelling artefact will exploit it more efficiently the better it gets. Spend the next month on replay validation and disruption modelling rather than on architecture search. The uncomfortable version of this rule: if the environment cannot reproduce the incumbent's realised KPIs, the correct next step is not a better policy but a better model of your own line.

What holds you here

The environment has never been validated against realised production, so a simulated gain is a claim about the model rather than about the line.

Highest-leverage next move

Replay a full quarter of logged decisions through the environment, publish the reproduction error per KPI, and only then compare a policy against the incumbent rule.

Cost of leaving

Effort
3–6 months
Team
One simulation or operations-research engineer, one ML engineer, the process owner
Risk
Medium — the real risk is spending a year inside a simulator nobody ever validates
To next stage
3–6 months

If this is you, the next step is

We replay your own decision log through your model and report the reproduction gap, KPI by KPI.

Validate your environment

Stage 3

Shadowed

22% of operators sit here

The policy runs on live plant state and proposes an action at every decision point while the incumbent still acts, and both the divergence and its counterfactual value are logged.

Shadow is the first stage that produces evidence about your plant rather than about a model of it, and it is astonishingly cheap for what it yields. The policy is served against live state, at the real cadence, seeing the states nobody simulated: the micro-stop at 03:40, the rework body that re-entered the sequence, the colour that went unavailable mid-shift. It proposes an action at every decision point. The incumbent rule still acts. Nothing in the plant changes, no approval is needed beyond read access, and within a fortnight the programme has something it has never had before — a paired record of what was chosen and what would have been chosen, on real states.

Two numbers come out of a shadow run, and they are this page's core instruments. Divergence rate — the share of decision points where the policy would have chosen differently — tells you whether there is anything to argue about at all. A divergence rate near zero is a genuine and useful result: the policy has rediscovered the rule, the rule is good, and the programme should stop, which is worth knowing in month three rather than month thirty. Counterfactual regret, estimated by off-policy evaluation over the logged decisions, tells you the sign and the size of the difference with an honest interval around it. Where the log is thin or the policy diverges often, that interval is wide, and a wide interval is information rather than a failure.

Shadow's failure mode is that it never ends. It is comfortable: no operational risk, a weekly number that trends the right way, a slide that survives any steering group. Meanwhile the sponsor moves on, the state feed drifts, and the policy silently starts scoring against a plant that has been rebalanced twice. The remedy is procedural and belongs in the shadow plan on day one — a date, an evidence threshold and a named person who decides to promote or to stop. The other thing to plan for is the shadow run's most valuable by-product: divergences that turn out to be defects in the state feed rather than disagreements about strategy. Reviewing the largest-regret cases with the process owner every week is where those surface, and each one is worth more than a month of training.

In practice

The divergence that was a data bug

In the first fortnight of shadow on a buffer-release decision, the policy diverged from the incumbent on roughly two decisions in five — far more than anyone expected from a well-tuned rule. The weekly review of the largest-regret cases found the cause in three sessions: the policy was reading bank contents from a feed that lagged reality by two conveyor positions, so it was consistently reasoning about a bank that had already moved on. The fortnight's real product was not a better policy but a corrected state feed, which the incumbent rule had been silently tolerating for years because a rule that only looks at the head of the queue never notices.

What it looks like

  • The policy is served against live MES and controller state at the real decision cadence
  • Every proposal is logged next to the incumbent's action and the realised outcome
  • Divergence rate and counterfactual regret are reported weekly — reward is not reported at all
  • A promotion decision exists, with pre-agreed evidence thresholds and a date on it

Diagnostic signals you can check this week

  • Ask for last week's divergence rate and the three highest-regret decision points, by name
  • Check that the policy reads exactly the state a live deployment would read, at the same latency — a shadow run on enriched data proves nothing
  • Ask whether the promotion criteria and the promotion date were written before shadow started
  • Ask who reviews divergent cases with the process owner, and how often — that review is where policy and data defects surface

Anti-pattern · Shadow as a permanent parking space

Because shadow carries no operational risk, it also carries no forcing function, and programmes settle into it for years. The tells are recognisable: the weekly report still exists but nobody reads the divergence review, the promotion criteria were never written down, and the answer to 'what would it take to switch this on?' is a discussion rather than a document. Fix it at the start, not at the end. Write the envelope specification, the evidence threshold and the promotion date into the shadow plan before the first proposal is logged, and treat a decision to stop as an equally successful outcome — because a policy that cannot beat a good rule is a finding, not a failure.

What holds you here

The policy has no write path and no agreed promotion criteria, so evidence accumulates indefinitely without ever converting into authority.

Highest-leverage next move

Write the action envelope and the promotion criteria, have the process owner and quality sign them, and put a date on the promotion decision.

Cost of leaving

Effort
3–6 months
Team
ML engineer, MES and controls integration engineer, process owner, plus quality for the promotion criteria
Risk
Low operationally, high organisationally — the real risk at this stage is stalling in it
To next stage
6–12 months

If this is you, the next step is

Evidence thresholds, promotion date and the envelope specification, agreed before the first proposal is logged.

Design a shadow run that ends

Stage 4

Bounded authority

10% of operators sit here

The policy writes its decision into the MES or the controller inside an envelope enforced outside the model, with a named approver, a drilled one-switch revert to the incumbent rule, and a change record in the quality system.

Bounded authority is the first stage at which the policy changes what the plant actually does, and it is defined entirely by what constrains it rather than by what it computes. The deliverable of this stage is the envelope, not the model. An envelope is a written artefact: the range each action may take, the rate at which it may change, the hysteresis that prevents oscillation, the conditions under which the write is suppressed entirely, and the signature of the process owner who agreed all of it. The filter that enforces the envelope runs outside the policy, in the same service that performs the write, so that a policy which returns nonsense — because a feature went stale, because a model version was mis-deployed, because the input distribution moved — cannot express the nonsense as an action.

Automotive already knows how to do this, which is the argument that gets the stage approved. A control plan states the window a parameter may occupy and the reaction plan when it leaves; a machine capability study establishes that the process can hold that window; a change to either goes through change control with evidence attached. A bounded policy is a new way of choosing a value inside a window that already existed and was already validated. That framing matters more than any technical detail when the proposal reaches a quality manager: the sentence is not 'we are deploying AI on the line', it is 'we are changing how a setpoint inside the existing control plan is selected, here is the evidence, here is the filter, here is the revert, and the safety validation of the cell is unchanged'. Programmes that lead with the model instead of the envelope get a longer queue and a worse answer.

The engineering that actually consumes the quarter is unglamorous and it is not modelling. It is the filter and its test suite, the fallback path and the switch that selects it, the write logging that records the policy version against every action, the rate limits and hysteresis, the suppression logic for degraded state feeds, and the monitoring that catches the day the policy starts hitting the envelope. That last number — envelope rejection rate — is the most useful single instrument at this stage. A rising rejection rate means either the policy is drifting or the plant has moved outside the conditions the envelope was written for, and it says so weeks before any KPI does. Watch it the way you would watch a control chart, because that is exactly what it is.

In practice

The setpoint that could only move so far

A paint booth's supply-air temperature and humidity setpoints are chosen every fifteen minutes by a policy trained on a year of logged booth conditions, ambient data and finish-defect outcomes. The envelope states that each setpoint may move by no more than a fixed increment per interval and may only sit inside the booth's validated process window; a filter in the write service rejects anything else and falls back to the schedule. Booth condition monitoring trips an automatic revert to the incumbent schedule if conditions leave the window for more than a set period. Nothing about the booth's fire and ventilation safety functions changed, because none of them was ever in the policy's path.

What it looks like

  • An action filter outside the policy rejects any action outside stated limits before it is written
  • Safety-rated functions are untouched — the protective stop is exactly what it was before the policy existed
  • The incumbent rule still runs and is one switch away, and the revert has been exercised deliberately
  • Policy version, reward specification, envelope and evidence are recorded as a process change with a named approver

Diagnostic signals you can check this week

  • Ask to see the envelope as a written artefact — ranges, rate limits, suppression conditions, and who signed them
  • Ask when the revert to the incumbent was last exercised on purpose, not in an incident
  • Ask for the envelope rejection rate trend over the last ten weeks
  • Ask whether the safety validation of the cell or process changed because of the policy. The correct answer is no

Anti-pattern · Widening the envelope so the policy can win

When the filter rejects a lot of proposals, the tempting fix is to loosen the limits, and it is tempting precisely because it works: the measured gain improves immediately. What has happened is that the envelope has quietly become a function of the policy's ambition rather than of the process's physics and its validated capability. Once that inversion happens the envelope stops being a constraint and becomes a dial, and the first genuinely bad action is not far away. Treat envelope changes exactly like control-plan changes: a written reason, capability evidence, a named approver, a version, and a record. If the physics genuinely permits a wider window, prove it the way you would have proved it without a policy involved.

What holds you here

Every additional policy is a bespoke negotiation with its own envelope, evidence pack and approver, so the second and third decisions cost as much as the first.

Highest-leverage next move

Standardise the envelope, the evidence pack and the revert into a reusable pattern, and open a control-policy register holding every deployed policy's version, bounds and revalidation date.

Cost of leaving

Effort
6–12 months
Team
ML engineer, controls and MES engineer, process owner, quality engineer, plus a safety engineer for the validation review
Risk
Higher — the policy now writes into a running process, so the filter, the revert and the evidence pack are the critical path
To next stage
12–24 months

If this is you, the next step is

Limits, rate limits, rejection monitoring, revert drill and the change record — the artefacts that make write access approvable.

Design and validate the envelope

Stage 5

Governed autonomy

3% of operators sit here

Several policies hold bounded authority under a versioned control-policy register, retraining and redeployment ride the plant's change gateway, and people set the reward and the bounds rather than the individual actions.

Governed autonomy is much narrower than the phrase suggests, and the narrowness is the point. It is not a self-optimising factory. It is an enumerated set of decisions — a sequencing decision, a release decision, a dispatch decision, a utilities setpoint — that a learned policy takes inside stated bounds without a per-decision approval, while everything outside those bounds escalates. Decisions with a safety function in their path, a homologation consequence, or unbounded commercial exposure are correctly held at bounded authority forever, and a plant that has thought clearly about this can say which decisions those are without hesitating. The stage is reached by subtraction: by naming what will never be automated, the remainder becomes governable.

The register is the artefact that defines the stage, and it is the same instinct as a control plan applied to a different object. For each deployed policy it holds the version in production, the reward specification it was trained against, the envelope it writes inside, the evidence that justified promotion, the approver, and — the field that does the most work — a revalidation date. A policy without a revalidation date is a policy nobody will notice has expired, and expiry in a plant is not a theoretical event: it happens on the Monday after a launch, a rebalance or a new colour. Everything else in the register exists so that an auditor, a customer representative or a new engineer can reconstruct why a machine did what it did on a given day, from documents rather than from the memory of whoever built it.

Sustaining this stage is a maintenance discipline rather than a summit, and it is the stage most likely to regress silently. Every product launch, line rebalance, new colour, new supplier, new shift pattern and new upstream release rule changes the distribution the policies were evaluated on, and nothing in the plant's normal rhythm will tell you that a policy's evidence has gone stale. Operators who sustain it wire policy revalidation into the launch gateway next to process capability and part-approval evidence, so a policy that has not been re-evidenced against the new product simply does not get write access on day one of production. The sober conclusion, and the one this page keeps returning to, is that being ready for more autonomy later is almost entirely a matter of doing the present work properly now: log the decision, validate the environment, bound the action, evidence the change.

In practice

The launch that expired four policies

A plant running bounded policies on paint sequencing, buffer release, tugger dispatch and booth conditioning adds a derivative model with a new body variant and two new colours. The launch gateway flags all four policies for revalidation because their evaluation windows predate the variant. Three are re-evidenced by replaying the ramp-up weeks and are back inside their envelopes within a month; the sequencing policy is not, because the new colours change the block economics enough that its regret estimate crosses zero. It reverts to the incumbent rule for the launch and returns to shadow — which is the system working exactly as designed, and would have been an unexplained KPI drift in a plant without a register.

What it looks like

  • A control-policy register lists every deployed policy with version, reward specification, envelope, evidence and revalidation date
  • Retraining and redeployment pass the same change gate as any other process change
  • Regret, divergence, envelope rejection and constraint violations are monitored continuously and page a named human
  • Reward specifications are reviewed at every product launch, line rebalance and process changeover

Diagnostic signals you can check this week

  • Ask to see the control-policy register and check whether every entry has a revalidation date that has not passed
  • Ask when the paging path for a constraint violation was last tested, and who answered
  • Check whether the last product launch triggered policy revalidation as a gateway item, or after somebody noticed
  • Ask whether an auditor could reconstruct a single automated action — state, policy version, envelope, approver — from records alone

Anti-pattern · Treating retraining as maintenance rather than as change

Once several policies are running well, retraining starts to feel like housekeeping: refresh the data, rerun the pipeline, redeploy. But a retrained policy is a different decision-maker with different failure modes, and deploying it without the evidence and the approval that the first version needed is a process change made without change control — the thing IATF-certified plants spend the rest of their time not doing. The discipline is unpopular and simple: every version passes the same gate, carries its own off-policy evidence, and enters the register with its own revalidation date. If that feels too heavy for the cadence you want, the honest conclusion is that the decision should not have been automated at that cadence.

What holds you here

Sustaining bounded authority across launches, changeovers and staff turnover is an organisational discipline rather than a build, and its decay is invisible until a KPI moves.

Highest-leverage next move

Wire policy revalidation into the launch gateway, and treat the reward specification and the envelope as versioned artefacts reviewed exactly like a control plan.

Cost of leaving

Effort
Continuous
Team
A small platform team plus a standing forum: process engineering, quality, controls, maintenance, safety
Risk
Concentrated and regulatory — low frequency, high consequence, with documentation duties attached

If this is you, the next step is

We reconstruct one automated decision from your records — state, version, envelope, approver — and hand you the gap list.

Audit a live policy end to end

Where automotive plants actually sit on the ladder

The distribution across the five stages, why the public record stops at Shadowed, and what the one well-documented bounded deployment — outside automotive — actually shows.

Most automotive plants are at Rule-bound or Simulated, and the population thins sharply above Shadowed. That distribution has a structural cause rather than a cultural one: the decisions RL suits are already served by rules that took decades to tune, the plants that could evaluate a replacement mostly do not log the decision, and the step that converts evidence into authority runs through a change-control regime that was designed — correctly — to make process changes slow. Nothing about that is a failure of ambition. It is what a certified, safety-critical, capital-intensive production system looks like when a new class of decision-maker asks for write access.

Distribution of automotive plants across the five control-authority stages

Illustrative distribution, not a survey: it is a reading of the published record — how little plant-floor reinforcement learning with write authority is documented anywhere, against how much simulation, virtual commissioning and robot-learning research OEMs publish. Treat the shape as the argument and the numbers as illustrative.

Share of plants

  • 34% — 1 · Rule-bound (the decision is not even logged)
  • 31% — 2 · Simulated
  • 22% — 3 · Shadowed
  • 10% — 4 · Bounded authority
  • 3% — 5 · Governed autonomy

Source: Illustrative; anchored to the published record of bounded industrial RL deployments

It is worth being blunt about the public record, because vendor material is not. No automotive manufacturer has published a plant-floor reinforcement-learning policy holding unattended write authority over a production decision. What OEMs do publish is the two stages below that, and they publish a lot of it: BMW Group (opens in a new tab) reports extensive virtual planning and simulation of production systems before physical build, Toyota (opens in a new tab) and Toyota Research Institute (opens in a new tab) publish robot-learning research conducted in laboratories rather than on lines, and General Motors (opens in a new tab) reports machine learning applied across manufacturing operations in overwhelmingly predictive and perceptual roles. That is a Simulated-and-Shadowed frontier, honestly reported, and reading it as evidence of learned control on production lines is the single most common misreading in this topic.

The scale term is what makes the small percentages worth chasing at all. ACEA (opens in a new tab) represents an industry producing vehicles in the tens of millions of units a year across its members' plants, and a paint shop makes a sequencing decision every time a body leaves the bank — a decision taken thousands of times a shift, every shift, for the life of the line. At that repetition, a small average improvement is a large annual number and a small average degradation is an equally large one, which is the real argument for evidence discipline. It is also why the sensible first RL decision in a plant is one that repeats constantly, is cheap to reverse, and has no safety function anywhere near it.

Which plant decisions are actually reinforcement-learning shaped

The candidacy screen: five conditions a decision must satisfy, the decisions in an automotive plant that pass and fail them, and the quadrant that tells you which controller belongs on the decision at all.

A plant decision is reinforcement-learning shaped when five conditions hold together, and failing any one of them makes RL the wrong tool rather than a harder version of the right one. The decision must be sequential, in that the consequence of choosing now depends on and affects what can be chosen later. It must repeat often enough to accumulate evidence — hundreds of times a shift, not weekly. Its outcome must already be measured by a system the plant trusts. Its action space must be bounded and enumerable, so that an envelope can be written around it. And the alternatives available at each decision point must be knowable, because without them there is nothing to compare against. Most plant AI proposals fail at least two of these, and the screen below is how we sort a wish list in an afternoon.

  • Sequential, with a delayed consequence

    Pulling a silver body out of the bank now is cheap; it becomes expensive twenty bodies later when the block ends early and a colour change lands in the middle of a run. If the consequence of an action is fully visible within the same cycle, the problem is a prediction or an optimisation, and both are cheaper to build, validate and explain than a policy.

  • Repeated at high frequency

    Evidence accrues per decision, not per week. A decision taken three thousand times a shift generates a usable off-policy estimate inside a month; one taken twice a shift will not generate one inside a year, and a policy that cannot be evaluated should not be deployed. Decision frequency is the first number to establish and the one most business cases omit.

  • Outcome already measured by a trusted system

    Colour changes and purge volume are in the paint MES. Starvation and blocking minutes are in the conveyor system. Kilowatt-hours are on the meter. If the outcome would have to be newly instrumented, that instrumentation is the project, and it is worth doing on its own merits before anyone trains anything.

  • Bounded, enumerable action space

    Choosing among the bodies physically present in a bank is bounded by construction. Choosing an arbitrary continuous setpoint is bounded only if someone writes the range, the rate limit and the suppression conditions down. An action space that cannot be written down cannot be filtered, and an action that cannot be filtered will not — and should not — be approved for write access.

  • Known alternatives at each decision point

    Off-policy evaluation compares against what was available, not against what is imaginable. A log that records the chosen body but not the other eleven in the bank supports monitoring and supports nothing else. This is the condition plants most often discover they fail, and it is fixed with a schema change rather than with a platform.

Plant decisionIs it RL-shaped?The incumbent you must beatSystem of recordKPI it movesVerdict
Paint-shop colour block sequencingYes — sequential, thousands of decisions a shift, consequences compound down the blockRule-based block builder in the paint MESPaint MES / selectivity bank controllerColour changes per shift, purge volume, block length, due-date deviationBest first RL decision in most plants — high frequency, cheap to reverse, no safety function in the path
Buffer and bank release between shopsYes — releasing now starves or floods a downstream shop twenty minutes laterFixed FIFO plus shift-leader overridesMES / conveyor controlStarvation and blocking minutes, OEE availabilityStrong stage 3–4 candidate once the manual overrides are logged as decisions
AMR and tugger dispatch for lineside supplyYes — sequential assignment under uncertainty, outcome measurable per callFleet manager rule set: nearest vehicle, priority queueFleet management system / MESLineside stockouts, empty travel, fleet utilisationStage 3–4; the fleet manager's own safety layer already supplies part of the envelope
Paint booth and compressed-air setpointsYes — continuous control with delayed thermal and quality consequenceSchedules and PID loops; MPC where it has been builtBuilding management system / utilities SCADAKilowatt-hours per unit, booth condition excursions, finish defectsStage 4 candidate, and only inside the validated process window
Robot motion and path refinementPartly — RL-shaped in simulation, but the production motion programme is a validated artefactOffline-programmed, validated robot pathsRobot controller / offline programming systemStation cycle time, spatter and reworkSimulate and then deploy as a re-validated programme — never as a live policy on the cell
Torque and process-parameter tuning at a stationNo — the consequence is immediate and the mapping is learnable from labelsControl-plan windows set by process engineeringMES / torque controllerNOK rate, reworkUse supervised models and designed experiments; exploration on a joint is not acceptable
Defect detection at inspectionNo — one-shot classification, no sequence, no action consequenceVision system with tuned thresholdsVision system / QMSFalse-reject rate, first-pass yieldSupervised learning. Reinforcement learning adds nothing but risk and vocabulary
The candidacy screen applied to an automotive plant. Note that three of the seven rows are decisions where reinforcement learning is the wrong tool — a sorted list should always contain some of those, and a list where everything passes has not been sorted.

The two rejected rows deserve as much attention as the accepted ones, because they are where credibility is won or lost with a plant audience. Torque tuning fails the screen on its first condition — the joint either passes or fails now, so there is nothing sequential to learn — and it fails on a second, harder ground: exploration on a safety-relevant joint is not an acceptable activity, whatever the expected value. Inspection fails even more clearly. Both are properly served by supervised models and by the analytical disciplines covered elsewhere in this series, and a programme that proposes reinforcement learning for either is signalling that it has chosen a technique before looking at the problem. The right conversation with process engineering is always the screen first and the technique second.

Which controller belongs on this decision

Plot how well you can write down the plant's dynamics against how long the consequence of an action takes to arrive. Reinforcement learning owns exactly one quadrant, and three of the four answers are cheaper, faster and easier to validate than a policy.

Reinforcement learning territory

  • Delayed consequence, no written dynamics
  • Paint sequencing, buffer release, fleet dispatch
  • Learn from the decision log; evaluate off-policy; shadow before authority

Model predictive control

  • Delayed consequence, dynamics you can state
  • Booth conditioning, oven profiles, energy scheduling
  • Cheaper to validate and far easier to explain to a safety engineer

Supervised learning

  • Immediate consequence, no written dynamics
  • Inspection, predictive quality, parameter recommendation
  • Predict the outcome and set the parameter; run a designed experiment, not exploration

Classical control and rules

  • Immediate consequence, dynamics you can state
  • Interlocks, PID loops, control-plan windows
  • Reinforcement learning here is expensive theatre with a validation bill
How long the consequence takes to arrive — top: Later, over many decisions, bottom: Immediately, within this cycle
A model of the dynamics you can write down — left: None — you only have logs, right: Good — dynamics and constraints are known

Notice what the top-right quadrant implies for sequencing conversations. If your process engineers can write down the dynamics and the constraints — and for booth conditioning, oven profiles and much of the utilities estate they often can — then model predictive control is very likely the better engineering answer, and saying so early buys enormous credibility for the cases where a policy genuinely is the right tool. Reinforcement learning's honest domain is the top-left: decisions where the plant's behaviour is a tangle of upstream release rules, breakdowns, product mix and human overrides that nobody has ever succeeded in writing down, but which the plant repeats often enough to learn from. That domain is real, it is valuable, and it is much smaller than the market for it.

Reward design: where automotive reinforcement learning actually fails

Not in the algorithm. In the objective — and specifically in the terms nobody wrote down, which is almost always a budget boundary rather than a mathematical oversight.

Automotive reinforcement-learning programmes fail at the reward far more often than at the algorithm, and they fail in a recognisable pattern: the reward contains the KPI of the department that sponsored the project, and the costs the policy discovers it can spend belong to departments that were not in the room. A sequencing policy rewarded on colour changes will happily hold a rare colour in the bank for two days, because due-date deviation is logistics' number. A release policy rewarded on OEE availability will run conveyors and drives at the top of their duty range, because component life is maintenance's number. Neither behaviour is a bug in the model. Both are the model doing exactly what was written down, with a persistence no human dispatcher would sustain.

  • Throughput bought with quality

    A reward on units per hour, with rework recorded in the quality system rather than the MES, is an invitation the policy will accept. It learns to push the line into conditions where the immediate count rises and the defect appears at end-of-line or, worse, in the field. The fix is not a smarter model; it is a first-time-through term with a real weight, read from the system that actually records rework.

  • Throughput bought with maintenance life

    Duty cycle, start–stop counts and thermal excursions are all things a policy can spend for immediate gain, and all things whose cost arrives months later in another budget. Include a duty term, or cap the behaviour in the envelope, or accept that the programme's second year will be spent arguing about a maintenance overspend that nobody can attribute.

  • The constraint satisfied by starving the constraint

    Told to minimise colour changes, a sequencer can achieve superb block lengths by simply never releasing bodies whose colour is scarce. Block length looks excellent; due-date deviation and bank residency quietly deteriorate. Any objective phrased as 'minimise X' needs a companion term on the resource X is being minimised out of, and a hard bank-residency limit in the envelope rather than a penalty in the reward.

  • A reward read from a signal a person can move

    If any term in the reward is influenced by a manual confirmation, a bypass switch or an operator-entered code, the policy will eventually learn to influence the operator rather than the process — by proposing actions that make the confirmation easier to give. Reward terms should read from process telemetry and system-of-record outcomes, never from a human's discretionary input.

  • Episode boundaries that hide the cost

    Ending an episode at the end of a shift teaches a policy that the state it hands over does not matter, so it learns to end shifts with a bank full of awkward bodies and an unbalanced sequence. Episodes must either span the handover or carry an explicit terminal-state value that prices the mess being left behind. This is the most common technical error on the list and the easiest to fix once seen.

TermSignSource systemRead cadenceWhat it prevents
Colour change countNegativePaint MESPer bodyNothing — this is the objective the project exists for
Purge volumeNegativePaint MES / solvent meteringPer changeoverCounting changeovers as equal when a light-to-dark change costs several times a dark-to-dark one
Due-date deviation, in hoursNegativeOrder system / MES build planPer body releasedPerfect blocks built by indefinitely deferring scarce colours
First-time-through at end-of-lineNegative on shortfallQuality systemPer shift, attributed per bodyThroughput bought with rework that lands in another department's ledger
Bank residency over the stated limitHard constraint, not a termSelectivity bank controllerContinuousA body starved indefinitely because the reward found it convenient
Oven and flash-off dwell windowHard constraint, not a termPaint process controlContinuousAny trade between finish quality and sequence efficiency — this one is physics
Terminal state value at shift handoverNegative on imbalanceBank state snapshotPer shift boundaryA policy that optimises its own shift by wrecking the next one
A worked reward specification for paint-shop colour sequencing — the shape every reward spec should take before training starts. Hard constraints are deliberately not reward terms: they belong in the envelope, where they cannot be traded away.

How to write a reward you can defend

  1. List the ledgers before the terms

    Ask which departments hold a KPI the decision could move: process engineering, quality, maintenance, logistics, energy, and the downstream shop. Each of them gets either a term with a sign and a weight, or an explicit written statement that the decision cannot affect them. There is no third option, and the exercise routinely finds a ledger nobody had considered.

  2. Separate constraints from preferences, and put constraints in the envelope

    Anything that must never happen — an oven dwell violation, a bank residency breach, a setpoint outside the validated window — is a hard constraint enforced by the action filter, not a penalty in the reward. Penalties are prices, and a policy that finds enough upside will pay them. This single distinction prevents most of the failures on the list above.

  3. Name the gaming modes in writing, before training

    For each term, write one sentence describing how a determined optimiser could satisfy it while making the plant worse, and what would detect that. This becomes the monitoring specification for shadow and for production, and it is the section of the document that process engineers engage with most, because they have watched people do exactly these things for years.

  4. Version it, sign it, and review it at every launch

    The reward specification is a controlled document with a version, an owner and an approval, in the same sense as a control plan. A new product, a new colour, a rebalanced line or a changed shift pattern can invalidate a weight that was correct last quarter — so it is reviewed at the launch gateway alongside process capability, not when someone notices a KPI drifting.

What the public record actually shows

Three publicly reported programmes, read against the ladder — and read honestly, because none of them is a plant-floor policy with write authority, and the operators do not claim it is.

The published automotive record on learned control stops at the ladder's second and third stages, and the operators themselves are precise about that even where the coverage is not. What the three programmes below demonstrate is genuinely valuable and genuinely bounded: an environment built at industrial fidelity, robot policies learned in a laboratory, and machine learning deployed at plant scale in perceptual and predictive roles. Read together they describe the frontier accurately — the assets that make bounded authority possible are being built, and the authority itself has not been publicly granted anywhere.

Three programmes read against the control-authority ladder

Outcomes as reported by the operators' own published material — verify against the linked source before reusing anything; we have not independently audited them. Images are generated library scenes, not operator photography, and no operator endorsement is implied.

Illustrative scene: a simulated automotive production hall used to plan and test line behaviour before physical buildBMW GroupGlobal OEM · premium vehicles · 30+ production sites12
Challenge
Planning and continuously changing a global production network in which layout changes, rebalances and launches traditionally had to be validated physically — late, expensive, and one plant at a time.
Approach
BMW has publicly reported building detailed virtual representations of plants and production systems and validating processes in simulation before physical commissioning, together with a portfolio of AI applications running on digitalised production data. In control-authority terms, that is investment in the environment: a model of the line at a fidelity where questions can be asked of it.
Reported outcome
BMW reports planning and validating production virtually ahead of physical build and extending simulation-first planning across its network, alongside in-plant AI applications built on the same production data.
What it shows about the curveAn industrial-fidelity environment is the precondition for stage 2, and it is a reusable asset that outlives any policy trained in it. It is not, on its own, evidence about a policy — which is why the ladder puts environment and evidence in different stages, and why replay validation sits between them.

BMW Group PressClub (opens in a new tab)

Illustrative scene: a robot cell being used for behaviour-learning research rather than production outputToyota / Toyota Research InstituteGlobal OEM · research institute · robot behaviour learning12
Challenge
Teaching robots dexterous, contact-rich behaviours that cannot be hand-programmed — the class of task where writing the dynamics down is precisely what nobody can do.
Approach
Toyota Research Institute publishes robot-learning research on teaching manipulation behaviours from demonstration and experience, conducted in laboratory settings. Separately, Toyota's own account of its production system rests on jidoka — equipment that stops itself when an abnormality occurs, so that a human decides what happens next.
Reported outcome
TRI reports learned robot behaviours in research settings; Toyota's published production philosophy continues to place a stopping authority in the equipment and the operator rather than in any optimiser.
What it shows about the curveThis is the separation the whole page turns on. A leading robot-learning programme and a production system whose first principle is that a machine must stop itself are not in tension — jidoka is the same idea as the action envelope, expressed sixty years earlier. Research capability is not plant authority, and the operator with the strongest learning programme is also the one most explicit about where the stop lives.

Toyota Research Institute (opens in a new tab)

Illustrative scene: plant-floor analytics screens monitoring robot and equipment condition across an assembly lineGeneral MotorsGlobal OEM · high-volume assembly · North American and global plants12
Challenge
Unplanned downtime and quality escapes across a large, heavily automated plant estate, where the cost of a stoppage is measured in vehicles per minute rather than in engineering hours.
Approach
GM publishes ongoing reporting on applying machine learning across its manufacturing operations, concentrated — as almost all published plant-floor AI is — on perception and prediction: condition monitoring of robots and equipment, inspection, and analytics that tell a maintenance or quality team what to do rather than doing it.
Reported outcome
GM reports machine learning used across manufacturing operations in monitoring, inspection and predictive roles, with the resulting actions taken by maintenance and production teams.
What it shows about the curveThe incumbent decision-makers on a car line are good, and they are supported by exactly this kind of analytics. Any reinforcement-learning proposal is competing with a tuned rule advised by a working predictive stack, not with a vacuum — which is why the divergence rate in a shadow run is so often near zero, and why that is a legitimate result rather than a disappointment.

GM Newsroom (opens in a new tab)

The safety envelope: what stands between a policy and the equipment

Five layers, annotated by the stage that first requires them — and one rule that never bends: the protective function is never learned, never inside the model, and never validated by the same evidence as the policy.

What stands between a learned policy and the equipment is a stack of five layers, only one of which is the model. From the bottom: safety-rated functions that were there before any policy existed and are unchanged by it; a process envelope that states the range, rate and suppression conditions any action must satisfy; the live decision path that computes a proposal and either writes it or falls back; a learning and evaluation layer that is entirely offline; and a governance layer that holds the versions, the evidence and the approvals. The single rule that organises all of it is that constraints are enforced by the layer beneath the one that could violate them — which is why a constraint written as a reward penalty is not a constraint at all, and why a policy that must be trusted for safety is an architecture error rather than a testing problem.

The five layers, annotated by the stage that first requires them

Each layer is annotated with the ladder stage that first makes it necessary. A programme trying to reach Bounded authority without the process-envelope and governance layers is not at stage 4 — it is at stage 3 with write access, which is a different and much worse thing.

  1. Safety-rated functions (present from day zero, unchanged)

    Stage 1+

    • Safety PLC and safety relaysProtective functions rated to the required performance level; no AI anywhere in this path
    • Robot safety functionsSafe zones, speed and separation monitoring — configured, validated, and untouched by any policy
    • Guarding, interlocks and E-stopsIndependent of every software decision, and the reason the policy question is never a safety question
  2. Process envelope

    Stage 4+

    • Control-plan windowThe validated range a parameter may occupy — the policy chooses inside it, never around it
    • Action filterDeterministic code in the write path: range, rate limit, hysteresis, suppression
    • Degraded-state suppressionStale feed, missing tag or failed health check suppresses the write and lets the incumbent stand
  3. Live decision path

    Stage 3+

    • State assemblyThe same state a live deployment reads, at the same latency, health-checked per decision
    • Policy serviceOne proposal per decision point, versioned, with a latency budget matched to the decision cadence
    • Incumbent fallbackThe rule still running, one switch away, and exercised on a schedule rather than in an incident
  4. Learning & evaluation (offline)

    Stage 2+

    • Decision logState, action, alternatives, outcome, joined on a stable decision identifier
    • Replay environmentValidated against realised KPIs, including breakdowns, rework re-entry and overrides
    • Off-policy evaluation harnessCounterfactual regret with intervals, re-run before every promotion
  5. Governance & change control

    Stage 4+

    • Control-policy registerVersion, reward spec, envelope, evidence, approver, revalidation date
    • Change record in the quality systemA policy change is a process change and carries the same evidence
    • Monitoring and pagingRegret, divergence, envelope rejection and constraint violations, with a named human on the other end

Pipeline described

  1. Safety-rated functions (present from day zero, unchanged) (stage 1+) — Safety PLC and safety relays: Protective functions rated to the required performance level; no AI anywhere in this path; Robot safety functions: Safe zones, speed and separation monitoring — configured, validated, and untouched by any policy; Guarding, interlocks and E-stops: Independent of every software decision, and the reason the policy question is never a safety question
  2. Process envelope (stage 4+) — Control-plan window: The validated range a parameter may occupy — the policy chooses inside it, never around it; Action filter: Deterministic code in the write path: range, rate limit, hysteresis, suppression; Degraded-state suppression: Stale feed, missing tag or failed health check suppresses the write and lets the incumbent stand
  3. Live decision path (stage 3+) — State assembly: The same state a live deployment reads, at the same latency, health-checked per decision; Policy service: One proposal per decision point, versioned, with a latency budget matched to the decision cadence; Incumbent fallback: The rule still running, one switch away, and exercised on a schedule rather than in an incident
  4. Learning & evaluation (offline) (stage 2+) — Decision log: State, action, alternatives, outcome, joined on a stable decision identifier; Replay environment: Validated against realised KPIs, including breakdowns, rework re-entry and overrides; Off-policy evaluation harness: Counterfactual regret with intervals, re-run before every promotion
  5. Governance & change control (stage 4+) — Control-policy register: Version, reward spec, envelope, evidence, approver, revalidation date; Change record in the quality system: A policy change is a process change and carries the same evidence; Monitoring and paging: Regret, divergence, envelope rejection and constraint violations, with a named human on the other end
Step-by-step insights
Safety-rated functions — the layer that is never in scope
Machinery-safety practice puts protective functions in certified hardware and logic, rated to a required performance level and validated independently of application software. A learned policy inherits none of that standing and must never be asked to. The practical design test is one sentence long: if the policy service stopped responding, or began returning random values within its output range, could anyone be hurt or could the equipment damage itself? If the answer is not an immediate no, the boundary has been drawn in the wrong place and no amount of testing repairs it. This is also why RL proposals aimed at robot motion in production cells belong in simulation and then in a re-validated motion programme — the validated artefact is the programme, not the learner.
The process envelope — the deliverable of stage 4
The envelope is a written document, not a configuration file: range per action, maximum change per interval, hysteresis to prevent oscillation, and the conditions under which the write is suppressed entirely. It is enforced by deterministic code inside the same service that performs the write, so a stale feature or a mis-deployed model version cannot express itself as an action. It is testable without understanding the policy, which is exactly what makes it signable by a quality or safety engineer. And its rejection rate is a leading indicator: when the filter starts rejecting more than it used to, either the policy has drifted or the plant has moved outside the conditions the envelope was written for, and both deserve a look before a KPI notices.
The live decision path — state parity is the whole game
The most damaging silent failure in a deployed policy is state divergence: the features the policy reads in production stop matching what it was evaluated on, because a tag was renamed, a feed lagged, a unit changed or an upstream rule started filtering. Health-check the state per decision and suppress rather than guess. The corollary is that a shadow run on enriched or reconstructed data proves nothing about production — shadow must read the identical state assembly a live deployment would read, at the same latency, or the evidence it generates is about a different system.
Learning and evaluation — offline, always
Nothing in this layer touches the plant, and the discipline is to keep it that way. Online exploration in a production automotive plant is essentially never justifiable: the cost of a bad action is measured in vehicles and the information gained is worth far less than the log you already have. Everything is learned from the decision log and from a replay environment validated against it, and every promotion is preceded by an off-policy estimate re-run on current data. The one legitimate exception is a bounded contextual-bandit arrangement on a genuinely reversible decision with a tiny action set, and even then the exploration budget belongs in the envelope, written down, with a rate limit.
Governance — the register is what makes the fifth stage exist
The register turns a set of deployments into an estate somebody can manage. Its most valuable field is the revalidation date, because expiry in a plant is an event with a cause — a launch, a rebalance, a new colour, a new supplier — and nothing in normal operations will announce that a policy's evidence has gone stale. Its second most valuable property is that it makes reconstruction possible: an auditor, a customer representative or a new engineer can answer why the machine did that on that day from records rather than from the memory of whoever built it. Both properties are the same instinct that produced control plans, applied to a new kind of decision-maker.

The best-documented industrial precedent for this architecture is not in automotive at all, and it is worth reading precisely because of that. Google DeepMind's account of putting an AI system in direct control of data-centre cooling (opens in a new tab) describes an arrangement that matches the stack above almost layer for layer: the learned system proposes actions, those actions are verified against constraints defined by the operators before they reach the plant, local control systems retain their own limits, and staff can hand control back to the conventional system at any moment. The reported saving is substantial and the architecture is unremarkable — which is exactly the point. What transfers to a car plant is the shape, not the number: thermal systems are forgiving, reversible and free of homologation consequence in a way that a paint oven, a robot cell or a fastening station is not.

It is also worth noticing how old this idea is inside the industry. Toyota's account of its own production system (opens in a new tab) describes jidoka — equipment that stops itself the moment an abnormality occurs, so that a person decides what happens next — as one of its two pillars, and that is structurally the same claim the safety envelope makes: the authority to stop lives outside the thing making the decisions, and it is not negotiable by performance. Teams that present a bounded policy to an automotive audience in that language tend to be understood immediately. Teams that present it as autonomy tend to spend two more meetings than they needed to.

Evidence: how to know a policy is better before it acts

The four instruments that make control authority arguable — divergence rate, counterfactual regret, envelope rejection and sim-to-real gap — with the formula, the source system and the stage each one becomes honest.

You know a policy is better than the rule it would replace by comparing them counterfactually on your own logged decisions, with an interval around the answer — never by comparing training curves or simulated scores. The method has a name outside manufacturing, off-policy evaluation, and one idea inside it: because the log records which action was taken in which state and what followed, a candidate policy can be scored on the decisions it would have taken differently, weighted by how much evidence exists for them. Where the log is rich and the divergence modest, the estimate is tight. Where the policy wants to do something the plant has rarely done, the interval opens up — and that widening is the most useful output, because it enumerates exactly the situations to watch when the policy first runs.

Four instruments carry the argument from shadow through to bounded authority, and they should be reported together every week, in the plant's own units. Divergence rate says whether there is a disagreement worth investigating at all. Counterfactual regret says which direction and how much, with its uncertainty. Envelope rejection rate says whether the policy is staying inside the bounds it was granted. Sim-to-real gap says whether the environment everything was trained in still resembles the line. Any one alone is misleading: a low regret estimate with a high divergence rate means the policy is doing something different for no benefit, and a healthy regret estimate with a widening sim-to-real gap means the number is about a plant that no longer exists.

InstrumentFormula / readSourceCadenceHonest from
Decision coverageDecisions logged with state, action and alternatives ÷ decisions takenMES decision logDailyStage 1
Sim-to-real gapReplayed KPI − realised KPI, per KPI, over a rolling quarterReplay environment vs MESMonthly, and before every retrainStage 2
Divergence rateDecision points where policy action ≠ incumbent action ÷ all decision pointsShadow logWeeklyStage 3
Counterfactual regretOff-policy estimate of policy value − realised incumbent value, with intervalDecision log + evaluation harnessWeekly during shadowStage 3
Envelope rejection rateFiltered actions ÷ proposed actions, split by which limit was hitAction filter logWeeklyStage 4
Policy write coverageDecisions written by the policy ÷ decisions in the policy's stated scopeMES / controller write logWeeklyStage 4
Reversion countAutomatic and manual reverts to the incumbent, with causeWrite service logPer event, reviewed monthlyStage 4
Constraint-violation countHard-constraint breaches reaching the plant — target zero, alarmedProcess control + filter logPer event, pagedStage 4
Evidence ageDays since the policy's off-policy estimate was last recomputed on current dataControl-policy registerWeeklyStage 5
Instrumentation build sheet for the four core control-authority KPIs plus the operational monitors. 'Honest from' is the ladder stage at which the number first measures something real; before that stage it is either uncomputable or meaningless.

The stage transitions themselves are verified by four readings, and each has a threshold that separates the stage beneath from the stage above. They are all readable from the same logs, which is the point — a plant that has instrumented one decision properly can answer all four without a new system.

ReadingSimulatedShadowedBounded authorityHow to read it
Where the policy runsIn a simulatorOn live state, no writeIn the write path, filteredWhich service holds the action the plant executes
Comparison usedReward curveOff-policy estimate on logged decisionsOff-policy estimate plus a live holdoutWhether the comparison is counterfactual and carries an interval
Constraint enforcementReward penaltyReward penalty, plus a written draft envelopeDeterministic filter outside the modelRead the code in the write path, not the training config
Revert pathNone neededNot applicable — the incumbent is still actingOne switch, drilled on a scheduleThe date the revert was last exercised on purpose
Verification readings for each transition on the control-authority ladder. All four come from telemetry rather than from self-report, which is why they survive a review.

Shadow-to-authority readiness checklist

If you cannot tick all seven, the policy is not ready for write access however good the regret estimate looks. Tick as you go — this list works without JavaScript.

0 of 7 ticked

Nothing ticked — you are not at the write-access conversation yet

That is the normal position, and the right response is not to work down this list in order. Pick one decision, log it properly for a quarter, and run the 90-day plan below. Six of these seven items fall out of doing that once, and the seventh is a document you will be able to write because you will finally know what the bounds should be.

Governance: a policy that writes is a process change

Automotive already has a regime for changing how a process behaves. The work is not inventing a new one — it is showing that a learned policy produces the same artefacts a control-plan change has always produced.

A learned policy that writes into production is a process change, and automotive has spent decades building the machinery for those — change control, capability evidence, part-approval discipline, process audits and customer-specific requirements. The productive move is to map the policy onto that machinery rather than to argue for an exception to it. In practice this means every artefact a quality manager expects for a control-plan change has a policy equivalent: what changed, why, what evidence supports it, who approved it, how it is monitored, and how it is reversed. Programmes that arrive with those six answers get a hearing. Programmes that arrive with a model card and an accuracy figure do not, and the difference is not fairness — it is that the second set of documents answers none of the questions the regime exists to ask.

Regime or frameworkWhat it asks forThe artefact a bounded policy producesWho signs
IATF 16949 change controlChanges to the process are planned, evidenced, approved and verified, with customer notification where requiredChange record carrying policy version, reward spec, envelope, off-policy evidence and revert planQuality manager and process owner
Control plan and reaction planThe window a parameter may occupy and what happens when it leavesThe written envelope — the same window, plus rate limits and suppression conditionsProcess engineering
VDA 6.3 process auditEvidence that the process is controlled and that changes are managed at the point they happenFilter logs, envelope rejection trend, reversion log and the register entryProcess owner, evidenced to the auditor
Machinery safety validationProtective functions rated, validated and independent of application softwareA design record showing the policy is outside every safety function's path, and validation unchangedSafety engineer
EU AI Act, Annex I regulated productsRisk management, data governance, logging, human oversight and technical documentation for high-risk AI in regulated productsRegister entry, decision and action logs, envelope, human-oversight design and evaluation evidenceWhoever holds product-compliance accountability
NIST AI Risk Management FrameworkGovern, map, measure and manage — an operating discipline rather than a certificateReward and gaming-mode documentation, the four instruments, monitoring and the paging pathThe standing policy forum
TISAX information-security assessmentAssessed handling of sensitive process and product information shared between partnersAccess, retention and exchange controls on the decision log and the replay environmentInformation security, with plant IT
The obligation map: what each regime asks for, and the artefact a properly bounded policy produces as a by-product. None of these frameworks prohibits a learned policy in a plant decision; each requires that the decision be bounded, evidenced and reconstructable.

Two of those rows are worth expanding because they are the ones teams most often get wrong. The regulatory clock first: the EU AI Act (opens in a new tab) became applicable in August 2026, and the European Commission's own account of the timeline gives AI embedded into regulated products under Annex I — the category that reaches machinery — an extended transition to 2 August 2028 following the AI Omnibus agreement. That is not permission to defer the work, because the obligations it describes are risk management, data governance, logging, human oversight and technical documentation, and every one of them is an artefact a bounded policy generates anyway. A plant that reaches stage 4 properly has most of the file already; a plant that reaches it improperly will be assembling it under a deadline. Being ready later is, again, mostly a matter of doing the present work properly.

The second is the simulator, which is almost always treated as a technical asset and is in fact an information-security one. A replay environment good enough to train a policy contains process windows, cycle times, defect patterns, supplier lot behaviour and volumes — a fairly complete description of how the plant makes money, frequently built or hosted with a partner. TISAX (opens in a new tab) exists precisely to make that kind of exchange assessable rather than negotiated case by case, and the same discipline applies to the decision log. On the quality side, IATF 16949 (opens in a new tab) sets the change-control expectations the register is designed to satisfy, VDA 6.3 (opens in a new tab) process audits probe exactly the point-of-change discipline the filter and its logs evidence, and the NIST AI Risk Management Framework (opens in a new tab) gives a vocabulary — govern, map, measure, manage — that maps cleanly onto the register, the candidacy screen, the four instruments and the paging path. Nothing here requires a new governance function. It requires the existing one to be handed documents it recognises.

What the control-policy register holds

  1. Identity and scope

    Which decision, on which line, at which cadence, and the explicit boundary of what the policy may and may not decide. Written so that a person who joins in two years can tell whether a new situation is inside scope without asking the team that built it.

  2. Version, reward specification and envelope

    The policy version in production, the reward specification it was trained against with its signs and weights, and the envelope it writes inside. All three are versioned documents with owners, because a change to any one of them changes what the plant will do.

  3. Evidence and approval

    The off-policy estimate with its interval, the shadow period's divergence and regret record, the holdout result if one exists, and the named approver. This is the section an auditor reads, and it is also the section that makes an internal argument reproducible six months after the enthusiasm has faded.

  4. Monitoring and paging

    Which instruments are watched, at what thresholds, and who is paged when one trips. A monitor with no named recipient is a dashboard, and a dashboard has never stopped an action.

  5. Revalidation date and expiry conditions

    The date by which the evidence must be recomputed, plus the events that expire it early — a launch, a line rebalance, a new colour or variant, a changed upstream release rule, a new shift pattern. Wire this into the launch gateway so it fires as a gateway item rather than after somebody notices a KPI drifting.

A 90-day plan: paint-shop colour sequencing, from Rule-bound to Shadowed

One decision, one paint shop, one quarter — and an honest ceiling. Ninety days buys evidence about your own line, not write access. Anyone promising the second in the first is selling something.

Ninety days is enough to move one decision from Rule-bound to Shadowed, and it is not enough to reach Bounded authority — that transition is gated by change control, envelope validation and a promotion decision that no plan can compress. The decision below is chosen deliberately: paint-shop colour block sequencing is the highest-frequency sequential decision in most car plants, its outcome is already measured in the paint MES, its action space is bounded by whatever is physically in the selectivity bank, and there is no safety function anywhere near it. If a plant is going to learn how to do this at all, this is the decision to learn it on. The quarter contains almost no modelling; it is instrumentation, replay validation, evaluation and a shadow run with a date on it.

Rule-bound → Shadowed on paint-shop sequencing, in one quarter

One shop, one bank, one named process owner. If a phase overruns its window, narrow the scope — one shift pattern, one product family — rather than extending the plan. The deliverable at day 90 is a shadow ledger and a promotion decision, not a deployed policy.

  1. Days 1–15

    Log the decision and baseline the incumbent

    Write a decision record every time the block-building rule fires: bank contents at that moment, the body chosen, every alternative physically available, and the identifiers that join it to the realised outcome. In parallel, baseline the incumbent from the paint MES — colour changes per shift, purge volume per changeover type, block-length distribution, due-date deviation and bank residency — and get the rule itself written down. Name the paint process owner; colour changes and purge are their numbers.

    A decision log that records alternatives, and a written incumbent baseline

  2. Days 16–45

    Build the replay environment and validate it against reality

    Build a discrete-event model of the bank, the oven and flash-off constraints and the upstream release behaviour — from the log rather than from a whiteboard — and include breakdowns, rework re-entry, colour unavailability and manual overrides. Then backtest the model itself: replay the previous quarter's real decisions and publish the reproduction error per KPI. Do not train anything until that error is stated and accepted.

    A published sim-to-real gap per KPI, and an environment nobody has to take on trust

  3. Days 46–70

    Write the reward, train, and evaluate off-policy

    Write and sign the reward specification first — colour changes and purge as the objective, due-date deviation and first-time-through as constraint terms with real weights, bank residency and oven dwell as hard constraints in the draft envelope rather than as penalties. Train, then score the candidate against the logged quarter using off-policy evaluation and report regret with an interval, plus the divergence rate it implies. Name each gaming mode and attach a monitor to it.

    A signed reward spec and a counterfactual regret estimate with its interval

  4. Days 71–90

    Shadow one shift pattern, with a promotion date

    Serve the policy against live bank state at the real cadence, reading exactly the state a deployment would read. It proposes; the rule still acts. Report divergence and regret weekly and review the three highest-regret cases with the process owner every week — that review is where state-feed defects surface. Write the promotion criteria and the promotion date before the first proposal is logged, and treat 'stop' as a legitimate outcome.

    A shadow ledger, a reviewed defect list, and a dated promotion decision

The order matters

  1. Log before you model

    Every day spent building a simulator before the decision log exists is a day spent guessing at what the states and alternatives look like. The log is also the only artefact on this plan that has value regardless of what happens next: it makes ordinary dispatch analysis possible, it makes vendor claims testable, and it is what any future policy will be evaluated against.

  2. Constraints before reward weights

    Decide what must never happen — oven dwell violated, a body held past its residency limit — and put those in the envelope where they cannot be traded away. Only then argue about weights. Teams that start with weights spend the quarter tuning a reward that a hard constraint would have settled in a sentence, and they ship a policy whose worst behaviour is a pricing decision.

  3. Shadow before authority, and put a date on it

    The shadow run is the cheapest evidence in the whole programme and the easiest place to get stuck. Write the promotion criteria and the promotion date into the plan on day one, agree them with quality, and hold to them. A shadow run that ends in a decision to stop has done its job and cost a quarter; one that runs for three years has cost far more than that and decided nothing.

What comes after day 90 is a different kind of quarter, and it is worth saying plainly so nobody plans it optimistically. Reaching Bounded authority on this decision means writing and signing the envelope, building and testing the filter in the write path, wiring the revert and drilling it, assembling the change record with quality, and getting a write-access decision from a process owner who is accountable for the shop's numbers. That is engineering plus a change-control cycle, and it typically runs one to two further quarters. The good news is that nearly all of it is reusable: the second decision on the same line inherits the envelope pattern, the evidence template, the revert mechanism and the register, which is exactly why stage 5 is about breadth rather than height.

Failure modes that send a policy backwards

Authority is not monotonic. Four regressions account for almost all of it, and three are invisible until a KPI moves.

Control authority is not monotonic: policies lose it, quietly and usually for reasons nobody logged. What makes these regressions dangerous is that the system keeps producing actions throughout — a stale policy does not fail loudly, it simply starts being wrong at a rate nobody is measuring, inside an envelope that is still holding, on a plant that has moved. Four patterns account for almost all of it, and each has a cheap preventive measure that costs less than the incident it avoids.

Likelihood: highImpact: high

The environment quietly stops matching the line

A rebalance, a new upstream release rule, an added colour or a changed shift pattern moves the plant away from the distribution the replay environment reproduces. Every subsequent evaluation is computed against a line that no longer exists, and because the environment still runs cleanly, nothing announces the divergence. This is the most common regression and the hardest to see from inside the programme.

PreventionRecompute the sim-to-real gap on a rolling window and gate every retrain on it, with the same seriousness as a capability study.

Likelihood: mediumImpact: high

The envelope is widened so the policy can win

Rejection rates are high, the measured gain is disappointing, and someone loosens a limit. It works immediately, which is why it recurs. The envelope stops being a statement about the process's validated capability and becomes a dial tuned to the policy's ambition — and the first genuinely bad action is then a matter of time rather than of probability.

PreventionEnvelope changes go through the same gate as a control-plan change: written reason, capability evidence, named approver, version and record.

Likelihood: highImpact: medium

Shadow becomes a permanent resting place

No risk, a weekly number and a slide that always survives. Meanwhile the sponsor changes, the state feed drifts, and the evidence being accumulated is about a configuration nobody would now deploy. The programme is not failing in any way anyone can point at, which is exactly why it can continue for years without a decision.

PreventionA promotion date and written evidence thresholds agreed before the first proposal is logged, with 'stop' recognised as a successful outcome.

Likelihood: mediumImpact: high

The reward stops matching the business after a product change

Weights that were correct for last year's mix are wrong for a new variant, a new colour set or a changed order pattern, and the policy pursues the old trade-off with undiminished persistence. The KPI drift is gradual and gets attributed to the launch, which is true in a sense that hides the cause.

PreventionReview the reward specification at every launch gateway alongside process capability, and expire the policy's evidence when the product mix changes.

Glossary

Hover a term for its definition — or expand the map full screen. The full definitions are written out below.

Reinforcement learning
A method for training a policy that chooses actions in sequence and learns from the consequences that follow, including delayed ones — as opposed to supervised learning, which learns from labelled examples of the correct answer.
Policy
The trained decision-maker itself: a function from the observed state of the plant to an action. In an automotive plant a policy is always scoped to one decision — which body to release, which tugger to dispatch, what setpoint to hold — never to a line.
Reward function
The numerical objective a policy maximises, written as signed, weighted terms read from named source systems. Its unwritten terms are the costs a policy will spend, which is why a reward specification is a controlled document rather than a training parameter.
Reward hacking
Satisfying the written objective while making the plant worse — building perfect colour blocks by never releasing scarce colours, or raising throughput by spending maintenance life. In a plant it is nearly always a budget boundary rather than a mathematical subtlety.
Off-policy evaluation (OPE)
Estimating how a candidate policy would have performed on decisions actually taken by a different policy — your incumbent rule — using the decision log. It produces an interval rather than a point, and the width of that interval is where the evidence is thin.
Counterfactual regret
The estimated difference between what a candidate policy would have achieved and what the incumbent actually achieved, over the same logged decisions. The only pre-deployment number that is about your line rather than about a model of it.
Divergence rate
The share of decision points at which the policy would have chosen differently from the incumbent. Near zero means the policy has rediscovered the rule — a legitimate and useful finding that should stop a programme early rather than late.
Sim-to-real gap
The measured difference between the KPIs a replay environment reproduces and the KPIs the plant actually realised over the same period. Unmeasured, it is the reason simulated gains routinely fail to appear on the line.
Action filter (action envelope)
Deterministic code in the write path that checks a proposed action against a written envelope — range, rate limit, hysteresis, suppression conditions — and rejects anything outside it. A constraint expressed only as a reward penalty is a price, not a constraint.
Shadow mode
Running the policy against live plant state at the real decision cadence while the incumbent still acts, logging both actions and the realised outcome. The cheapest source of genuine evidence on the ladder, and the easiest stage to get stuck in.
Model predictive control (MPC)
Choosing actions over a horizon from an explicit written model of the dynamics and constraints. Where such a model exists it is usually the better engineering answer than a learned policy: cheaper to validate, easier to explain and predictable outside its data.
Control-policy register
The governed list of every deployed policy with its version, reward specification, envelope, evidence, approver and revalidation date. The artefact that makes stage 5 exist, and the reason an expired policy gets noticed at a launch gateway rather than in a KPI review.

Frequently asked questions

The questions plant, process and quality teams ask most often when a reinforcement-learning proposal reaches them.

What is reinforcement learning in an automotive plant?

It is a method for training a policy that repeatedly observes plant state, chooses an action and learns from the consequences, including delayed ones. In a car plant it fits a narrow family of sequential decisions: paint-shop colour sequencing, buffer and bank release between shops, AMR and tugger dispatch, and some utilities setpoints. It does not fit inspection, defect classification or immediate parameter setting, which are supervised-learning problems. The distinguishing feature is that a policy chooses an action rather than producing a prediction, which is why its deployment is governed by control authority rather than by accuracy.

Where does reinforcement learning actually work in a car plant today?

In simulation and in shadow, mostly. No automotive manufacturer has publicly documented a plant-floor policy holding unattended write authority over a production decision. What is published is virtual planning and simulation of production systems, robot-learning research conducted in laboratories, and a great deal of machine learning in perceptual and predictive roles. The best-evidenced deployment of a learned control policy with real write authority anywhere is Google DeepMind's data-centre cooling work, which sits outside manufacturing and in far more forgiving physics. Treat any vendor claim of production RL in a plant as a question about which stage of the ladder they mean.

Is reinforcement learning safe to run on a production line?

It is safe when the safety question never reaches it. Protective functions — stops, interlocks, guarding, robot speed and separation monitoring — stay in certified hardware and logic, rated and validated independently, and a policy is never in their path. Above that sits an action filter enforcing a written envelope of ranges, rate limits and suppression conditions, running outside the model in the write path. The design test is one sentence: if the policy started returning random values inside its output range, could anyone be hurt or could the equipment damage itself? Anything other than an immediate no means the boundary is drawn wrongly.

Do we need a digital twin before we can use reinforcement learning?

You need a decision log first and an environment second, and the order is not interchangeable. A simulator built before the log exists encodes guesses about which states occur and which alternatives are available, and a policy trained in it becomes fluent in situations the plant never occupies. Build the log, then build a replay environment from it, then validate the environment by replaying a real quarter and publishing the reproduction error per KPI. An environment that cannot reproduce what your incumbent rule actually achieved has no standing to predict what a new policy would achieve.

What is the difference between reinforcement learning and model predictive control?

MPC chooses actions over a horizon from an explicit written model of the dynamics and constraints; RL learns a policy from experience where no such model can be written. Where your process engineers can state the dynamics — booth conditioning, oven profiles, much of the utilities estate — MPC is usually the better answer: cheaper to validate, far easier to explain to a safety engineer, and predictable outside its data because it has none. RL's honest domain is decisions tangled with upstream release rules, breakdowns, product mix and human overrides that nobody has succeeded in writing down, but which repeat often enough to learn from.

Why is reward design the hardest part?

Because a policy pursues exactly what is written with a persistence no person sustains, and the terms that are missing are the ones it will spend. In a plant the missing terms are almost always a budget boundary: a sequencer rewarded on colour changes spends due-date performance, which is logistics' number; a release policy rewarded on availability spends component life, which is maintenance's. The remedy is procedural rather than mathematical — list the ledgers the decision can touch, give each a term or an explicit written exclusion, put anything that must never happen in the envelope rather than in the reward, and name the gaming modes before training.

How do you evaluate a policy without letting it run?

With off-policy evaluation on your own decision log, followed by a shadow run. Off-policy evaluation scores a candidate on the decisions your incumbent rule actually took, weighted by how much evidence exists for the alternatives, and produces an interval rather than a point estimate. Shadow then serves the policy against live state at the real cadence while the incumbent still acts, producing a divergence rate and a regret estimate on real states nobody simulated. Together they answer the only question that matters before write access: on this line, over this period, would it have been better, and how sure are we?

Does the EU AI Act apply to a learned policy controlling a machine?

It can, and the timeline is now specific. The Act became applicable in August 2026, and the European Commission's account of the phase-in gives AI embedded into regulated products under Annex I — the category reaching machinery — an extended transition to 2 August 2028 after the AI Omnibus agreement. The obligations it describes are risk management, data governance, logging, human oversight and technical documentation. A policy deployed properly at Bounded authority produces all of those as by-products: the register entry, the decision and action logs, the envelope, the human-oversight design and the evaluation evidence. Confirm your own classification with whoever holds product-compliance accountability.

How does IATF 16949 change control apply to a policy that retrains?

A retrained policy is a different decision-maker, so it is a process change and it carries a process change's evidence. That means a change record with the policy version, the reward specification it was trained against, the envelope it writes inside, its off-policy evidence, a named approver and a revert plan — the same six answers a control-plan change has always required. The common failure is treating retraining as housekeeping because the pipeline is automated. If the change discipline feels too heavy for the retraining cadence you want, the honest conclusion is that the decision should not be automated at that cadence.

Should a policy explore in production?

Essentially never in an automotive plant. Exploration means deliberately taking an action believed to be suboptimal in order to learn from it, and on a line the cost is measured in vehicles while the information is usually already available in the decision log. Learn offline from the log and from a validated replay environment. The one defensible exception is a tightly bounded contextual-bandit arrangement on a genuinely reversible decision with a very small action set — and even then the exploration budget belongs in the written envelope with a rate limit, not in the training configuration.

How long does it take to go from a rule-based line to a policy with bounded authority?

About a quarter to reach Shadowed on one decision, and typically one to two further quarters to reach Bounded authority. The first ninety days are instrumentation, replay validation, reward specification and a shadow run with a promotion date. What follows is engineering plus a change-control cycle: the envelope written and signed, the filter built and tested in the write path, the revert wired and drilled, the change record assembled with quality, and a write-access decision from an accountable process owner. Nearly all of that is reusable, which is why the second decision costs a fraction of the first.

What is the most common reason an automotive RL project fails?

The decision was never logged, so nothing could be evaluated and the argument defaulted to whoever presented most persuasively. Second most common is a simulated gain that nobody replayed against real states, which evaporates on contact. Third is a single-term reward that bought its headline KPI with a cost in another department's ledger. All three are failures of instrumentation and specification rather than of modelling, which is why the first ninety days of a credible programme contain almost no model development at all.

About the author

Atomic Loops Engineering

Industrial AI practice

Atomic Loops builds production AI systems for manufacturing, logistics and energy operators — vision inspection, predictive quality, sequencing and decision support running against live plant data, integrated into the MES and controls layer rather than delivered as dashboards.

  • · Production deployments across body shop, paint, machining and final assembly
  • · Decision instrumentation and off-policy evaluation run jointly with process engineering
  • · Integration-first delivery: MES and controller write-back, action filtering, drilled revert
  • · 13 cited sources on this page

Sources

  1. Richard S. Sutton and Andrew G. BartoReinforcement Learning: An Introduction (opens in a new tab)
  2. Google DeepMindSafety-first AI for autonomous data centre cooling and industrial control (opens in a new tab)
  3. NISTAI Risk Management Framework (opens in a new tab)
  4. European CommissionRegulatory framework for AI (EU AI Act) and its application timeline (opens in a new tab)
  5. International Automotive Task ForceIATF 16949 and global oversight (opens in a new tab)
  6. VDA QMCQuality management standards and process audits (VDA 6.x) (opens in a new tab)
  7. ENX AssociationTISAX — trusted information security assessment exchange (opens in a new tab)
  8. ACEAEuropean automotive industry data (opens in a new tab)
  9. BMW GroupBMW Group PressClub (virtual production planning and plant AI) (opens in a new tab)
  10. Toyota Motor CorporationToyota corporate newsroom (opens in a new tab)
  11. Toyota Motor CorporationToyota Production System (jidoka and just-in-time) (opens in a new tab)
  12. Toyota Research InstituteRobot behaviour-learning research (opens in a new tab)
  13. General MotorsGM Newsroom (machine learning in manufacturing operations) (opens in a new tab)

Find out what one decision is actually worth — before anything writes

We take one sequential decision on your line, check what is genuinely logged about it, replay a quarter through whatever environment you have, and produce a counterfactual estimate with its interval and a shadow plan with a promotion date. You keep the instrumentation specification and the findings whether or not we build anything.

Published · Last updated

Benchmark request

Tell us where to send it

Benchmark for this page

Used once, to send this benchmark and follow it up personally. No newsletter, no automated sequences.