Manufacturing (Automotive)AI Implementation & Best Practices
Reinforcement learning in automotive plants: what a learned policy is allowed to control
Reinforcement learning is a method for training a policy that chooses actions in sequence and learns from the consequences they produce, rather than from labelled examples. In an automotive plant it fits a narrow set of repeated, measurable, bounded decisions — and how much authority that policy holds is a governance question long before a modelling one.

Key takeaways
- Reinforcement learning fits sequential plant decisions whose consequence arrives later and whose outcome the plant already records — paint-shop colour sequencing, buffer release between shops, AMR dispatch, utilities setpoints. Perception and prediction problems are not RL problems, and treating them as such is the most common category error in plant AI.
- Maturity here is control authority, not model quality. The ladder runs Rule-bound, Simulated, Shadowed, Bounded authority, Governed autonomy — and the public automotive record stops at Simulated and Shadowed. No OEM has published a plant-floor policy holding unattended write authority.
- The safety envelope is always enforced outside the policy. Protective functions stay in the certified hardware and logic they already live in; the policy proposes a value inside a window the safety system would hold even if the policy output were random. A constraint expressed only in a reward term is not a constraint.
- The honest KPI is counterfactual, never a training curve. Divergence rate against the incumbent rule plus off-policy evaluation on the logged decisions — with an interval, not a point estimate — is what tells you whether the policy would have been better on your line.
- Reward design is where automotive RL fails, and it fails as a budget boundary: a single-term reward buys throughput by spending quality, maintenance life, energy or due-date performance that lands in another department's ledger. Write the constraint terms and the gaming modes before the first training run.
Abbreviations used on this page
- RL
- Reinforcement learning
- MDP
- Markov decision process (the state–action–reward formalism RL assumes)
- OPE
- Off-policy evaluation (scoring a policy on decisions it did not take)
- MPC
- Model predictive control
- PLC
- Programmable logic controller
- MES
- Manufacturing execution system
- AMR
- Autonomous mobile robot (lineside material supply)
- OEE
- Overall equipment effectiveness
- FTT
- First-time-through rate
- EOL
- End-of-line (test)
- PL d
- Performance level d — a machinery safety-function rating under ISO 13849
- IATF
- International Automotive Task Force (IATF 16949)
What reinforcement learning is in an automotive plant — and what it is not
A definition, the loop as it actually exists on a line, and the thesis of this page: on a car line, reinforcement learning is a control-authority question long before it is a modelling question.
Reinforcement learning is a way of training a decision-maker — a policy — that repeatedly observes the state of a system, chooses an action, and learns from the consequences that follow, including consequences that arrive many steps later. It is not trained on labelled examples of the right answer, because for most sequential problems nobody knows the right answer; it is trained on outcomes. In an automotive plant that description fits a specific and fairly small family of decisions: which body to pull next out of the paint shop's selectivity bank, when to release a buffer between shops, which tugger to send to which lineside call, what to do with a booth setpoint for the next quarter of an hour. It fits none of the plant's perception problems and few of its prediction problems, and the single most common category error in industrial AI is using it where a supervised model or a well-posed optimisation would do the job with a hundredth of the risk.
Three neighbours are worth separating from it precisely, because plant conversations blur them constantly. Supervised learning maps an input to a label — this weld is porous, this claim is this failure mode — and its consequence is immediate and observable, which is why inspection and prediction are supervised problems. Model predictive control chooses actions over a horizon too, but it does so from an explicit written model of the dynamics and constraints, and where such a model exists it is almost always the better engineering choice: cheaper to validate, easier to explain to a safety engineer, and behaves predictably outside its training data because it has none. Reinforcement learning earns its place in the gap between the two — decisions whose consequence is delayed and whose dynamics nobody can write down, but which the plant repeats often enough, and logs well enough, to learn from experience.
Reinforcement learning is learning what to do — how to map situations to actions — so as to maximize a numerical reward signal.
That definition contains the whole problem of putting it in a plant. "Maximize a numerical reward signal" is a promise to pursue whatever is written down with more persistence, and less common sense, than any rule or any person. A shift leader told to reduce colour changes will not starve the assembly line of a due model to do it; a policy will, unless due-date deviation is a term in its reward with a sign and a weight. And "map situations to actions" is a promise to act, which is what separates this page from every other AI topic in this series. A vision model that is wrong produces a false reject and an operator's irritation. A policy that is wrong produces an action, in a plant, on a machine — which is why the rest of this page is largely about what stands between the policy's output and the equipment, and why the ladder that organises it measures authority rather than accuracy.
Decision value released along the control-authority ladder
The curve is not linear, and its flat section is longer than most programmes plan for. Value stays close to zero through Rule-bound and Simulated — where simulated gains are claims about a model — and only begins to release at Shadowed, the first stage that produces evidence about your own line. The steep segment is Bounded authority, where an action is actually written; Governed autonomy adds breadth rather than height, because the second and third policies reuse the envelope, the evidence pattern and the register built for the first.
Decision value released by stage
- Stage 1 · Rule-bound — 34% of operators. Every sequential decision on the line is made by a fixed rule or a person, and the decision itself is not recorded — so no policy can be evaluated and no baseline exists to beat.
- Stage 2 · Simulated — 31% of operators. A model of the decision exists — a replay of the log, a discrete-event model of the line or a robot simulation — and policies train in it, but nothing has ever run against the plant.
- Stage 3 · Shadowed — 22% of operators. The policy runs on live plant state and proposes an action at every decision point while the incumbent still acts, and both the divergence and its counterfactual value are logged.
- Stage 4 · Bounded authority — 10% of operators. The policy writes its decision into the MES or the controller inside an envelope enforced outside the model, with a named approver, a drilled one-switch revert to the incumbent rule, and a change record in the quality system.
- Stage 5 · Governed autonomy — 3% of operators. Several policies hold bounded authority under a versioned control-policy register, retraining and redeployment ride the plant's change gateway, and people set the reward and the bounds rather than the individual actions.
Curve shape: logistic, plotted from the stage data above. Distribution: Consistent with the published record of bounded industrial RL deployments.
The reinforcement-learning loop as it actually exists on a line
Three lanes, and only the middle one touches the plant. The training lane is entirely offline: it learns from the decision log and is evaluated counterfactually. The live lane runs left to right at the real decision cadence, and every action passes an action filter before any write occurs. The governance lane owns the two artefacts that make the whole thing approvable — the reward and envelope specification, and the change record — while the safety-rated stop sits outside the loop altogether and is unchanged by any of it.
- Data & feeds
- AI / model
- Where value leaks
- System-of-record action
- Human in the loop
The process, in words
- The training lane never touches the plant. It reads the decision log — state, action taken, alternatives available, realised outcome — replays it through an environment that has been validated against real production, trains a policy, and scores that policy counterfactually. A policy is promoted out of this lane only on evidence, never on a training curve.
- The live lane runs at the real decision cadence. Plant state feeds both the incumbent rule and the learned policy; the policy proposes one action; the action filter checks it against a written envelope and either passes it to the MES or controller write, or rejects it and lets the incumbent's action stand. Every write carries the policy version, and every realised outcome goes back into the log.
- The governance lane contains nothing that is learned. The reward and envelope specification is written and signed before training begins, and it constrains both the training lane and the filter. Every write produces a change-controlled record — policy version, evidence, approver, revalidation date.
- The safety-rated functions sit outside the loop by design. Protective stops, interlocks, guarding and robot speed-and-separation monitoring are unchanged by the presence of a policy, act on the equipment regardless of what any software proposed, and are validated independently. If a proposed architecture requires the policy to be trusted for safety, the architecture is wrong.
Step-by-step insights
- The decision log — the artefact that makes everything else possible
- The record has four parts and all four are load-bearing: the state at the moment of the decision, the action that was actually taken, the alternatives that were genuinely available, and the outcome that followed. Most plants have the fourth and none of the first three. The alternatives matter more than teams expect, because off-policy evaluation is a comparison against the options that existed, and a log without them can only tell you what happened, never what could have. A stable decision identifier joining all four is what turns a stream of events into a dataset — and it is the same discipline that the unit-level data backbone imposes elsewhere in the plant, which is why programmes with a serialised production record reach usable logs months faster.
- The replay environment — validate the model before you trust the policy
- The environment is the second-most-abused artefact in industrial RL, after the reward. The correct first use of it is not training but backtesting itself: run last quarter's real decisions through it and check that it reproduces the realised KPIs within a tolerance you publish. Nominal-flow models are the usual failure — no breakdowns, no rework re-entry, no manual overrides, no changeover — and a policy trained in one becomes fluent in states the plant never occupies. Adding disruption almost always reduces the simulated gain and increases its truthfulness, which is an uncomfortable conversation and the right one.
- Off-policy evaluation — the only honest pre-deployment number
- Off-policy evaluation estimates how a new policy would have performed on decisions taken by a different policy — your incumbent rule. It is the industrial equivalent of a holdout, and it comes with an interval that widens where the log is thin or where the new policy diverges hardest. That widening is the useful part: it tells you exactly which regions of the decision space you have no evidence about, which is the list of situations to watch during shadow. Treat a point estimate with no interval as a marketing artefact rather than an evaluation, whether it comes from a vendor or from your own team.
- The incumbent rule — the baseline that never leaves
- The incumbent stays live for the whole life of the policy, not just during shadow. It is the fallback the filter reverts to on a rejected action, the destination of the one-switch revert, and the comparison the regret number is computed against. Programmes that decommission the rule at promotion lose all three at once, and discover on the first bad week that they have nothing to fall back to and no way to say how bad the week actually was. Keeping a rule alive costs almost nothing; recreating one under pressure costs a quarter.
- The action filter — a constraint in the reward is not a constraint
- This is the single most important architectural claim on the page. Constraints expressed as reward penalties are preferences: a policy that finds enough upside will pay the penalty, and a policy whose inputs have gone stale will violate them without knowing. The filter is deterministic code, outside the model, in the same service that performs the write, and it checks range, rate of change, hysteresis and suppression conditions before anything reaches the plant. It is also the artefact that makes the deployment reviewable, because a quality or safety engineer can read it, test it and sign it without understanding the policy at all.
- Safety-rated functions — the layer no policy may enter
- Machinery-safety practice puts protective functions in certified hardware and logic, rated for the required performance level, validated independently of whatever application software exists. A learned policy changes none of that: it may choose a setpoint inside a validated window, but it may not be the thing that stops a machine, opens a guard or relaxes speed and separation limits. The practical test is simple and it belongs in every design review — if the policy stopped working, or started returning random values, would anyone be hurt? If the answer is anything other than an immediate no, the boundary has been drawn in the wrong place.
The control-authority ladder: five stages
For each stage: what it looks like on the ground, the diagnostic signals a reviewer can check in an afternoon, the anti-pattern that traps plants there, and what leaving actually costs.
The ladder below measures how much authority a learned policy holds over a real decision, from Rule-bound — where the decision is not even recorded — to Governed autonomy, where several policies write inside stated bounds under a versioned register. It deliberately does not measure model quality, because model quality is not what makes a policy safe, useful or approvable in an automotive plant; a mediocre policy inside a well-designed envelope is a manageable process change, and an excellent policy with unbounded write access is an incident waiting for its Tuesday. Each stage is written for a practitioner: the hallmarks are observable conditions, the diagnostic signals are checks you can run against your own MES and controls estate this week, and the anti-pattern is the specific mistake most often made trying to leave that stage.
One thing to notice as you read. The stages get harder in the opposite order to the one most programmes assume. Stages 1 and 2 are engineering problems with known solutions and can be bought with effort. Stages 3, 4 and 5 are evidence and governance problems, and they are gated by people who will not be persuaded by a training curve — process owners, quality engineers, safety engineers, and eventually a customer auditor. That is why a plant with a mediocre model and an excellent envelope will ship, and a plant with a state-of-the-art policy and no written bounds will not.
Select a stage
Every stage's full detail is in the page source — the selector only changes which panel is visible, so nothing here depends on JavaScript to exist.
Stage 1
Rule-bound
34% of operators sit here
Every sequential decision on the line is made by a fixed rule or a person, and the decision itself is not recorded — so no policy can be evaluated and no baseline exists to beat.
Rule-bound is the correct default and the honest starting point for almost every automotive plant, and it deserves respect before critique. The rules are usually good: decades of industrial engineering are compressed into a paint shop's block-building heuristic, a conveyor's release logic and a fleet manager's priority queue, and they have been tuned against real breakdowns by people who watched the line. A learned policy that beats them is doing something genuinely difficult, not something obvious. The problem at stage 1 is not that the rules are bad. It is that the plant keeps no record of them operating, so there is no way for anyone — including the people who wrote them — to ask whether a different choice would have been better.
The tell is the decision log, and it is a five-minute diagnostic. Ask for one week of one specific decision: which body the paint sequencer pulled next out of the selectivity bank, which tugger the fleet manager assigned to which lineside call, when the shift leader released the buffer between body shop and paint. Ask for it with the state at the moment of the decision and the alternatives that were physically available. At stage 1 the realised outcome is in the MES — colour changes per shift, starvation minutes, kilowatt-hours — and the decision is nowhere. The plant knows what happened and not what was chosen, which is precisely the wrong half of the pair for any technique that learns from consequences.
This stage is cheap to leave for one decision and expensive to leave for a plant, and the difference is worth being disciplined about. Instrumenting a single decision point is a controls-and-MES task measured in weeks: write the state, the action, the alternatives and a stable decision identifier to a log every time the rule fires. The common mistake is to treat that logging as the first phase of an AI project and therefore to fund it, staff it and justify it as one. It is process instrumentation, it belongs to industrial engineering as much as to data science, and it pays for itself in ordinary dispatch analysis — how often does the rule take the option the shift leader would have overridden? — long before any policy exists.
In practice
The sequencer nobody could question
A paint shop's block-building rule decides which body leaves the selectivity bank next, several thousand times a week. The shift's colour-change count and purge consumption are in the MES and reviewed every morning. Which bodies were sitting in the bank at 09:14, and which one the rule chose, is recorded nowhere. When a vendor proposes a learned sequencer, the only feasibility question that matters — how much better could anything have done on last quarter's real bank states? — cannot be answered at all. The decision reverts to whoever presents most persuasively, which is exactly how plants end up buying a policy they can neither evaluate nor revert.
What it looks like
- Sequencing, release and dispatch decisions live in PLC logic, an MES rule set or a shift leader's judgement
- The decision is not logged — only the shift's aggregate KPIs survive
- Reinforcement learning appears in strategy decks and vendor demonstrations, never against plant state
- Nobody can say what the incumbent rule chose last Tuesday, or what else was available to choose
Diagnostic signals you can check this week
- Ask for one week of one decision with its state and available alternatives; if the answer is a KPI report, you are here
- Ask for the incumbent rule in writing. Undocumented PLC logic is common, and it is the baseline any policy has to beat
- Ask anyone to state the decision's frequency — decisions per shift is the unit RL economics are actually computed in
- Ask whether any AI in the plant today chooses an action, or only classifies and predicts. Almost always the latter
Anti-pattern · Buying a policy before you own the log
Vendors will offer a sequencer, a dispatcher or an energy optimiser pre-trained on somebody else's plant, and at stage 1 there is no way to test the claim. Without your own decision log you cannot evaluate the policy before it is deployed, cannot detect the week it stops working, and cannot revert with evidence rather than with a shrug. The log is the asset that makes every later vendor conversation an evaluation instead of an argument — and it is yours whether or not you ever train anything. Build it first, and build it for the decision you care about rather than for the plant.
What holds you here
The decision is not recorded with the state it was taken in, so no policy can be evaluated and there is no measurable incumbent to beat.
Highest-leverage next move
Pick one repeated sequential decision and log state, action taken, alternatives available and realised outcome against a stable decision identifier.
Cost of leaving
- Effort
- 6–12 weeks
- Team
- One controls engineer and one MES engineer, part-time, plus the process owner for the decision
- Risk
- Low — logging is additive and nothing in production changes
- To next stage
- 3–6 months
If this is you, the next step is
A four-week scope: pick the decision, define the state and action record, and start the log.
Stage 2
Simulated
31% of operators sit here
A model of the decision exists — a replay of the log, a discrete-event model of the line or a robot simulation — and policies train in it, but nothing has ever run against the plant.
Stage 2 is where automotive reinforcement learning looks most convincing and is least informative. There is a training curve that climbs, a simulated improvement quoted as a double-digit percentage, and a demonstration that runs faster than real time in front of a steering group. Every part of that is real work and none of it is evidence about the plant: a reward improvement inside an environment is a statement about the environment. The gap between the two is the sim-to-real gap, and at stage 2 it is almost never quantified — not because teams are dishonest, but because quantifying it requires the decision log that stage 1 was supposed to produce.
The discipline that separates a useful stage 2 from a decorative one is environment validation, and it runs backwards from what people expect. Before you evaluate a policy in the environment, evaluate the environment itself: replay last quarter's real decisions through it and check that it reproduces the realised KPIs — colour changes per shift, starvation and blocking minutes, kilowatt-hours per unit, due-date deviation — within a stated tolerance, and publish the tolerance. An environment that cannot reproduce what the incumbent rule actually achieved has no standing to tell you what a new policy would achieve. This single check retires more bad RL proposals than any amount of model review.
There is a second, quieter trap: the environment usually omits exactly the things that make a real automotive line hard. Breakdowns and micro-stops, rework bodies re-entering the sequence out of order, the shift leader's manual overrides, the Thursday changeover, the supplier delivery that arrives on the wrong pallet, the colour that is temporarily unavailable because a tank is being cleaned. A model of the line on a good day trains a policy for a plant that does not exist, and the policy's advantage is often precisely a fluency in states the plant never occupies. Adding disruptions to the environment usually reduces the simulated gain and increases its usefulness, which is an uncomfortable trade to present and the right one to make.
In practice
The gain that did not survive contact
A sequencing policy showed a large simulated reduction in colour changes against the plant's block-building rule, sustained across a long training run. When the same policy was replayed against the previous quarter's actual bank contents, most of the advantage disappeared: the gain had been earned in states where four bodies of the same colour were simultaneously available in the bank, which the simulator generated freely and the upstream release rule makes rare. The residual advantage on real states was real, small, and worth pursuing — but it was a different number, attached to a different argument, and only the replay could tell them apart.
What it looks like
- A replay or simulation environment exists, and a policy trains in it to a better reward than the incumbent rule
- The environment's fidelity against realised production has never been measured as a number
- Results are reported as reward curves and simulated percentage gains
- No live plant system reads the policy's output
Diagnostic signals you can check this week
- Ask what the sim-to-real gap is, as a number, on the KPI the business case rests on
- Check whether the environment contains breakdowns, rework re-entry and manual overrides, or only nominal flow
- Check that the incumbent rule is implemented inside the same environment — without it there is no comparison, only a score
- Ask whether the simulated gain carries a confidence interval or is quoted as a single figure
Anti-pattern · Optimising the policy instead of the environment
When simulated gains fail to reproduce, the reflex is to train harder — a bigger network, a longer run, a different algorithm, more reward shaping. It almost never helps, because the fidelity of the environment sets the ceiling on everything trained inside it, and a policy that exploits a modelling artefact will exploit it more efficiently the better it gets. Spend the next month on replay validation and disruption modelling rather than on architecture search. The uncomfortable version of this rule: if the environment cannot reproduce the incumbent's realised KPIs, the correct next step is not a better policy but a better model of your own line.
What holds you here
The environment has never been validated against realised production, so a simulated gain is a claim about the model rather than about the line.
Highest-leverage next move
Replay a full quarter of logged decisions through the environment, publish the reproduction error per KPI, and only then compare a policy against the incumbent rule.
Cost of leaving
- Effort
- 3–6 months
- Team
- One simulation or operations-research engineer, one ML engineer, the process owner
- Risk
- Medium — the real risk is spending a year inside a simulator nobody ever validates
- To next stage
- 3–6 months
If this is you, the next step is
We replay your own decision log through your model and report the reproduction gap, KPI by KPI.
Stage 3
Shadowed
22% of operators sit here
The policy runs on live plant state and proposes an action at every decision point while the incumbent still acts, and both the divergence and its counterfactual value are logged.
Shadow is the first stage that produces evidence about your plant rather than about a model of it, and it is astonishingly cheap for what it yields. The policy is served against live state, at the real cadence, seeing the states nobody simulated: the micro-stop at 03:40, the rework body that re-entered the sequence, the colour that went unavailable mid-shift. It proposes an action at every decision point. The incumbent rule still acts. Nothing in the plant changes, no approval is needed beyond read access, and within a fortnight the programme has something it has never had before — a paired record of what was chosen and what would have been chosen, on real states.
Two numbers come out of a shadow run, and they are this page's core instruments. Divergence rate — the share of decision points where the policy would have chosen differently — tells you whether there is anything to argue about at all. A divergence rate near zero is a genuine and useful result: the policy has rediscovered the rule, the rule is good, and the programme should stop, which is worth knowing in month three rather than month thirty. Counterfactual regret, estimated by off-policy evaluation over the logged decisions, tells you the sign and the size of the difference with an honest interval around it. Where the log is thin or the policy diverges often, that interval is wide, and a wide interval is information rather than a failure.
Shadow's failure mode is that it never ends. It is comfortable: no operational risk, a weekly number that trends the right way, a slide that survives any steering group. Meanwhile the sponsor moves on, the state feed drifts, and the policy silently starts scoring against a plant that has been rebalanced twice. The remedy is procedural and belongs in the shadow plan on day one — a date, an evidence threshold and a named person who decides to promote or to stop. The other thing to plan for is the shadow run's most valuable by-product: divergences that turn out to be defects in the state feed rather than disagreements about strategy. Reviewing the largest-regret cases with the process owner every week is where those surface, and each one is worth more than a month of training.
In practice
The divergence that was a data bug
In the first fortnight of shadow on a buffer-release decision, the policy diverged from the incumbent on roughly two decisions in five — far more than anyone expected from a well-tuned rule. The weekly review of the largest-regret cases found the cause in three sessions: the policy was reading bank contents from a feed that lagged reality by two conveyor positions, so it was consistently reasoning about a bank that had already moved on. The fortnight's real product was not a better policy but a corrected state feed, which the incumbent rule had been silently tolerating for years because a rule that only looks at the head of the queue never notices.
What it looks like
- The policy is served against live MES and controller state at the real decision cadence
- Every proposal is logged next to the incumbent's action and the realised outcome
- Divergence rate and counterfactual regret are reported weekly — reward is not reported at all
- A promotion decision exists, with pre-agreed evidence thresholds and a date on it
Diagnostic signals you can check this week
- Ask for last week's divergence rate and the three highest-regret decision points, by name
- Check that the policy reads exactly the state a live deployment would read, at the same latency — a shadow run on enriched data proves nothing
- Ask whether the promotion criteria and the promotion date were written before shadow started
- Ask who reviews divergent cases with the process owner, and how often — that review is where policy and data defects surface
Anti-pattern · Shadow as a permanent parking space
Because shadow carries no operational risk, it also carries no forcing function, and programmes settle into it for years. The tells are recognisable: the weekly report still exists but nobody reads the divergence review, the promotion criteria were never written down, and the answer to 'what would it take to switch this on?' is a discussion rather than a document. Fix it at the start, not at the end. Write the envelope specification, the evidence threshold and the promotion date into the shadow plan before the first proposal is logged, and treat a decision to stop as an equally successful outcome — because a policy that cannot beat a good rule is a finding, not a failure.
What holds you here
The policy has no write path and no agreed promotion criteria, so evidence accumulates indefinitely without ever converting into authority.
Highest-leverage next move
Write the action envelope and the promotion criteria, have the process owner and quality sign them, and put a date on the promotion decision.
Cost of leaving
- Effort
- 3–6 months
- Team
- ML engineer, MES and controls integration engineer, process owner, plus quality for the promotion criteria
- Risk
- Low operationally, high organisationally — the real risk at this stage is stalling in it
- To next stage
- 6–12 months
If this is you, the next step is
Evidence thresholds, promotion date and the envelope specification, agreed before the first proposal is logged.
Stage 4
Bounded authority
10% of operators sit here
The policy writes its decision into the MES or the controller inside an envelope enforced outside the model, with a named approver, a drilled one-switch revert to the incumbent rule, and a change record in the quality system.
Bounded authority is the first stage at which the policy changes what the plant actually does, and it is defined entirely by what constrains it rather than by what it computes. The deliverable of this stage is the envelope, not the model. An envelope is a written artefact: the range each action may take, the rate at which it may change, the hysteresis that prevents oscillation, the conditions under which the write is suppressed entirely, and the signature of the process owner who agreed all of it. The filter that enforces the envelope runs outside the policy, in the same service that performs the write, so that a policy which returns nonsense — because a feature went stale, because a model version was mis-deployed, because the input distribution moved — cannot express the nonsense as an action.
Automotive already knows how to do this, which is the argument that gets the stage approved. A control plan states the window a parameter may occupy and the reaction plan when it leaves; a machine capability study establishes that the process can hold that window; a change to either goes through change control with evidence attached. A bounded policy is a new way of choosing a value inside a window that already existed and was already validated. That framing matters more than any technical detail when the proposal reaches a quality manager: the sentence is not 'we are deploying AI on the line', it is 'we are changing how a setpoint inside the existing control plan is selected, here is the evidence, here is the filter, here is the revert, and the safety validation of the cell is unchanged'. Programmes that lead with the model instead of the envelope get a longer queue and a worse answer.
The engineering that actually consumes the quarter is unglamorous and it is not modelling. It is the filter and its test suite, the fallback path and the switch that selects it, the write logging that records the policy version against every action, the rate limits and hysteresis, the suppression logic for degraded state feeds, and the monitoring that catches the day the policy starts hitting the envelope. That last number — envelope rejection rate — is the most useful single instrument at this stage. A rising rejection rate means either the policy is drifting or the plant has moved outside the conditions the envelope was written for, and it says so weeks before any KPI does. Watch it the way you would watch a control chart, because that is exactly what it is.
In practice
The setpoint that could only move so far
A paint booth's supply-air temperature and humidity setpoints are chosen every fifteen minutes by a policy trained on a year of logged booth conditions, ambient data and finish-defect outcomes. The envelope states that each setpoint may move by no more than a fixed increment per interval and may only sit inside the booth's validated process window; a filter in the write service rejects anything else and falls back to the schedule. Booth condition monitoring trips an automatic revert to the incumbent schedule if conditions leave the window for more than a set period. Nothing about the booth's fire and ventilation safety functions changed, because none of them was ever in the policy's path.
What it looks like
- An action filter outside the policy rejects any action outside stated limits before it is written
- Safety-rated functions are untouched — the protective stop is exactly what it was before the policy existed
- The incumbent rule still runs and is one switch away, and the revert has been exercised deliberately
- Policy version, reward specification, envelope and evidence are recorded as a process change with a named approver
Diagnostic signals you can check this week
- Ask to see the envelope as a written artefact — ranges, rate limits, suppression conditions, and who signed them
- Ask when the revert to the incumbent was last exercised on purpose, not in an incident
- Ask for the envelope rejection rate trend over the last ten weeks
- Ask whether the safety validation of the cell or process changed because of the policy. The correct answer is no
Anti-pattern · Widening the envelope so the policy can win
When the filter rejects a lot of proposals, the tempting fix is to loosen the limits, and it is tempting precisely because it works: the measured gain improves immediately. What has happened is that the envelope has quietly become a function of the policy's ambition rather than of the process's physics and its validated capability. Once that inversion happens the envelope stops being a constraint and becomes a dial, and the first genuinely bad action is not far away. Treat envelope changes exactly like control-plan changes: a written reason, capability evidence, a named approver, a version, and a record. If the physics genuinely permits a wider window, prove it the way you would have proved it without a policy involved.
What holds you here
Every additional policy is a bespoke negotiation with its own envelope, evidence pack and approver, so the second and third decisions cost as much as the first.
Highest-leverage next move
Standardise the envelope, the evidence pack and the revert into a reusable pattern, and open a control-policy register holding every deployed policy's version, bounds and revalidation date.
Cost of leaving
- Effort
- 6–12 months
- Team
- ML engineer, controls and MES engineer, process owner, quality engineer, plus a safety engineer for the validation review
- Risk
- Higher — the policy now writes into a running process, so the filter, the revert and the evidence pack are the critical path
- To next stage
- 12–24 months
If this is you, the next step is
Limits, rate limits, rejection monitoring, revert drill and the change record — the artefacts that make write access approvable.
Stage 5
Governed autonomy
3% of operators sit here
Several policies hold bounded authority under a versioned control-policy register, retraining and redeployment ride the plant's change gateway, and people set the reward and the bounds rather than the individual actions.
Governed autonomy is much narrower than the phrase suggests, and the narrowness is the point. It is not a self-optimising factory. It is an enumerated set of decisions — a sequencing decision, a release decision, a dispatch decision, a utilities setpoint — that a learned policy takes inside stated bounds without a per-decision approval, while everything outside those bounds escalates. Decisions with a safety function in their path, a homologation consequence, or unbounded commercial exposure are correctly held at bounded authority forever, and a plant that has thought clearly about this can say which decisions those are without hesitating. The stage is reached by subtraction: by naming what will never be automated, the remainder becomes governable.
The register is the artefact that defines the stage, and it is the same instinct as a control plan applied to a different object. For each deployed policy it holds the version in production, the reward specification it was trained against, the envelope it writes inside, the evidence that justified promotion, the approver, and — the field that does the most work — a revalidation date. A policy without a revalidation date is a policy nobody will notice has expired, and expiry in a plant is not a theoretical event: it happens on the Monday after a launch, a rebalance or a new colour. Everything else in the register exists so that an auditor, a customer representative or a new engineer can reconstruct why a machine did what it did on a given day, from documents rather than from the memory of whoever built it.
Sustaining this stage is a maintenance discipline rather than a summit, and it is the stage most likely to regress silently. Every product launch, line rebalance, new colour, new supplier, new shift pattern and new upstream release rule changes the distribution the policies were evaluated on, and nothing in the plant's normal rhythm will tell you that a policy's evidence has gone stale. Operators who sustain it wire policy revalidation into the launch gateway next to process capability and part-approval evidence, so a policy that has not been re-evidenced against the new product simply does not get write access on day one of production. The sober conclusion, and the one this page keeps returning to, is that being ready for more autonomy later is almost entirely a matter of doing the present work properly now: log the decision, validate the environment, bound the action, evidence the change.
In practice
The launch that expired four policies
A plant running bounded policies on paint sequencing, buffer release, tugger dispatch and booth conditioning adds a derivative model with a new body variant and two new colours. The launch gateway flags all four policies for revalidation because their evaluation windows predate the variant. Three are re-evidenced by replaying the ramp-up weeks and are back inside their envelopes within a month; the sequencing policy is not, because the new colours change the block economics enough that its regret estimate crosses zero. It reverts to the incumbent rule for the launch and returns to shadow — which is the system working exactly as designed, and would have been an unexplained KPI drift in a plant without a register.
What it looks like
- A control-policy register lists every deployed policy with version, reward specification, envelope, evidence and revalidation date
- Retraining and redeployment pass the same change gate as any other process change
- Regret, divergence, envelope rejection and constraint violations are monitored continuously and page a named human
- Reward specifications are reviewed at every product launch, line rebalance and process changeover
Diagnostic signals you can check this week
- Ask to see the control-policy register and check whether every entry has a revalidation date that has not passed
- Ask when the paging path for a constraint violation was last tested, and who answered
- Check whether the last product launch triggered policy revalidation as a gateway item, or after somebody noticed
- Ask whether an auditor could reconstruct a single automated action — state, policy version, envelope, approver — from records alone
Anti-pattern · Treating retraining as maintenance rather than as change
Once several policies are running well, retraining starts to feel like housekeeping: refresh the data, rerun the pipeline, redeploy. But a retrained policy is a different decision-maker with different failure modes, and deploying it without the evidence and the approval that the first version needed is a process change made without change control — the thing IATF-certified plants spend the rest of their time not doing. The discipline is unpopular and simple: every version passes the same gate, carries its own off-policy evidence, and enters the register with its own revalidation date. If that feels too heavy for the cadence you want, the honest conclusion is that the decision should not have been automated at that cadence.
What holds you here
Sustaining bounded authority across launches, changeovers and staff turnover is an organisational discipline rather than a build, and its decay is invisible until a KPI moves.
Highest-leverage next move
Wire policy revalidation into the launch gateway, and treat the reward specification and the envelope as versioned artefacts reviewed exactly like a control plan.
Cost of leaving
- Effort
- Continuous
- Team
- A small platform team plus a standing forum: process engineering, quality, controls, maintenance, safety
- Risk
- Concentrated and regulatory — low frequency, high consequence, with documentation duties attached
If this is you, the next step is
We reconstruct one automated decision from your records — state, version, envelope, approver — and hand you the gap list.
Where automotive plants actually sit on the ladder
The distribution across the five stages, why the public record stops at Shadowed, and what the one well-documented bounded deployment — outside automotive — actually shows.
Most automotive plants are at Rule-bound or Simulated, and the population thins sharply above Shadowed. That distribution has a structural cause rather than a cultural one: the decisions RL suits are already served by rules that took decades to tune, the plants that could evaluate a replacement mostly do not log the decision, and the step that converts evidence into authority runs through a change-control regime that was designed — correctly — to make process changes slow. Nothing about that is a failure of ambition. It is what a certified, safety-critical, capital-intensive production system looks like when a new class of decision-maker asks for write access.
Distribution of automotive plants across the five control-authority stages
Illustrative distribution, not a survey: it is a reading of the published record — how little plant-floor reinforcement learning with write authority is documented anywhere, against how much simulation, virtual commissioning and robot-learning research OEMs publish. Treat the shape as the argument and the numbers as illustrative.
Share of plants
- 34% — 1 · Rule-bound (the decision is not even logged)
- 31% — 2 · Simulated
- 22% — 3 · Shadowed
- 10% — 4 · Bounded authority
- 3% — 5 · Governed autonomy
Source: Illustrative; anchored to the published record of bounded industrial RL deployments
It is worth being blunt about the public record, because vendor material is not. No automotive manufacturer has published a plant-floor reinforcement-learning policy holding unattended write authority over a production decision. What OEMs do publish is the two stages below that, and they publish a lot of it: BMW Group (opens in a new tab) reports extensive virtual planning and simulation of production systems before physical build, Toyota (opens in a new tab) and Toyota Research Institute (opens in a new tab) publish robot-learning research conducted in laboratories rather than on lines, and General Motors (opens in a new tab) reports machine learning applied across manufacturing operations in overwhelmingly predictive and perceptual roles. That is a Simulated-and-Shadowed frontier, honestly reported, and reading it as evidence of learned control on production lines is the single most common misreading in this topic.
The scale term is what makes the small percentages worth chasing at all. ACEA (opens in a new tab) represents an industry producing vehicles in the tens of millions of units a year across its members' plants, and a paint shop makes a sequencing decision every time a body leaves the bank — a decision taken thousands of times a shift, every shift, for the life of the line. At that repetition, a small average improvement is a large annual number and a small average degradation is an equally large one, which is the real argument for evidence discipline. It is also why the sensible first RL decision in a plant is one that repeats constantly, is cheap to reverse, and has no safety function anywhere near it.
Which plant decisions are actually reinforcement-learning shaped
The candidacy screen: five conditions a decision must satisfy, the decisions in an automotive plant that pass and fail them, and the quadrant that tells you which controller belongs on the decision at all.
A plant decision is reinforcement-learning shaped when five conditions hold together, and failing any one of them makes RL the wrong tool rather than a harder version of the right one. The decision must be sequential, in that the consequence of choosing now depends on and affects what can be chosen later. It must repeat often enough to accumulate evidence — hundreds of times a shift, not weekly. Its outcome must already be measured by a system the plant trusts. Its action space must be bounded and enumerable, so that an envelope can be written around it. And the alternatives available at each decision point must be knowable, because without them there is nothing to compare against. Most plant AI proposals fail at least two of these, and the screen below is how we sort a wish list in an afternoon.
Sequential, with a delayed consequence
Pulling a silver body out of the bank now is cheap; it becomes expensive twenty bodies later when the block ends early and a colour change lands in the middle of a run. If the consequence of an action is fully visible within the same cycle, the problem is a prediction or an optimisation, and both are cheaper to build, validate and explain than a policy.
Repeated at high frequency
Evidence accrues per decision, not per week. A decision taken three thousand times a shift generates a usable off-policy estimate inside a month; one taken twice a shift will not generate one inside a year, and a policy that cannot be evaluated should not be deployed. Decision frequency is the first number to establish and the one most business cases omit.
Outcome already measured by a trusted system
Colour changes and purge volume are in the paint MES. Starvation and blocking minutes are in the conveyor system. Kilowatt-hours are on the meter. If the outcome would have to be newly instrumented, that instrumentation is the project, and it is worth doing on its own merits before anyone trains anything.
Bounded, enumerable action space
Choosing among the bodies physically present in a bank is bounded by construction. Choosing an arbitrary continuous setpoint is bounded only if someone writes the range, the rate limit and the suppression conditions down. An action space that cannot be written down cannot be filtered, and an action that cannot be filtered will not — and should not — be approved for write access.
Known alternatives at each decision point
Off-policy evaluation compares against what was available, not against what is imaginable. A log that records the chosen body but not the other eleven in the bank supports monitoring and supports nothing else. This is the condition plants most often discover they fail, and it is fixed with a schema change rather than with a platform.
| Plant decision | Is it RL-shaped? | The incumbent you must beat | System of record | KPI it moves | Verdict |
|---|---|---|---|---|---|
| Paint-shop colour block sequencing | Yes — sequential, thousands of decisions a shift, consequences compound down the block | Rule-based block builder in the paint MES | Paint MES / selectivity bank controller | Colour changes per shift, purge volume, block length, due-date deviation | Best first RL decision in most plants — high frequency, cheap to reverse, no safety function in the path |
| Buffer and bank release between shops | Yes — releasing now starves or floods a downstream shop twenty minutes later | Fixed FIFO plus shift-leader overrides | MES / conveyor control | Starvation and blocking minutes, OEE availability | Strong stage 3–4 candidate once the manual overrides are logged as decisions |
| AMR and tugger dispatch for lineside supply | Yes — sequential assignment under uncertainty, outcome measurable per call | Fleet manager rule set: nearest vehicle, priority queue | Fleet management system / MES | Lineside stockouts, empty travel, fleet utilisation | Stage 3–4; the fleet manager's own safety layer already supplies part of the envelope |
| Paint booth and compressed-air setpoints | Yes — continuous control with delayed thermal and quality consequence | Schedules and PID loops; MPC where it has been built | Building management system / utilities SCADA | Kilowatt-hours per unit, booth condition excursions, finish defects | Stage 4 candidate, and only inside the validated process window |
| Robot motion and path refinement | Partly — RL-shaped in simulation, but the production motion programme is a validated artefact | Offline-programmed, validated robot paths | Robot controller / offline programming system | Station cycle time, spatter and rework | Simulate and then deploy as a re-validated programme — never as a live policy on the cell |
| Torque and process-parameter tuning at a station | No — the consequence is immediate and the mapping is learnable from labels | Control-plan windows set by process engineering | MES / torque controller | NOK rate, rework | Use supervised models and designed experiments; exploration on a joint is not acceptable |
| Defect detection at inspection | No — one-shot classification, no sequence, no action consequence | Vision system with tuned thresholds | Vision system / QMS | False-reject rate, first-pass yield | Supervised learning. Reinforcement learning adds nothing but risk and vocabulary |
The two rejected rows deserve as much attention as the accepted ones, because they are where credibility is won or lost with a plant audience. Torque tuning fails the screen on its first condition — the joint either passes or fails now, so there is nothing sequential to learn — and it fails on a second, harder ground: exploration on a safety-relevant joint is not an acceptable activity, whatever the expected value. Inspection fails even more clearly. Both are properly served by supervised models and by the analytical disciplines covered elsewhere in this series, and a programme that proposes reinforcement learning for either is signalling that it has chosen a technique before looking at the problem. The right conversation with process engineering is always the screen first and the technique second.
Which controller belongs on this decision
Plot how well you can write down the plant's dynamics against how long the consequence of an action takes to arrive. Reinforcement learning owns exactly one quadrant, and three of the four answers are cheaper, faster and easier to validate than a policy.
Reinforcement learning territory
- Delayed consequence, no written dynamics
- Paint sequencing, buffer release, fleet dispatch
- Learn from the decision log; evaluate off-policy; shadow before authority
Model predictive control
- Delayed consequence, dynamics you can state
- Booth conditioning, oven profiles, energy scheduling
- Cheaper to validate and far easier to explain to a safety engineer
Supervised learning
- Immediate consequence, no written dynamics
- Inspection, predictive quality, parameter recommendation
- Predict the outcome and set the parameter; run a designed experiment, not exploration
Classical control and rules
- Immediate consequence, dynamics you can state
- Interlocks, PID loops, control-plan windows
- Reinforcement learning here is expensive theatre with a validation bill
Notice what the top-right quadrant implies for sequencing conversations. If your process engineers can write down the dynamics and the constraints — and for booth conditioning, oven profiles and much of the utilities estate they often can — then model predictive control is very likely the better engineering answer, and saying so early buys enormous credibility for the cases where a policy genuinely is the right tool. Reinforcement learning's honest domain is the top-left: decisions where the plant's behaviour is a tangle of upstream release rules, breakdowns, product mix and human overrides that nobody has ever succeeded in writing down, but which the plant repeats often enough to learn from. That domain is real, it is valuable, and it is much smaller than the market for it.
Reward design: where automotive reinforcement learning actually fails
Not in the algorithm. In the objective — and specifically in the terms nobody wrote down, which is almost always a budget boundary rather than a mathematical oversight.
Automotive reinforcement-learning programmes fail at the reward far more often than at the algorithm, and they fail in a recognisable pattern: the reward contains the KPI of the department that sponsored the project, and the costs the policy discovers it can spend belong to departments that were not in the room. A sequencing policy rewarded on colour changes will happily hold a rare colour in the bank for two days, because due-date deviation is logistics' number. A release policy rewarded on OEE availability will run conveyors and drives at the top of their duty range, because component life is maintenance's number. Neither behaviour is a bug in the model. Both are the model doing exactly what was written down, with a persistence no human dispatcher would sustain.
Throughput bought with quality
A reward on units per hour, with rework recorded in the quality system rather than the MES, is an invitation the policy will accept. It learns to push the line into conditions where the immediate count rises and the defect appears at end-of-line or, worse, in the field. The fix is not a smarter model; it is a first-time-through term with a real weight, read from the system that actually records rework.
Throughput bought with maintenance life
Duty cycle, start–stop counts and thermal excursions are all things a policy can spend for immediate gain, and all things whose cost arrives months later in another budget. Include a duty term, or cap the behaviour in the envelope, or accept that the programme's second year will be spent arguing about a maintenance overspend that nobody can attribute.
The constraint satisfied by starving the constraint
Told to minimise colour changes, a sequencer can achieve superb block lengths by simply never releasing bodies whose colour is scarce. Block length looks excellent; due-date deviation and bank residency quietly deteriorate. Any objective phrased as 'minimise X' needs a companion term on the resource X is being minimised out of, and a hard bank-residency limit in the envelope rather than a penalty in the reward.
A reward read from a signal a person can move
If any term in the reward is influenced by a manual confirmation, a bypass switch or an operator-entered code, the policy will eventually learn to influence the operator rather than the process — by proposing actions that make the confirmation easier to give. Reward terms should read from process telemetry and system-of-record outcomes, never from a human's discretionary input.
Episode boundaries that hide the cost
Ending an episode at the end of a shift teaches a policy that the state it hands over does not matter, so it learns to end shifts with a bank full of awkward bodies and an unbalanced sequence. Episodes must either span the handover or carry an explicit terminal-state value that prices the mess being left behind. This is the most common technical error on the list and the easiest to fix once seen.
| Term | Sign | Source system | Read cadence | What it prevents |
|---|---|---|---|---|
| Colour change count | Negative | Paint MES | Per body | Nothing — this is the objective the project exists for |
| Purge volume | Negative | Paint MES / solvent metering | Per changeover | Counting changeovers as equal when a light-to-dark change costs several times a dark-to-dark one |
| Due-date deviation, in hours | Negative | Order system / MES build plan | Per body released | Perfect blocks built by indefinitely deferring scarce colours |
| First-time-through at end-of-line | Negative on shortfall | Quality system | Per shift, attributed per body | Throughput bought with rework that lands in another department's ledger |
| Bank residency over the stated limit | Hard constraint, not a term | Selectivity bank controller | Continuous | A body starved indefinitely because the reward found it convenient |
| Oven and flash-off dwell window | Hard constraint, not a term | Paint process control | Continuous | Any trade between finish quality and sequence efficiency — this one is physics |
| Terminal state value at shift handover | Negative on imbalance | Bank state snapshot | Per shift boundary | A policy that optimises its own shift by wrecking the next one |
How to write a reward you can defend
List the ledgers before the terms
Ask which departments hold a KPI the decision could move: process engineering, quality, maintenance, logistics, energy, and the downstream shop. Each of them gets either a term with a sign and a weight, or an explicit written statement that the decision cannot affect them. There is no third option, and the exercise routinely finds a ledger nobody had considered.
Separate constraints from preferences, and put constraints in the envelope
Anything that must never happen — an oven dwell violation, a bank residency breach, a setpoint outside the validated window — is a hard constraint enforced by the action filter, not a penalty in the reward. Penalties are prices, and a policy that finds enough upside will pay them. This single distinction prevents most of the failures on the list above.
Name the gaming modes in writing, before training
For each term, write one sentence describing how a determined optimiser could satisfy it while making the plant worse, and what would detect that. This becomes the monitoring specification for shadow and for production, and it is the section of the document that process engineers engage with most, because they have watched people do exactly these things for years.
Version it, sign it, and review it at every launch
The reward specification is a controlled document with a version, an owner and an approval, in the same sense as a control plan. A new product, a new colour, a rebalanced line or a changed shift pattern can invalidate a weight that was correct last quarter — so it is reviewed at the launch gateway alongside process capability, not when someone notices a KPI drifting.
What the public record actually shows
Three publicly reported programmes, read against the ladder — and read honestly, because none of them is a plant-floor policy with write authority, and the operators do not claim it is.
The published automotive record on learned control stops at the ladder's second and third stages, and the operators themselves are precise about that even where the coverage is not. What the three programmes below demonstrate is genuinely valuable and genuinely bounded: an environment built at industrial fidelity, robot policies learned in a laboratory, and machine learning deployed at plant scale in perceptual and predictive roles. Read together they describe the frontier accurately — the assets that make bounded authority possible are being built, and the authority itself has not been publicly granted anywhere.
Three programmes read against the control-authority ladder
Outcomes as reported by the operators' own published material — verify against the linked source before reusing anything; we have not independently audited them. Images are generated library scenes, not operator photography, and no operator endorsement is implied.
BMW GroupGlobal OEM · premium vehicles · 30+ production sites12
- Challenge
- Planning and continuously changing a global production network in which layout changes, rebalances and launches traditionally had to be validated physically — late, expensive, and one plant at a time.
- Approach
- BMW has publicly reported building detailed virtual representations of plants and production systems and validating processes in simulation before physical commissioning, together with a portfolio of AI applications running on digitalised production data. In control-authority terms, that is investment in the environment: a model of the line at a fidelity where questions can be asked of it.
- Reported outcome
- BMW reports planning and validating production virtually ahead of physical build and extending simulation-first planning across its network, alongside in-plant AI applications built on the same production data.
- What it shows about the curveAn industrial-fidelity environment is the precondition for stage 2, and it is a reusable asset that outlives any policy trained in it. It is not, on its own, evidence about a policy — which is why the ladder puts environment and evidence in different stages, and why replay validation sits between them.
Toyota / Toyota Research InstituteGlobal OEM · research institute · robot behaviour learning12
- Challenge
- Teaching robots dexterous, contact-rich behaviours that cannot be hand-programmed — the class of task where writing the dynamics down is precisely what nobody can do.
- Approach
- Toyota Research Institute publishes robot-learning research on teaching manipulation behaviours from demonstration and experience, conducted in laboratory settings. Separately, Toyota's own account of its production system rests on jidoka — equipment that stops itself when an abnormality occurs, so that a human decides what happens next.
- Reported outcome
- TRI reports learned robot behaviours in research settings; Toyota's published production philosophy continues to place a stopping authority in the equipment and the operator rather than in any optimiser.
- What it shows about the curveThis is the separation the whole page turns on. A leading robot-learning programme and a production system whose first principle is that a machine must stop itself are not in tension — jidoka is the same idea as the action envelope, expressed sixty years earlier. Research capability is not plant authority, and the operator with the strongest learning programme is also the one most explicit about where the stop lives.
General MotorsGlobal OEM · high-volume assembly · North American and global plants12
- Challenge
- Unplanned downtime and quality escapes across a large, heavily automated plant estate, where the cost of a stoppage is measured in vehicles per minute rather than in engineering hours.
- Approach
- GM publishes ongoing reporting on applying machine learning across its manufacturing operations, concentrated — as almost all published plant-floor AI is — on perception and prediction: condition monitoring of robots and equipment, inspection, and analytics that tell a maintenance or quality team what to do rather than doing it.
- Reported outcome
- GM reports machine learning used across manufacturing operations in monitoring, inspection and predictive roles, with the resulting actions taken by maintenance and production teams.
- What it shows about the curveThe incumbent decision-makers on a car line are good, and they are supported by exactly this kind of analytics. Any reinforcement-learning proposal is competing with a tuned rule advised by a working predictive stack, not with a vacuum — which is why the divergence rate in a shadow run is so often near zero, and why that is a legitimate result rather than a disappointment.