
Why Governance Must Come Before AI in Leadership Decisions
14 min read
The wrong order
A CHRO brought a leadership assessment platform to the talent committee, confident in what it could offer. The vendor's dashboard was polished, the competency taxonomy precise, and the security and data-residency questions were answered without hesitation. The General Counsel asked one question. What had the board seen that would allow it to govern the system if a promotion decision made on its outputs were later challenged? The room did not have an answer. That question should have come first.
I have watched the order go wrong more than once. The board sees the demo before it sees the controls. Procurement sees the vendor pack before NomCo sees the governance burden. HR sees more data and assumes the evidential position has improved. It has not. More data does not remove the governance problem; it sharpens it.
This is the sequence error underneath a great deal of current AI deployment in leadership work. Boards debate whether the output looks useful, then ask legal and governance teams to clean up the implications afterwards. By the time a board is debating whether to trust the output, the governance work should already have been done.
In “Where Leadership Analytics Goes Wrong and How to Fix It”, I introduced the Claim Boundary: a published statement of what a leadership system does claim, does not claim, and cannot yet claim on the evidence available. That is the statement of limit. This article asks a stricter question one step earlier. What has to exist before the board is willing to let that claim into the room at all?
The distinction matters. A Claim Boundary tells the board where the evidence stops. A Regression Pack tells the board whether the system is governed well enough for the board to rely on anything inside that boundary. D is about honesty; E is about permission. They are not the same discipline.
This is one of seven disciplines I am publishing on evidence-based leadership decisions. Each examines a different point at which the evidence and the confidence diverge.
What the governance literature already established
The literature is clearer on this than current deployment behaviour suggests.
Bogen and Rieke showed that employers and vendors can adopt automated hiring tools without giving buyers or affected people a sufficiently clear view of what is being measured, how judgments are formed, and where the risks sit (Bogen & Rieke, 2018). That disclosure gap is not accidental. It is structural. Providers invest heavily in documenting what a system can do but rarely invest with equal precision in documenting where it should not be relied on. Raghavan and colleagues deepened the point: hiring-system promises about fairness and objectivity often run further than the practices beneath them can support (Raghavan et al., 2020).
The accountability literature moves in the same direction. Mitchell proposed model cards because high-impact systems need a standard way to state intended use, limits, and relevant caveats (Mitchell et al., 2019). Raji argued for internal algorithmic auditing because accountability gaps do not close themselves (Raji et al., 2020). NIST's AI Risk Management Framework begins with mapping and governing risk, not with a dashboard or a vendor promise (NIST, 2023).
Other high-stakes domains already know the sequence problem is real. SR 11-7 and PRA SS1/23 do not ask financial institutions to use a model first and document it later, they assume the reverse. Model use is a governance issue before it is a performance issue. COSO's internal control framework makes a related point in plainer language: a control environment is not something you describe after the process runs. It is what makes the process trustworthy while it runs (COSO, 2013). Reporting disciplines such as TRIPOD and TRIPOD+AI apply the same logic to prediction models (Collins et al., 2015; Collins et al., 2024). The same logic extends to model design itself. In settings where the stakes are material, explanation after the fact is a weaker control than systems that remain inspectable while the decision is being made (Rudin, 2019).
Those literatures are saying something simple. Governance is not the paragraph at the back of the deck; it is the precondition of responsible use.
Why documentation after deployment is not governance
Boards often receive the wrong pack. They receive a procurement summary, an information security response, perhaps a privacy note, perhaps later a model card, and perhaps later still an assurance deck once the pilot is already underway. All of that may be helpful, but none of it resolves the sequencing problem.
Documentation after deployment explains what was built, while governance before deployment limits what the system is allowed to do inside a decision. Those are different objects with different purposes. The distinction matters because leadership decisions have social lock-in. Once a system's output has entered a succession discussion, helped frame a shortlist, or shaped who counts as credible, the board is no longer judging in a clean room. It is judging under the influence of an output already being treated as meaningful. Governance at that point is being asked to ratify a use that is socially underway.
A board does not need thicker description after the fact. It needs a smaller, sharper pre-deployment record that constrains behaviour. What is the system for here, exactly? What is it not for? What evidence classes support the output? When must the system abstain? How will the board know if the output should stop being relied on?
The regulatory direction of travel is the same, and the AI Act is phasing in. Transparency obligations have applied since 2 August 2026, the UK's Data (Use and Access) Act is in force, and the Annex III high-risk obligations that cover employment uses are set to apply in December 2027.
Boards should not wait for a compliance date; the burden of proof starts before reliance.
Regression Pack
This is the missing object. I use the phrase deliberately. Boards do not need a glossy AI policy once the system is already running. They need a pre-deployment pack that shows how the system behaves under challenge, how its claim narrows when evidence weakens, and what happens if live use falls below the standard the board was promised.
It is not a model card clone. Model cards help. Internal audits help. Security reviews help. But boards need one governance object they can use in procurement, approval, and oversight without becoming technical specialists. The first page of the pack should be the Claim Boundary, introduced in my previous article “Where Leadership Analytics Goes Wrong and How To Fix It”. The remainder governs first reliance, ongoing challenge, and conditions for withdrawal.
Seven elements turn it into practice:
Intended use
The exact decision context must be named. Search-universe expansion is different from shortlist pressure-testing. Shortlist pressure-testing is different from committee support in succession. A board should not have to infer the use from marketing language.
Excluded use
The pack should state, in plain English, what the system must not be used for. Not a sole basis for an employment decision. Not a covert inference about personal qualities. Not a replacement for committee judgment.
Evidence classes
The board should see what classes of evidence are in play. Public record. Structured inputs. Documented track record. Human review. Role-relevant outcomes. Without that, the board cannot judge whether the claim is proportionate.
Abstention conditions
Low coverage, weak construct match, role ambiguity, sparse evidence, conflicting signals. These are not edge cases. They are normal conditions in leadership work. A system that never abstains is not complete. It is undisciplined.
Error history
What has the system been wrong about? What weak spots are known? What changed after challenge or correction? Boards do not need a perfection story. They need an error culture.
Challenge tests
What would materially alter the conclusion? Changed role requirements? Source corrections? Stronger contradictory evidence? If the provider cannot say what would change the output, the board is dealing with a conclusion that has not been properly stress-tested.
Rollback conditions
At what point does the board stop relying on the output? This is the most neglected element. A system should not enter a leadership decision unless the board also knows what would cause it to suspend use, narrow use, or revert to human-only judgment.
What boards should require before AI touches real decisions
A Regression Pack does not rescue a weak system. It does not validate a vague construct. A strong pack cannot make a weak claim stronger than the evidence allows. Governance is necessary. It is not sufficient.
I need to be precise about regulation. This is not legal advice. UK data protection and recruitment governance, alongside the AI Act, all move in the same direction. Documentation, due diligence, meaningful oversight, and a workable record of challenge are not optional clean-up work once the tool is live, they are design decisions.
This argument would weaken on two conditions: that governance retrofitted after deployment creates controls as robust, inspectable, and board-usable as systems designed with governance from the start, and that documentation produced after the fact constrains real-world use as tightly as a pack approved before first reliance.
Before the next AI-enabled tool enters a live succession, appointment, or leadership review, the board should require the pack: ahead of the demo, the press-ready summary, and the model explainer that arrives after the pilot has already run.
The board should be able to see, before first use, what the system is for, what it is not for, what evidence classes support it, where it abstains, what it has got wrong, how it will be challenged, and what would cause the board to stop relying on it. That is the minimum burden of proof for letting AI near a consequential leadership decision.
And once that discipline is in place, the next question becomes harder again. Not whether the system is governed, but whether the leaders using it can themselves govern AI in role. That question cannot be answered by self-report; it has to be answered by record.
Governance is not what comes after trust breaks, it is what makes reliance thinkable in the first place.
Written by James Nash.
First published on inBeta.io. Co-published on Substack. Summer 2026.
Series: The Seven®, by James Nash. © Copyright 2026 inBeta. inBeta, Optics, Divergence and The Seven are all trademarks of inBeta Ltd

James Nash
James is the founder of inBeta. He has spent fifteen years working with boards and senior leadership teams at global and publicly listed companies on succession, talent, capability, and leadership governance. He holds executive education from Saïd Business School, University of Oxford, in Artificial Intelligence (including Audit and Ethics), Executive Leadership, Strategic Innovation, and Executive Finance. He founded inBeta because he kept watching boards make their most important decisions on instinct, narrative, and incomplete information, and believed the evidence base existed to do it differently. James is a certified AI Auditor, AI Ethicist, and AI Professional (CAIA, CAIE, CAIP; Oxethica), and a certified practitioner in CliftonStrengths (Gallup), Hogan (including PBC 360), FIRO-B, and Cultural Intelligence (CQC).
METHODS APPENDIX
This article forms part of my thinking on evidence-based leadership decisions, a series of pieces I am surfacing through 2026, arguing for a governance standard for consequential people decisions rather than a single technical method. The appendix discloses the principles behind that standard at a level appropriate for board review. It does not disclose scoring formulae, thresholds, or controlled parameters. I have built a system in this market, and the standard set out here applies to my own work before it applies to anyone else's. AI tools from Anthropic and SpaceXAI were used in preparing this series, under my direction and review. The arguments, the practitioner observations, and the judgments are mine, and I take full responsibility for the final text. No AI system is an author of this work.
Construct
Regression Pack. A board-facing pre-deployment governance pack that states intended use, excluded use, evidence classes, abstention conditions, error history, challenge tests, and rollback conditions before an AI system enters a leadership decision. A governance discipline, not a model card clone.
My intended use
To help boards, governance committees, CHROs, Company Secretaries, and General Counsel require the pre-deployment controls needed before AI enters a consequential leadership decision.
My excluded uses
My writing and thought leadership are my own and do not evaluate any specific vendor or product. This article does not prescribe a technical method. It does not claim that governance alone makes a weak system reliable. It does not provide legal advice on AI Act compliance, employment law, or data protection.
Abstention conditions
The standard I have written about applies to AI-enabled systems used in consequential leadership decisions involving identifiable individuals and material board or committee judgment. It may not apply in the same form to low-stakes internal productivity tools, purely descriptive analytics, or non-individual aggregate reporting.
Source classes
Three classes of evidence. First, peer-reviewed and official work on hiring-system disclosure, model interpretability, model reporting, and AI accountability: Bogen and Rieke (2018), Raghavan et al. (2020), Rudin (2019), Mitchell et al. (2019), and Raji et al. (2020). Second, governance and control frameworks from AI, model-risk, and internal-control domains: NIST AI RMF 1.0 (2023), SR 11-7 (2011), PRA SS1/23 (2023), COSO (2013), TRIPOD (2015), and TRIPOD+AI (2024). Third, regulatory context and practitioner observation from my own work in board and procurement settings, where I have repeatedly seen deployment pressure run ahead of the governance the board would later need to defend.
Bibliography
Bogen, M., & Rieke, A. (2018). Help Wanted: An Examination of Hiring Algorithms, Equity, and Bias. Upturn. https://www.upturn.org/work/help-wanted/
Collins, G. S., Reitsma, J. B., Altman, D. G., & Moons, K. G. M. (2015). Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD): The TRIPOD Statement. Annals of Internal Medicine, 162(1), 55–63. https://doi.org/10.7326/M14-0697
Collins, G. S., et al. (2024). TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods. BMJ, 385, e078378. https://doi.org/10.1136/bmj-2023-078378
COSO. (2013). Internal Control: Integrated Framework. Committee of Sponsoring Organizations of the Treadway Commission. https://www.coso.org/guidance-on-ic
European Commission. (2026). Navigating the AI Act. https://digital-strategy.ec.europa.eu/en/faqs/navigating-ai-act
European Union. (2024). Regulation (EU) 2024/1689, Artificial Intelligence Act. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
Federal Reserve. (2011). SR 11-7: Supervisory Guidance on Model Risk Management. https://www.federalreserve.gov/supervisionreg/srletters/sr1107.htm
Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Hutchinson, B., Spitzer, E., Raji, I. D., Vasserman, L., & Gebru, T. (2019). Model Cards for Model Reporting. FAT* '19. https://doi.org/10.1145/3287560.3287596
NIST. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). https://doi.org/10.6028/NIST.AI.100-1
Prudential Regulation Authority. (2023). SS1/23: Model Risk Management Principles for Banks. https://www.bankofengland.co.uk/prudential-regulation/publication/2023/may/model-risk-management-principles-for-banks-ss
Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices. FAT* '20. https://doi.org/10.1145/3351095.3372828
Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., Hutchinson, B., Smith-Loud, J., Theron, D., & Barnes, P. (2020). Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing. FAT* '20. https://doi.org/10.1145/3351095.3372873
Rudin, C. (2019). Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nature Machine Intelligence, 1, 206–215. https://doi.org/10.1038/s42256-019-0048-x
UK Public General Acts. (2025). Data (Use and Access) Act 2025. https://www.legislation.gov.uk/ukpga/2025/18/contents
More articles
Our events






