Keeping an eye on AI in financial services: A closer look at the autonomy spectrum and the changing role of the human
This website will offer limited functionality in this browser. We only support the recent versions of major browsers like Chrome, Firefox, Safari, and Edge.
In a recent post which gave an overview of The Mills Review, there was a promise of a series of deeper dives. This post looks at the question of managing autonomy, which The Mills Review positions as one of the most fundamental challenges for the financial services ecosystem right now.
Introduction
The AI autonomy spectrum is a framework through which humans move from “operators” who use AI to assist them with defined tasks to “observers” who monitor the outcomes of systems that act continuously within determined boundaries. It assists in analysis of the evolution of the human role across the spectrum together with the evolving risks and benefits.
The Mills Review suggests that the financial services ecosystem is likely to move rapidly along the spectrum driven forward by advances in AI capability, compute, data and digital infrastructure, and that it will reshape significantly in response.
An essay about agents
The source of the spectrum used by the FCA is an essay published last summer by the Knight First Amendment Institute at Colombia University “Levels of Autonomy for AI Agents”. The intention behind this academic piece (the Essay) is to “contribute meaningful, practical steps towards responsible deployed and useful AI agents in the real world”.
The Essay argues that while autonomy is a double-edged sword it could be possible to enable autonomous agents to be useful, reliable and safe, by intentionally giving them framework within which they have given capabilities, target deployment environments, refined user experiences, communication protocols, and resolution mechanisms.
The Essay defines five levels of autonomy, each characterised by the role that the user takes, and within each level analyses the ways in which the user can control and interact with the agent. This five-level framework is proposed as a way to reach the place where agents are actively in use and are safely and responsibly deployed.

Credit for figure 1. Levels of Autonomy for AI Agents | Knight First Amendment Institute.
The predictions about what agents will learn about us and do for us seem to become more realistic by the day as part of an overarching drive to automate tasks. However, they significantly amplify the positives v negatives picture and, as the Essay highlights, “make it more important and more difficult to anticipate harms”. For this reason, the Essay argues for autonomy to be by design i.e. intentional and not just consequential.
Autonomy by design
Autonomy by design is an autonomy that is user-centred and has its abilities to act without a user designed by reference to:
The user has five possible roles:
The role of the user is therefore intended to be an considered part of the framework and decision making, just as is the level of autonomy, which should be based on informed decisions, on the intended use, and on the desired outcomes.
Autonomy certificates
The Essay proposes an agent governance mechanism referred to “autonomy certificates”. These certificates are proposed as a tool for “risk assessment, safety framework design, and multi-agent systems engineering”. The Essay hypothesises that autonomy certificates would be issued by a governing body and live within an agent’s metadata where they would be visible to other agents and to developers. Certificates would prescribe maximum levels of autonomy and set technical specifications i.e. they would be a way of communicating the intention behind the design of an agent to other stakeholders in the agent ecosystem to assist, for example, with risk assessments, improving safety framework design, and predicting which agents will work with each other and how.
Definitions
The Essay helpfully attempts to remove room for confusion with some foundational definitions:
The concept of “involvement” is described as a multifaceted concept which “includes a spectrum of actions from direct control to light supervision”.
The spectrum
At the lowest level of autonomy, the user is always in charge of and owns the workflow. The agent is utilised only for on demand support, for example, summarising a report or searching for information on the web. At this level, the agent does not take any actions (even though it may be capable of doing so). An agent operating at this level of autonomy could be used in “high-stakes, high-expertise workflows where autonomous agent activities can be particularly costly if inaccurate, and/or where the lack of user involvement can easily lead to accountability concerns and legal consequences”.
The second level of autonomy envisages a close collaboration between user and agent where both can work on their own tasks, and both can “leverage each other’s capabilities and knowledge”. This level of autonomy requires clear and effective user-agent communication protocols and full user visibility into the agent’s work and users who are able to (re)take control of the agent’s work at any time including, for example, if the user noticed that the agent was being unproductive, doing something too risky, hallucinating, or if the user wanted to complete part of the task themselves (for example, if there was a value in it being completed by a human, even if the agent could do it).
At the third level of autonomy, more responsibility shifts to the agent. The user moves towards “feedback, preferences, and higher-level directional guidance rather than hands-on collaboration”. This may have consequences for the user including the possibility of losing the ability to take control of the agent or losing the ability to edit the agent’s outputs. At this level of autonomy, the agent will have some knowledge of the user’s expertise and preferences. The agent will therefore have a “learning curve” and “training period” to reach maximum utility.
At the fourth level of autonomy the user becomes more passive and will only interact with the agent when the agent “encounters a blocker it cannot resolve on its own” such as entering a passkey that it does not have or signing off on a consequential action. The Essay suggests that this level of agents is ideal for “tasks with high amounts of lower-stakes decision-making—the agent’s automated decision-making can improve workflow efficiency and relieve users of excessive cognitive load, while erroneous decisions do not impose significant risks”. Risks highlighted at this level of autonomy include:
At the highest level of autonomy where the user exists as an observer the agent is fully autonomous and has no need for user involvement. These agents “plan and execute tasks over long time horizons and make all decisions on their own. When they run into blockers, they repeatedly iterate on solutions until resolution or modify their approach to avoid running into the blocker in the first place”. At this level of autonomy, the user will be able to monitor the agent but will be unable to change its trajectory. The user’s only control mechanism at this level of autonomy is the “off-switch that shuts off all agent activity”.
The risks associated with fifth level agents are many and will likely prompt some of us to ask existential questions about why we would even want level five agents at all. Some of the risks include:
Why this matters
If, as per the prediction, more of financial services becomes assisted by or performed by agents, the role of the human in financial services will inevitably have to change.
Important decisions will need to be made about which elements of financial services are suitable for agents, at what levels of autonomy, and where humans are still needed.
As use cases move along the length of the spectrum, the questions that must be answered will become increasingly challenging questions, largely driven by oversight shifting from a focus on individual tasks and decisions, to a new system of setting boundaries and permissions around tasks completed and decisions made at scale and at machine-speed.
The risks and challenges are many and include what evidence will look like, how accuracy can be ensured, appropriate levels of reliance, the development of over-reliance and the loss of essential skills, including the ability to challenge robustly.
The role of the human in the financial services ecosystem will shift as human ability to exercise oversight changes and significant shifts in governance capability and risk management will need to happen. These changes will demand significant organisational shifts, workforces will need to re- and up-skill in AI competencies so that they are able to understand AI enabled financial services workflows, robustly and effectively scrutinise and challenge AI outputs, interpret AI behaviours, and compile and explain new forms of automated evidence.
The Mills Review points to an emergence of use cases across the autonomy spectrum and an appetite for automation. Significantly, it also highlights the prerequisites of improved “reliability, data access, verification and identity, liability and redress, consumer support, and consumer trust” as building blocks for an agentic infrastructure, including the development of regulator-led trusted agent protocols.
The FCA’s vision of the regulatory future is influenced by its acknowledgement that as AI moves along the autonomy spectrum, regulation as we know it will “come under strain”, and that supervision will need to go on its own agentic journey to ensure that it is capable of both “firm based supervisory efficiency, and for monitoring and detecting system-wide risks”.
The Mills Review demonstrates that the foundations of the regulatory infrastructure come under strain at different points of the autonomy spectrum:
Succinctly put by The Mills Review itself, the agentic vision for financial services obliges “the regulator itself… to move down the same autonomy spectrum as firms, consumers and the wider economy towards AI as Collaborator and AI as Approver”, and in an agentic financial services ecosystem the “AI agentic supervisory model would operate continuously across the FCA’s remit, monitoring, flagging, triaging and surfacing firm and system-wide signals”. This is, at present, a work in progress.
The conclusions made by The Mills Review about the impact of the journey along the autonomy spectrum, drive directly into its priority recommendations, including:
Conclusion
Arguably, firms and regulators alike, are currently sitting at different points on the autonomy spectrum. Significantly, without clarity on what the regulatory foundations are capable of holding and where exactly on the trajectory all relevant stakeholders are.
To conclude with some words from the Foreword of The Mills Review:
“Practically speaking, we don’t know how far this goes and full autonomy seems unlikely to sit well with the current regulatory framework or UK societal appetites. But we can prepare, and I believe an Agentic Supervisory Model giving the FCA the AI-enabled tools to monitor and detect emerging system-wide risks is a necessity. Striking the balance between enabling delegation and managing autonomy is the central challenge I believe the FCA is now committed to addressing. What emerges is set out in the executive summary, which includes our recommendations. Taken as a whole, I believe the Review offers a practical agenda for the responsible development of AI in retail financial services which enables our world-leading industry to continue to grow and innovate, and meet the needs of all consumers. Decisions on taking them forward rest with the FCA Board and Executive, and I wish them and the FCA every success in this endeavour. Go boldly where...”
Our thought leadership
You can read more updates like this by subscribing to our monthly financial services regulation update by clicking here, clicking here for our AI blog, and here for our AI newsletter.
If you would like to discuss how current or future regulations impact what you do with AI, please contact me, Tom Whittaker, or Martin Cook. You can meet our financial services experts here and our AI experts here.
“Our descriptions and examples of each autonomy level in our framework showcase autonomy as a deliberate design decision for AI agents. By posing open questions at each level, we show that powerful capabilities enable effective agents across all levels, not just the higher ones, as suggested by prior frameworks. Our framework further demonstrates that increasing agent autonomy involves making nuanced tradeoffs— more autonomy does not simply mean a better agent.”
https://knightcolumbia.org/content/levels-of-autonomy-for-ai-agents-1
Want more Burges Salmon content? Add us as a preferred source on Google to your favourites list for content and news you can trust.
Update your preferred sourcesBe sure to follow us on LinkedIn and stay up to date with all the latest from Burges Salmon.
Follow us