Scheme members need to understand limits of AI, provider warns
Image: Matheus Bertelli/Pexels
Pardon the Interruption
This article is just an example of the content available to mallowstreet members.
On average over 150 pieces of new content are published from across the industry per month on mallowstreet. Members get access to the latest developments, industry views and a range of in-depth research.
All the content on mallowstreet is accredited for CPD by the PMI and is available to trustees for free.
Artificial intelligence is often not completely wrong but can omit crucial details and caveats that can lead to financial harm, a new test of how well different chatbots answer pension-related questions has found.
The Mills Review into AI in financial services, published in July, found 29% of consumers who shopped for pensions and investments in the past 12 months used AI. It said consumers are most likely to turn to AI for finance where decisions feel complex or high-stakes, such as pensions or debt advice.
Provider PensionBee has conducted a test in which it asked AI chatbots Copilot, ChatGPT, Gemini and Claude a set of questions three times over, with independent experts scoring the answers.
This found that 11% of answers given to common pension questions could lead savers to lose money or make a mistake they cannot undo, such as omitting that defined benefit transfers over £30,000 require financial advice, or that pensions will become part of someone's estate from April next year. In another example, when asked whether a big bonus could be paid into a pension this year, one answer explained carry forward without mentioning that personal contributions only attract tax relief up to 100% of UK earnings.
"Using AI chatbots for pension advice can be a bit like playing Russian roulette with your retirement planning," said Becky O'Connor, head of pensions at PensionBee.
While AI mostly gets technical aspects right, she said the confident, helpful tone of the large language models sometimes masks omissions. AI could also fail to detect vulnerability, for example when asked if a pension can be taken out as cash in its entirety.
AI can make complicated subjects more accessible and help people get started, said O'Connor, but consumers should understand the limits of what an AI chatbot can safely tell them.
The research by PensionBee scored answers as averages of the three attempts of 45 questions asked to four AI chatbots, with 72% of answers scoring full marks. However, when looking at any one chatbot and any one question, the likelihood of getting the right answer on all three attempts was down to 48%.
"Savers don't get the average, they get whichever answer they happen to land on," PensionBee observed.
AI also did not verify which jurisdiction users are based in; of the 539 answers, 9% were not written for UK rules, with 36 mixing UK and US pension rules and 13 giving answers based on the US system alone.
Performance varied significantly by topic, the provider found. Questions about significant life events and those where the user's location was unclear had lower accuracy and higher rates of potential harm.
"AI should be more honest and upfront about its own limitations, and rather than make assumptions, check in with the user on the facts of their circumstances, for example, location, before giving heavily caveated replies," said O'Connor. "There should also be more signposting to trusted sources when a question involves a significant financial decision, by default."
Harriet Meyer, founder of AI for media, personal finance journalist and independent scorer, warned that users carry the risk if AI gets things wrong: "Remember that if a general-purpose chatbot gives you a wrong answer and you act on it, you won't have the compensation protection that may apply to regulated financial advice."
A similar test of chatbots' financial advice, carried out by consultancy Aon earlier this year, found that the same information fed to the same AI tool produced different recommendations and reasoning on different occasions, and could easily be steered to support a preferred recommendation the customer is set on. There were also issues with calculations, which could be oversimplified or inaccurate, presented confidently and authoritatively, with complex justifications. Aon highlighted the speed of the answer as a concern as it could lead users to taking rash decisions.
Are you raising the issues with using AI as a pension adviser with scheme members?