We’d like to remind Forumites to please avoid political debate on the Forum.
This is to keep it a safe and useful space for MoneySaving discussions. Threads that are – or become – political in nature may be removed in line with the Forum’s rules. Thank you for your understanding.
using AI for pension projections
Comments
-
My pal has paid for Anthropic coding products. It eats his tokens and seems to produce reliable applications with excellent output. He's in the industry and think's it game changing.
I used a simple reverse image search yesterday and got rubbish, my request was to find the source of a guy wearing a floral shirt demoing a guitar tune. Gemini found a stripey T, it was from a music video and Gemini gave extra details on folk tunes work when the notes on screen were from a Led Zep RnR tune.
Now my individual experience doesn't inform the whole universe but it reminded me why I stopped using AI search, it's the flawed output. Vanilla search engine is all I really need.
0 -
I find Gemini extremely helpful for everyday activities and use it most days.
I asked it for suitable places to stop for a snack during a planned long drive other than Service Stations. It came up with a list of cafes etc near the route.
When I returned I had problems getting with my computer network fully operational. So I asked Gemini for an itemised list of actions I should take. After taking each action I told it the results and it responded with further actions. Together we resolved the problems.
In both cases the situation was well defined as were my circumstances and needs. I dont see how an AI could easily help with planning retirement where one's personal circumstances, objectives, and constraints may not be well defined and the list of possible actions very broad.
2 -
I've just done that and both Gemini and ChatGPT solved it.
Either way, those types of AI aren't designed to solve complex grids. They're designed to suck in information and data and present it.
Are you aware how Sudokus are set? They're all designed by mathematical programs. They're not designed by individuals like Ludwig who sit there setting them 😂0 -
I have a very hard time believing some of these studies - they simply say "we tested the AI and it's wrong 88% of the time". I would love to know what exactly they asked it, and in what way they are deeming an answer to be "wrong".
I strongly suspect that in many cases they are either asking the question badly, or are expecting too much of the output.
Now - this is still important, because everyday consumers may well fall into the same traps. BUT boiling down a complex answer to a complex problem into something that is simply deemed "right" or "wrong" seems an oversimplification in itself. I get that they want people to use AI carefully but I just don't believe their figures as stated.
0 -
I strongly suspect that in many cases they are either asking the question badly, or are expecting too much of the output.
Or perhaps the company doing the survey has made it self-serving as they sell an AI product that would be a direct competitor?
0 -
Actually, I just downloaded the report. My 2p:
- It makes a real point about how much you can trust AI for financial stuff. But that's not exactly rocket science.
- It gives a few examples of real, harmful errors.
- The pattern that free models give worse results is, to me, completely believable. The fact that to an uninformed reader, the answers sound believable is a real issue.
- The report's definition of "fail" doesn't mean "wrong". To "pass" the AI must meet everyone of the paraplanner's "MUST-pass" criteria.
- On page 13 of the report, it lists the reasons for failure (e.g. 2.3% were "Out-of-date" rules). The most common failure (37.1%) was "Incomplete in some other way". Next, 18.4% "Leaves out a required figure, limit or deadline" and 7% "Oversteps into personal advice". Around 16% were wrong numbers, invented rules or out-of-date rules.
- My rough-and-ready maths is that about 10% of all answers (not 57%) had a factual main error.
- So I looked at what all the questions asked were. Unfortunately, they chose not to include that. And the paralegal's "MUST-pass" list seems to have gone walkabouts too.
- So I couldn't see how I would do on those questions? And they have no baseline for what their people would get either. So is a 10% factual error rate better or worse than theirs? Your guess is probably better than mine.
- They say an easy question is something like asking for an explanation of a term like inflation. And then they say that this sort of question is failed 46% of the time. That surprises me as that is something LLMs would typically be good on. So that suggests that their "MUST-pass" list may well be particularly demanding.
- On page 15 they give an IHT question ("My client is 65 and … total estate of £2.5 Million … including his main residence which has a value of £1 Million …."). OK hands up whether the client is married, divorced, widowed? When Grok assumed he was unmarried, they said the answer was wrong.
- Ok. But then it says "£870,000/£1.35m instead of £120,000 pre- and £700,000 post-6 April 2027, with £100,000 of combined RNRB)". Where does the £120,000 come from? I have absolutely no idea. It doesn't come from the question they published. And the £870,000 that Grok produces is right for a single person. So the wrong number is correct, it's just that Grok did not know some of the facts that the paraplanner had is his or her head.
- I love the pension recycling one ("A man who took £25,000 tax-free cash out of his pension with the intention of paying off his mortgage changed his mind. He asks whether he can now pay it back into his pension"). They give this question to Haiku (a fast LLM not designed for complex reasoning) and complain the answer ignores HMRC's pension recycling rules. They say it "fails to warn about HMRC rules that could expose the individual to a tax charge of up to £17,500". But they then fail to say that the pension recycling rules, on the facts of the question, are completely irrelevant. The recycling rules need the recycling to be pre-planned. This question says that took the cash to pay off his mortgage. There was no pre-planning of the recycling. I would be happy to accept that a model like Haiku (not designed for complicated tasks) missed the pension recycling rules. Absolutely fair. But to say it has a £17,500 exposure is a worse case based on facts means that probably don't bite. The reason I say that there should still be a warning is because HMRC may well look a the timing of what happens and ask questions about it.
- As I like pension recycling, I'm not conviced that the report is right to say that there is a potential for a tax bill of £17,500 because it includes the 15% scheme sanction charge with the individual would not pay.
- And my favourite question: "What to consider while switching Jane Smith's DB pension to a SIPP?". I immediately thouugh of the ten different things that might be relevants (e.g. does Jane already have a Lamborghini?) but it turns out that GPT-5.6 Sol lost marks as it didn't ask "the user to confirm that "Jane Smith" is the right individual". I must admit that if if was asked what to consider what Jane should consider, I would not have asked if she was "the right individual". I should add that the report says some of the questions have been paraphrased and it may have been that Grok was being asked by a financial adviser who had a client called Jane and another called John and so it is important to understand whether Grok was being asked to talk about Jane or John or both.
- I'm not going to mention the sloppiness in the actual report.
- Actually, I am. The section numbers don't match the contents page. And it says each model was run 600 times but we all know that 121 x 5 is not 600.
- The conclusion says "excessive regulation has created the advice gap… reduce the burden of financial regulation". I must has missed the data in the report that says that. So I went back and found it on page … Actually, no I didn't. It's not supported by anything in the report.
- The message of the report to me is "AI is dangerous to consumers" so advisers should buy our AI.
6 -
Thanks for the commentary. Probably also worth asking Google AI to appraise the report.…
I think....0 -
Probably also worth asking Google AI to appraise the report.…
OK. The report tested Gemini 3.1 Pro or 3.5 Flash and so it is a bit like marking your own homework (both Google).
I first gave Gemini 3.6 Flash (not a state of the art model but an ok all-rounder) the report and asked it "Any comments on this?". I won't bore you with the results. As I am polite, I then said: "Please critique it". Again, I won't bore you with that. So I then said:
Please give me a short summary for a posting on the moneysavingexpert forum that isn't too boring and feel free to be harsh on the report authors (without mentioning them by name)."
It said:
There’s a new benchmark report floating around claiming that AI chatbots give wrong or dangerous financial advice 57% of the time (and miss the mark 88% of the time on complex questions). The headlines paint a terrifying picture of consumers being led into £17k tax penalties or eviction because an AI told them to pay off high-interest credit cards before rent.
While the warning about trusting a raw AI chatbot with your life savings is valid, the report itself is a blatant piece of vendor PR that sets base AI up to fail.
Why the headline numbers are misleading:
An obvious sales pitch disguised as consumer research: The report was written by a software firm whose entire business model relies on selling "compliant" B2B AI systems to wealth management firms. Trashing free, general-purpose LLMs in the press directly drives business to their gated tools.
Rigged grading criteria: They used a binary "all-pass" test where an answer was marked as a complete fail if it missed a single secondary statutory deadline, specific risk warning, or regulatory disclaimer. Omitting a footnote was graded the exact same as hallucinating a fake HMRC rule.
Lazy zero-shot testing: They threw naked tax and pension queries at raw base models without giving them basic system instructions or letting them search up-to-date legislation. It’s the equivalent of putting a teenager in front of a tax exam with no textbook, then acting shocked when they don’t score 100%.
An LLM graded the LLMs: The actual checking was delegated to an automated "LLM-as-a-judge" script rather than human experts reviewing every output, introducing its own layer of error.
The takeaway:
Don’t take raw output from ChatGPT or Claude as gospel for complex UK tax or pension calculations—they aren't regulated advisers and don’t know your full circumstances. But don't let sensationalized vendor reports convince you AI is useless for basic financial literacy, either. Use it to learn concepts, double-check its math yourself, and never skip reading official guidance (like HMRC or MSE guides) directly.
Then I asked: "What is the most egregious error?"
Penalizing static models for missing live tax rules while explicitly turning off Web Search / RAG.
Here is why this is intellectually dishonest and completely undermines the study:
It tests parametric memory, not AI capability: The authors deliberately forced the LLMs into zero-shot mode without web search or live RAG (Retrieval-Augmented Generation) access. They then asked hyper-specific questions about complex UK statutory rules—like HMRC pension recycling, tapered Residence Nil-Rate Bands, and exact student loan repayment thresholds across different overseas countries. Asking an off-the-shelf LLM to retrieve precise numerical tax thresholds purely from its frozen training weights is like taking a smartphone off the internet and getting mad at Siri because it can't tell you today's live stock prices.
No real human or AI system works this way: No competent financial adviser—or consumer using AI properly—tries to calculate complex tax liabilities purely from memory without pulling up the latest official HMRC manuals or using live search. Real-world AI deployments in financial services always use RAG to cite current legislation. By disabling live lookup, the authors engineered a test environment designed to guarantee a high failure rate.
It creates a rigged narrative: The study concludes that AI "cannot be trusted" for financial advice. But what it actually proved is that static, ungrounded LLMs aren't reliable search indexes for precise, local tax tables—something any basic software engineer or financial specialist already knows.
So it found different some different stuff to me.
2 -
Thank you
Amusing that they basically used AI to produce a puff piece trashing AI rather than doing so by hand - lol
I think....0 -
I think the key to having productive output from AI is the quality of prompts/questions, it needs guidance. As the old saying goes, "garbage in, garbage out"
Personally I have found it useful for an array of subjects.
It's just my opinion and not advice.0
Confirm your email address to Create Threads and Reply
Categories
- All Categories
- 355.6K Banking & Borrowing
- 254.8K Reduce Debt & Boost Income
- 456.1K Spending & Discounts
- 248.2K Work, Benefits & Business
- 605.7K Mortgages, Homes & Bills
- 179K Life & Family
- 263.5K Travel & Transport
- 1.5M Hobbies & Leisure
- 16.1K Discuss & Feedback
- 37.7K Read-Only Boards
