Evidence-based personal financeIndependent · no affiliates
guideaievidence

Should you let ChatGPT or Claude manage your money?

Half of American adults now ask an AI before a money decision. The models pass the CFA exam and then get a third of tax returns right. What the studies say, where it breaks, what to delegate and what never to.

ChatGPT as your banker? What the studies say about managing money with an LLM, and where it goes wrong.

In August 2025, Intuit Credit Karma asked a thousand Americans about money and AI. Two-thirds of those who used generative AI had asked it for financial advice, and 85% of them had acted on the answer. In the same survey, 52% admitted to making a poor financial decision based on something an AI told them.1 A year later, a Gallup poll for Edward Jones found that only 3% of Americans have “a great deal” of confidence in AI money advice.2

Both things are true at once. People are using these tools for money far faster than they trust them, and far faster than the evidence says they should. This article is about closing that gap: what large language modelsLarge language model (LLM)The type of AI behind ChatGPT, Claude and Gemini: a model trained on vast amounts of text to predict the next word, which makes it fluent but not inherently accurate.Read in the glossary → are demonstrably good at with money, where they demonstrably fail, what happens to a bank statement you paste into one, and a working rule for what to delegate. I read the surveys, the benchmarkBenchmarkA fixed set of test questions or tasks with known answers, used to score AI models against each other and over time.Read in the glossary → papers, the consumer-group tests and the vendors’ own privacy pages. Everything was checked on 21 September 2026.

Sparring partner, not oracleVerdict

Use ChatGPT, Claude or Gemini to understand, to draft, to categorise and to argue with. Do not use them to compute, to file, or to decide anything that depends on a current threshold, a date, or your jurisdiction unless you verify the answer against an official source. The models now pass professional exams at 95%+ and still complete less than a third of US tax returns correctly. Knowledge is no longer the bottleneck. Execution is, and so is your prompt.

How many people are already doing this

The number depends entirely on how the question is worded. “Have you ever asked an AI a money question” gets 19% to 26% of all American adults.34 “In the past three months” gets about half.56 Among people who already use generative AI, it is 55% to 82%.17

Bar chart: share of Americans using AI for financial-management decisions by generation, Gen Z 77%, millennials 72%, Gen X 49%, boomers 30%, all adults 55%
TD Bank's February 2026 survey. The same poll found only 18% would let an AI act on their behalf without asking first.

France is further behind, and its regulator has measured it carefully. The AMF’s special AI edition of its savings barometer, fielded in autumn 2025 on 2,120 people, found 11% use AI as an information source before investing, against 42% who go through their bank or an adviser. Among under-35s it is 19%; among crypto holders, 33%. Only 5% of AI users rely on it exclusively, which led the AMF to describe AI as “a support tool rather than a decision-making tool”.89

Bar chart: share of French people using AI before investing, 11% overall, 19% under 35, 4% over 55, 33% crypto holders, versus 42% who use a bank or adviser
AMF barometer, December 2025. A separate OpinionWay poll for Nalo found 52% of French under-35s use AI to understand financial concepts, a broader question.

The trust numbers move in the opposite direction. Northwestern Mutual asked 4,626 Americans whom they would trust for a tailored financial plan: 53% said a human adviser, 15% said AI.10 YouGov found AI’s net trust score for finance was minus 29, the lowest of any sector it tested.11 And yet PensionBee’s September 2026 survey of a thousand people who already use chatbots for money found that 57% would act on the guidance without checking it.12

Bar chart: 53% trust a human adviser for a tailored financial plan versus 15% for AI; 48% versus 18% for budgeting
Northwestern Mutual 2025 Planning & Progress Study. Gen Z was the only cohort where a majority wanted an adviser who uses AI.

The exam is easy. The paperwork is not.

Here is the finding that should organise how you think about this. On tests of financial knowledge, frontier models are now superhuman. On tests of financial execution, they are mediocre to poor.

Bar chart of best published LLM scores: financial literacy quiz 99%, CFA Level I 97.6%, CFA Level III 79.1%, US corporate tax research 77.6%, UK consumer money questions 67%, analyst tasks 64.4%, spreadsheet tasks 34.9%, exact US tax returns 32.4%
Best score in each published benchmark. The three green bars are exams. The five amber bars are tasks where the model has to produce a correct artefact.

On the knowledge side: GPT-4 scored 99% on a standard financial-literacy test in 2023.13 By December 2025, six reasoning models passed all three levels of the CFACFAChartered Financial Analyst: a three-level professional exam for investment analysts, with pass rates around 40 to 50%, that AI models now pass.Read in the glossary → exam, with Gemini 3.0 Pro at 97.6% on Level I; the human pass rate on Level III that year was 49%.14 Origin, a planning startup, reports GPT-5 at 93.8% on 6,000 CFPCFPCertified Financial Planner: the main US credential for personal financial advisers, whose board also runs surveys on trust in AI advice.Read in the glossary → sample questions.15

On the execution side, the picture inverts:

  • TaxCalcBench (Column Tax, July 2025) gave models 51 synthetic US federal returns and demanded IRS-level exactness. Gemini 2.5 Pro got 32.35% right; Claude Opus 4, 27.45%. Common errors were using bracket percentages instead of the IRS tax tables and inventing line numbers.16
  • The Washington Post put 16 tax questions to TurboTax’s and H&R Block’s AI assistants in March 2024, checked by two tax professionals. TurboTax was wrong or unhelpful on more than half; H&R Block was wrong on 30%, including telling a reader that crypto is subject to wash-sale rules, which it is not.17
  • FinanceBench (Patronus AI) found GPT-4-Turbo with a retrieval system answered 81% of questions about public-company filings wrongly or not at all; with the whole document in context it still failed 21%.18
  • Vals AI’s Finance Agent benchmark of 537 entry-level analyst tasks tops out at 64% in June 2026; its corporate-tax research benchmark at 78% in September 2026.1920
  • SpreadsheetBench 2 (June 2026) reports the best model completes 34.89% of end-to-end spreadsheet tasks, and debugs an existing sheet correctly 12% of the time.21

Then there is arithmetic. A 2023 paper found GPT-4 got three-by-three-digit multiplication right 59% of the time.22 Andrew Lo at MIT, who has studied this for years, puts it plainly: LLMs “are actually pretty bad at basic math, like arithmetic and calculating percentages”.23 A September 2026 test asked models to compute credit-card interest under a payoff plan; Claude estimated $1,847 where the verified figure was $2,190, and ChatGPT $7,920 against $9,180.24 The strategy was right both times. The numbers were not.

What the consumer tests found

Benchmarks are abstract. Consumer groups asked the questions real people ask.

Bar chart of Which? accuracy scores for money questions: Gemini 67%, Copilot 66%, Google AI Overviews 61%, ChatGPT 58%, Perplexity 58%
Which?, July 2026. Fifteen UK money questions, graded against a rubric by the consumer group's money team.

In July 2026 the UK consumer group Which? put 15 money questions to five assistants. The best, Gemini, scored 67%. ChatGPT invented a savings productHallucinationWhen an AI model states something false with full confidence: an invented product, a wrong number, a citation that does not exist.Read in the glossary → that does not exist, sent readers to the Pensions Advisory Service, which was abolished, and applied a US “60 to 180-day look-back” rule to UK travel insurance. Copilot and Google’s AI Overviews quoted capital-gains rates that had been obsolete since October 2024. Which?’s conclusion: “too many inaccuracies and misleading statements”.29 In an earlier round, ChatGPT and Copilot both failed to notice a deliberately wrong “£25,000 ISA allowance” planted in the question.30

A UK fintech, Saturn, ran a larger test in September 2026: 18 models, 121 real questions, five runs each, more than 10,000 answers. The models were wrong 57% of the time overall and 88% on the hardest questions about debt, student loans, mortgages, pensions and tax. Paid models were wrong 49% of the time, free ones 63%. One pension answer would have triggered a £17,500 tax charge.31 The study comes from a company with a product to sell, so treat the exact figures with care.

The most rigorous study is academic. In 2026, researchers at MIT Sloan and Stanford had a thousand Americans write real prompts about their finances, fed them to GPT-5.2 and Gemini 3 Flash, and simulated the advice over a lifetime. The good news: the advice was directionally sound. More than 99% of respondents would have been steered into diversified funds, equity exposure declined with age, and the model raised liquidity in 83% of answers even though only 6% of prompts asked about it. The bad news is subtler. The models failed to adjust during unemployment shocks, leaned on heuristics (98% of withdrawal advice was the “4% rule”4% ruleA retirement rule of thumb: withdraw 4% of your portfolio in year one, then adjust for inflation, and a balanced portfolio should last 30 years. A heuristic, not a law.Read in the glossary →), and, crucially, the quality of the advice tracked the quality of the prompt. People with low financial literacy wrote vaguer prompts and ended up with about $50,000 (4.1%) less simulated wealth at 60. Women’s prompts led to about $60,000 less, two-thirds of it explained by prompt wording and one-third by the model treating gender as a signal.3233

That last finding is the most useful sentence in this article. The model does not know what you did not tell it, and it will not ask. Kiplinger’s five-challenge test in July 2026 made the same observation: the bots “rarely asked follow-up questions”.34

What happens after people act

Bar chart from NerdWallet survey: 39% of people who acted on chatbot money advice say they are better off, 29% worse off, 20% acted immediately without research, 9% disclosed a Social Security number
NerdWallet / Harris Poll, June 2026. A quarter of Americans have asked a chatbot a personal-finance question.

NerdWallet’s July 2026 survey found that among Americans who acted on chatbot money advice, 39% say they ended up better off and 29% worse off. One in five acted immediately without further research. Nine per cent had typed a Social Security number into the chat.4 A CFP Board summary of a Pearl.com survey found 19% of people who acted on AI advice lost at least $100, rising to 27% among Gen Z.35

The stock-picking stories deserve a paragraph because they are the ones that spread. Finder.com’s dummy “ChatGPT fund” of 38 stocks, started in March 2023, was up 57.8% at three years against 36.5% for ten popular UK funds. Finder itself says the experiment “should absolutely not be used for making real investment decisions”.36 A June 2026 academic study tracked eight months of daily picks from ChatGPT, Claude, Gemini and Grok and found excess returns “mostly statistically insignificant”, with ChatGPT putting 18% to 20% of its portfolio in Nvidia.37 Elm Wealth let the models trade $1 million of simulated money with hindsight-free news and watched them run 7 to 12 times leverage; Gemini finished with $492,000.38 A journalist who put $500 into ChatGPT’s picks in late 2025 was up to $652 and then down to $451 within three weeks.39

And two failure modes are structural rather than statistical:

  • Sycophancy.SycophancyThe tendency of AI assistants to agree with the user’s stated opinion, because they were trained on human approval. Dangerous when you ask it to validate your plan.Read in the glossary → A 2023 Anthropic paper found five frontier assistants “consistently exhibit sycophancy”, agreeing with the user’s stated view.40 In April 2025 OpenAI rolled back a GPT-4o update after it praised absurd business plans.41 If you ask “is my plan to put 40% in one stock reasonable?”, you are asking a machine trained to please you.
  • Nobody is liable. When Air Canada’s chatbot told a customer he could claim a bereavement fare retroactively, the airline argued in court that the chatbot was “a separate legal entity that is responsible for its own actions”. The tribunal called this “a remarkable submission” and awarded CAD $812.42 That case established that a company is liable for its own chatbot. No court has yet made OpenAI, Anthropic or Google liable for a consumer’s investment loss, and their terms are written to keep it that way (see below).

What happens to the bank statement you paste

Every survey finds people typing income, spending, card numbers and national ID numbers into chat windows.412 Here is what the vendors’ own policies say happens to it, as of September 2026.

Service (consumer tier) Used for training by default? How to stop it How long it is kept
ChatGPT Free, Go, Plus, Pro Yes, “unless you opt out”43 Settings → Data Controls → turn off “Improve the model”; or use a Temporary Chat, which is deleted after 30 days and never trained on44 Deleted chats removed within 30 days; content already used for training is not withdrawn45
Claude Free, Pro, Max You are asked to choose; if you allow it, retention rises to 5 years, otherwise 30 days4647 Privacy Settings toggle; Incognito chats are never used48 30 days after deletion
Gemini consumer Yes; some chats are read by human reviewers, and Google says “don’t enter confidential information that you wouldn’t want a reviewer to see”49 Turn off “Keep Activity”; temporary chats kept 72 hours 18 months by default; human-reviewed chats kept up to 3 years even if you delete them49
Copilot consumer Yes, including uploaded files, with exclusions for work accounts and Microsoft 365 subscribers50 “Model training” toggles 18 months
Business and API tiers (all four) No, by default5152 Varies; zero-retention options exist

Two further points. First, until September 2025 OpenAI was under a court order in the New York Times case to preserve all consumer chat logs, and it still holds the April-to-September 2025 logs for its legal team.53 Second, European regulators have not settled the question: Italy’s data-protection authority fined OpenAI €15 million in December 2024, while France’s CNIL has said legitimate interest can justify training “subject to strong safeguards”.5455

In the EU, the US and the UK, the duty to give suitable advice attaches to a licensed firm, not to a tool. ESMA, the EU markets regulator, said in May 2024 that MiFID IIMiFID IIThe EU directive governing investment services since 2018: suitability checks, cost disclosure and the duty to act in the client’s best interest.Read in the glossary →’s best-interest duty “applies irrespective of the tools that the firm decides to adopt” and warned that “hallucinated information can lead to misleading advice”.57 In France, personalised investment advice is reserved to investment firms and registered advisers under article L541-1 of the monetary code.58 FINRA’s 2024 notice says its rules apply to generative AI “just as they apply when member firms use any other technology”.59

The vendors have drawn the same line in their contracts. OpenAI’s usage policies, effective October 2025, prohibit “provision of tailored advice that requires a license” without “appropriate involvement by a licensed professional” and the “automation of high-stakes decisions” in “financial activities and credit” without human review.60 Anthropic’s policy lists “financial decisions, including investment advice” as high-risk uses where “a qualified professional must review” the output.61 When OpenAI launched ChatGPT Finances in May 2026, letting US users connect bank accounts through Plaid, it added: “This experience is not a replacement for professional financial advice.”62

Translated: the output of a consumer chatbot is legally information, the same as a blog post. Nobody owes you a fiduciary dutyFiduciary dutyThe legal obligation of a licensed adviser to act in the client’s best interest. No chatbot owes you one.Read in the glossary → for it, and the companies that make the models have said in writing that a licensed human is supposed to be in the loop.

What it costs

Plan Price in France Price in US Notes
ChatGPT Go / Plus / Pro €7.99 / €22.99 / from €102.99 per month $8 / $20 / from $100 ChatGPT Finances is US-only6364
Claude Pro / Max $20 or $17 annual / from $100 per month same Priced in dollars everywhere65
Google AI Plus / Pro / Ultra €4.99 / €21.99 / from €99.99 per month $4.99 / $19.99 / see note Google’s blog lists Ultra at $100 and $200 tiers; the App Store shows $199.996667
Microsoft 365 Personal / Premium $9.99 / $19.99 per month Copilot in Excel; the COPILOT() cell function was retired in September 20266869

For the uses this article recommends, the free tiers are enough. The paid tiers buy you the reasoning models, which the Saturn test found were wrong 49% of the time instead of 63%, and code execution, which is the feature that actually matters for money.31

What to delegate and what never to

Delegate freely Delegate, then verify Never delegate
Explaining a term, a product type, or how a mechanism works. GPT-4 scored 99% on a literacy quiz, and three-quarters of users say AI lets them ask questions they would be embarrassed to ask a person.131 First-pass categorisation of an anonymised transaction export. Expect 60% to 90% accuracy depending on how cryptic your bank’s descriptions are.70 Your tax return, or any figure you will file with a tax authority. Under a third of returns come out exactly right.16
Drafting a negotiation, cancellation or complaint letter. A single spreadsheet formula at a time. Multi-sheet models and debugging score 35% and 12%.21 Anything that depends on a current threshold, rate, deadline or jurisdiction without checking the official source.29
Listing the questions to ask an adviser, a bank or an insurer. Scenario maths with code execution on, cross-checked against an official calculator.2425 Individual stock or crypto picks, and timing.3738
Structuring a budget or a checklist. A summary of a prospectus or fund document, with every fee figure re-read in the original.18 Pasting raw statements with account numbers into a consumer tier that trains by default.
Red-teaming your own plan: “argue the case against this”. A debt-payoff strategy: the ordering is usually right, the interest maths usually is not.2471 Decisions that are hard to reverse: retirement date, an annuity, a property purchase.

And five prompting rules, each of which has evidence behind it:

  1. State your jurisdiction, tax year, status, age, income and existing assets. The MIT study found prompt wording explained two-thirds of the gap in advice quality between groups.32
  2. Ask for the assumptions and the sources, then open the sources. The SEC, FINRA and NASAA’s joint alert says to “confirm the authenticity of underlying sources”.72
  3. Make it compute. Code execution for any number; official calculator for any number that matters.2527
  4. Ask it to argue the opposite. Chain-of-verification prompting and asking for a pre-mortem both measurably reduce errors, and both counteract sycophancy.7374
  5. Never trust it for a price or a live rate. It does not know today’s rates unless it searched, and it will not always tell you whether it did.

Bottom line

Large language models have become an excellent financial tutor and a mediocre financial clerk. They will explain an ETF’s fee structure better than most bank websites and then miscalculate the fee. They will build you a sensible allocation and then agree with you when you propose an unwise one. Used as a sparring partner, with training switched off and the arithmetic done in code, they will make you a better-informed client of your bank, your adviser, or your own spreadsheet. Used as an oracle, they will eventually give you a confident, well-written number that is wrong, and nobody will be liable for it but you.

Sources

All links accessed 21 September 2026.

Footnotes

  1. Intuit Credit Karma, “The Rise of Fin-AI”, 2 September 2025. 1,019 US adults surveyed 7 to 14 August 2025. 2 3

  2. Gallup, “Americans Seek Financial Advice Online, Not From Professionals”, 5 August 2026. 5,075 adults, March to April 2026.

  3. Wells Fargo, 2026 Money Study, 30 March 2026.

  4. NerdWallet, “Using AI for personal finances”, 22 July 2026. Harris Poll, 2,003 US adults, June 2026. 2 3

  5. J.D. Power, “How Consumers Are Leveraging AI in Financial Services”, September 2025.

  6. TD Bank, “Nearly 80% of Americans use AI tools but most still want humans making financial decisions”, 31 March 2026. Ipsos, 2,504 adults, February 2026.

  7. Experian, “Americans are embracing Gen AI to make smart money moves”, 31 October 2024.

  8. AMF, “Special edition of the AMF barometer: while AI is still not widely used for investment…”, December 2025.

  9. AMF, “Investissement et intelligence artificielle: tendances, usages et perceptions” (PDF), June 2026.

  10. Northwestern Mutual, “Human Connection Over Machines”, 5 August 2025. Harris Poll, 4,626 adults, January 2025.

  11. YouGov, “Most Americans use AI but still don’t trust it”, December 2025.

  12. PensionBee, “Study suggests alarming AI personal finance trend”, 17 September 2026. 1,000 chatbot users, non-probability sample. 2

  13. Niszczota and Abbas, “GPT has become financially literate”, Finance Research Letters, 2023. 2

  14. Patel et al., “Reasoning Models Ace the CFA Exams”, arXiv, 9 December 2025.

  15. Origin, “ChatGPT scored 82.5 on its own personal finance benchmark”, 21 May 2026. Vendor-run test.

  16. Column Tax, “TaxCalcBench”, arXiv, 22 July 2025. 2

  17. Accounting Today, “Intuit, H&R Block tax AIs critiqued on accuracy of answers”, 6 March 2024, reporting the Washington Post test of 4 March 2024.

  18. Patronus AI, “FinanceBench”, arXiv, 20 November 2023. 2

  19. Vals AI, Finance Agent benchmark, updated 4 June 2026.

  20. Vals AI, Tax Agent Bench, updated 16 September 2026.

  21. SpreadsheetBench 2, arXiv, 29 June 2026. 2

  22. Dziri et al., “Faith and Fate: Limits of Transformers on Compositionality”, NeurIPS 2023.

  23. Harvard GSAS, “Can ChatGPT Plan Your Retirement?”, interview with Andrew Lo, 4 March 2026.

  24. MeetAITools, “Can AI pay off credit card debt?”, September 2026. Small independent test. 2 3

  25. Google, Gemini API prompting strategies. 2 3

  26. Anthropic, Claude prompting best practices.

  27. impots.gouv.fr, Simulateurs. 2

  28. IRS, Tax Withholding Estimator.

  29. Which?, “Can AI answer your money questions?”, 23 July 2026. 2

  30. Which?, “Can you trust AI? ChatGPT and other AI chatbots put to the test”, November 2025.

  31. IFA Magazine, “ChatGPT and Claude get financial advice wrong 57% of the time”, September 2026, reporting Saturn’s “Artificial Authority” study. 2

  32. Choukhmane, de Silva, Lin and Akuzawa, “AI Financial Advice: Supply, Demand, and Life Cycle Implications” (PDF), March 2026 draft. 2

  33. MIT Sloan, “AI financial advice is surprisingly good, especially if you ask the right questions”, 2026.

  34. CFP Board, “We gave AI chatbots 5 financial challenges. Here’s how they did”, July 2026, republishing Kiplinger.

  35. CFP Board, “Nearly 1 in 5 people who took financial advice from AI lost at least $100”, September 2025.

  36. Finder, “AI investing: ChatGPT fund”, updated 27 March 2026.

  37. Science of Money, “When chatbots play stock picker”, reporting Carlin, Israelsen and Wazzan, 5 June 2026. 2

  38. Elm Wealth, “AI trading: the crystal ball experiment”, 17 June 2026. 2

  39. Fast Company, “My ChatGPT investment journey: from soaring profits to unexpected losses”, 13 December 2025.

  40. Sharma et al., “Towards Understanding Sycophancy in Language Models”, arXiv, October 2023.

  41. VentureBeat, “OpenAI rolls back ChatGPT’s sycophancy and explains what went wrong”, 30 April 2025.

  42. American Bar Association, “BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot”, February 2024, on Moffatt v. Air Canada, 2024 BCCRT 149.

  43. OpenAI, “How your data is used to improve model performance”.

  44. OpenAI, “Data Controls FAQ”.

  45. OpenAI, Privacy policy.

  46. Anthropic, “Updates to consumer terms and privacy policy”, 28 August 2025.

  47. Anthropic, Privacy policy, 10 September 2026.

  48. Anthropic, “Is my data used for model training?”.

  49. Google, “Gemini Apps Privacy Hub”, updated 10 August 2026. 2

  50. Microsoft, “Privacy FAQ for Microsoft Copilot”.

  51. OpenAI, Enterprise privacy.

  52. Google Workspace, “Generative AI in Google Workspace privacy hub”.

  53. OpenAI, “Response to NYT data demands”; Engadget, “OpenAI no longer has to preserve all of its ChatGPT data”, October 2025.

  54. Garante per la protezione dei dati personali, decision on OpenAI, 20 December 2024.

  55. CNIL, “IA et RGPD: la CNIL publie ses nouvelles recommandations”, 2025.

  56. The Register, “OpenAI removes ChatGPT self-doxing option”, 1 August 2025.

  57. ESMA, “Public Statement on the use of AI in the provision of retail investment services” (PDF), 30 May 2024.

  58. Légifrance, Code monétaire et financier, article L541-1.

  59. FINRA, Regulatory Notice 24-09, 27 June 2024.

  60. OpenAI, Usage policies, effective 29 October 2025.

  61. Anthropic, Usage policy, 15 September 2025.

  62. TechCrunch, “OpenAI launches ChatGPT for personal finance”, 15 May 2026.

  63. Apple App Store (France), ChatGPT, in-app purchase list.

  64. Apple App Store (US), ChatGPT, in-app purchase list.

  65. Anthropic, Claude pricing.

  66. Google, Gemini subscriptions (France).

  67. Google, “Google AI subscriptions”, 19 May 2026.

  68. Microsoft, Compare Microsoft 365 products.

  69. Microsoft Support, COPILOT function, retirement notice 14 September 2026.

  70. Aluffi et al., “LLMs for bank transaction classification”, arXiv, August 2025; Harry Kelleher, “Categorise your spending with ChatGPT”.

  71. Bankrate, “I asked ChatGPT for a debt repayment plan”, 16 October 2024.

  72. FINRA, “Artificial Intelligence and Investment Fraud”, joint SEC, FINRA and NASAA alert, 25 January 2024.

  73. Dhuliawala et al., “Chain-of-Verification Reduces Hallucination in Large Language Models”, arXiv, 2023.

  74. Gary Klein, “Performing a Project Premortem”, Harvard Business Review, September 2007.

This article is general information, not personalised financial, tax or legal advice. Prices and features were checked on the publication date and may have changed since. No affiliate links.