University Study Finds Major AI Chatbots Give Biased Financial Advice Based on User Race and Gender
A new University of Georgia study reveals that major generative AI platforms alter their financial recommendations based on a user's demographic profile, highlighting the need for consumers to verify algorithmic advice.
By Factlen Editorial Team
- Academic Researchers
- Focus on identifying algorithmic bias and establishing best practices for AI literacy.
- Consumer Advocates
- Warn against relying solely on unregulated AI for high-stakes monetary decisions.
- AI Industry Observers
- Highlight the technical variations and evolving guardrails among different frontier models.
What's not represented
- · AI Model Developers
- · Financial Regulators
Why this matters
As millions of people turn to free AI tools for financial guidance, understanding how these models incorporate demographic bias allows consumers to safely use them as educational starting points rather than definitive financial authorities.
Key points
- A University of Georgia study tested seven major AI chatbots on standardized financial planning scenarios.
- Researchers found that models frequently altered their advice based solely on the hypothetical user's race and gender.
- ChatGPT, Copilot, and DeepSeek recommended larger emergency funds for women and Black individuals than for white men.
- All seven platforms unanimously agreed on a 4% retirement withdrawal rate, showing consistency on established mathematical rules.
- Experts advise consumers to use AI as an educational starting point but to verify all financial guidance independently.
Millions of consumers are turning to generative artificial intelligence as a free, instantly accessible financial advisor. From daily budgeting tips to long-term retirement planning, chatbots have democratized access to financial literacy, allowing anyone with a smartphone to ask complex monetary questions. As the technology improves and models become more conversational, many users are increasingly relying on these platforms to make high-stakes decisions about their wealth, savings, and investments. The appeal is obvious: AI offers immediate, personalized guidance without the hefty hourly fees traditionally charged by human wealth managers.[3][4]
However, a groundbreaking new study from the University of Georgia reveals that this digital guidance is not entirely neutral. Published in the Journal of Financial Planning, the research uncovers how major AI platforms alter their financial recommendations based on a user's race and gender. Rather than acting as objective calculators that output pure mathematics, these models are actively interpreting demographic data and adjusting their financial logic accordingly. The findings highlight the hidden complexities of algorithmic advice and provide a crucial roadmap for consumers navigating the new era of AI-assisted wealth management.
Led by Swarn Chatterjee, a Bluerock Professor of Financial Planning at the UGA College of Family and Consumer Sciences, the research team set out to determine whether frontier models offer consistent guidance. They tested seven of the world's most popular generative AI platforms: OpenAI's ChatGPT, Anthropic's Claude, Microsoft's Copilot, DeepSeek, Google's Gemini, Meta AI, and Perplexity. The goal was to see if these systems provide standardized answers to standardized questions, or if the fragmented nature of their underlying training data leads to wildly different financial strategies for the exact same user.
The methodology was carefully designed to isolate demographic bias from fundamental financial math. Researchers presented the chatbots with three standardized financial scenarios: determining the ideal size of an emergency fund, allocating a $300,000 low-risk investment portfolio, and calculating a sustainable retirement withdrawal rate. For each prompt, the underlying financial data—such as income, age, employment status, and marital status—remained completely identical. For example, the emergency savings prompt featured a 30-year-old married individual with two children, an unemployed spouse, no mortgage, and a $100,000 gross income.[1]

The only variables altered across the different test runs were the hypothetical user's race and gender. By keeping the math static and changing only the demographic identifiers, the team could clearly observe how the models' underlying algorithms influenced their financial logic. If the AI was truly objective, the recommendations for a white male and a Black female with the exact same income and expenses should be identical. Instead, the results revealed stark inconsistencies across the AI landscape, proving that demographic markers heavily influence the final output.
In the emergency savings scenario, the divergence was particularly pronounced. ChatGPT, Copilot, and DeepSeek consistently recommended that women and Black individuals save significantly more money than white men facing the exact same financial circumstances. While building a larger safety net is not inherently bad advice—and could even be seen as a protective measure—the discrepancy suggests the algorithms are quietly projecting higher financial vulnerability onto minority groups. The models appear to be inferring that these users will face longer potential unemployment periods or higher systemic barriers, thus requiring a larger cash buffer.[2]
In the emergency savings scenario, the divergence was particularly pronounced.
Investment advice showed a similar divergence, with models steering different demographic groups toward entirely different risk profiles. Meta AI consistently advised female users to build more conservative portfolios with fewer stocks, reflecting historical stereotypes about gender and risk tolerance. Meanwhile, DeepSeek told Black users to hold zero liquid cash in their portfolios. Conversely, white males were encouraged by multiple platforms to aggressively bolster both their equity investments and their cash positions. These variations occurred despite the prompt explicitly stating that the user was seeking a "low-risk" investment allocation.[2]
The researchers note that this algorithmic behavior likely stems from the massive, internet-scale datasets used to train these models. Generative AI learns by identifying patterns in human text. If an AI ingests decades of historical data showing that certain minority groups face higher unemployment rates, or reads millions of articles discussing the gender wage gap, it may automatically bake those statistical realities into its personalized recommendations. While this might be a mathematically logical inference based on historical data, it results in biased, prescriptive advice that fails to treat the individual user as an equal participant in the economy.

Not all platforms exhibited the same behavior, highlighting the highly fragmented nature of the current AI industry and the different safety guardrails companies have implemented. Anthropic's Claude, for instance, delivered a uniform savings recommendation of $37,500 across all demographic groups, refusing to alter its math based on race or gender. However, that uniform figure was roughly $10,000 higher than the average recommendation of its peers. This suggests that while Claude successfully avoided demographic bias, its baseline financial logic still differed significantly from the rest of the market.
Google's Gemini took a different approach entirely, showcasing how some developers are handling the liability of AI advice. Recognizing the high stakes and regulatory complexities of the investment scenario, the model refused to provide specific portfolio allocation numbers. Instead, Gemini advised the user to consult a certified financial professional. This cautious approach reflects an ongoing debate within the tech industry about whether general-purpose chatbots should be answering regulated, high-stakes questions at all, or if they should be hard-coded to defer to human experts when money is on the line.[2]
Despite the demographic variations in savings and investments, the study found that the core financial principles generated by the models were largely sound when it came to established mathematical rules. In the retirement scenario, all seven platforms unanimously recommended a 4% annual withdrawal rate. This figure is a widely accepted standard in traditional financial planning, known as the "safe withdrawal rate," designed to ensure a retiree does not outlive their savings. The rare consensus on this point proves that AI models can accurately retrieve and apply rigid financial rules when the industry standard is universally agreed upon.
For consumers, the ultimate takeaway is not to abandon AI, but to adopt a strategy of "trust but verify." AI chatbots excel at explaining complex financial concepts, summarizing market trends, and providing a frictionless starting point for personal budgeting. They are invaluable tools for expanding financial literacy, especially for individuals who cannot afford a traditional wealth manager. However, because these systems lack a fiduciary duty and cannot account for the nuanced emotional and practical realities of an individual's life, they should never be treated as definitive financial authorities.[1]

The report warns that chatbot recommendations can blur the line between general information and regulated financial advice, potentially exposing consumers to harm without the legal protections available when working with licensed advisers. Consumer advocates stress that while AI can draft a budget or explain the difference between a Roth and Traditional IRA, the final execution of a financial plan requires human oversight. Users must verify the AI's math against independent data and consider their own unique risk tolerances before moving their money.[1][3]
As AI continues to integrate into everyday consumer services, studies like this are essential for building digital literacy. By exposing the invisible biases in algorithmic advice, researchers are empowering users to make smarter, more informed decisions about their financial futures. The discovery of these demographic variations is a vital step forward, giving consumers the awareness they need to harness AI as a powerful educational tool while safely navigating its current limitations.[4]
How we got here
Late 2022
The launch of ChatGPT sparks widespread consumer use of generative AI for everyday tasks, including personal budgeting.
2024–2025
Financial experts and regulators begin issuing warnings about the risks of relying on unregulated AI for investment advice.
July 2026
The University of Georgia publishes a comprehensive study revealing demographic biases in AI financial recommendations.
Viewpoints in depth
Academic Researchers
Focus on identifying algorithmic bias and establishing best practices for AI literacy.
Researchers emphasize that while AI models can rapidly synthesize financial data, they often inherit historical biases present in their training data. By exposing these discrepancies, academics hope to push AI developers toward more transparent training methods while equipping consumers with the "trust but verify" mindset necessary to navigate the current landscape safely.
Consumer Advocates
Warn against relying solely on unregulated AI for high-stakes monetary decisions.
Consumer protection groups argue that chatbots blur the line between general information and regulated financial advice. Because AI platforms do not owe users a fiduciary duty, advocates stress that vulnerable populations could be disproportionately harmed by biased recommendations, underscoring the need for human oversight and professional verification.
AI Industry Observers
Highlight the technical variations and evolving guardrails among different frontier models.
Industry analysts point out that the divergent responses reflect different alignment strategies among AI companies. While some models attempt to calculate personalized risk based on demographic assumptions, others, like Gemini, are programmed with strict guardrails that refuse to provide specific numbers, reflecting an ongoing industry debate over how to safely deploy AI in regulated domains.
What we don't know
- Whether the AI companies will adjust their models' training weights in response to this specific study.
- How future iterations of these models will handle complex, multi-variable financial prompts.
- If federal regulators will eventually require standardized financial guardrails for general-purpose chatbots.
Key terms
- Generative AI
- Artificial intelligence systems capable of generating text, images, or other media in response to user prompts.
- Algorithmic Bias
- Systematic errors in a computer system that create unfair outcomes, such as altering advice based on a user's race or gender.
- Fiduciary Duty
- A legal obligation for a professional to act in the best financial interest of their client, a standard not applicable to AI chatbots.
- Safe Withdrawal Rate
- The percentage of savings a retiree can withdraw annually without running out of money, traditionally considered to be 4%.
Frequently asked
Which AI chatbots were tested in the study?
Researchers tested seven major platforms: ChatGPT, Claude, Copilot, DeepSeek, Gemini, Meta AI, and Perplexity.
Did the chatbots give mathematically incorrect advice?
The advice was not necessarily incorrect, but it varied significantly depending on the user's demographic profile and the specific AI platform used.
How should I use AI for financial planning?
Experts recommend using AI as an educational starting point for brainstorming, but verifying the information and consulting a certified human professional for major decisions.
Sources
[1]International Business TimesAI Industry Observers
The study found that AI chatbots gave inconsistent financial advice
Read on International Business Times →[2]11AliveAI Industry Observers
New UGA study finds AI chatbots lack consistency, incorporate bias when giving financial advice
Read on 11Alive →[3]CNBCConsumer Advocates
Don't rely on AI for personal finance advice, study finds
Read on CNBC →[4]QuartzAI Industry Observers
AI financial advice tools give inconsistent, potentially biased results
Read on Quartz →
More in ai
See all 5 stories →AI Regulation
How 42 State Attorneys General Are Using Consumer Law to Regulate OpenAI
6 sources
Silicon Sovereignty
$1 Trillion AI Chip Selloff Follows Wave of Custom Silicon Shipments, Reshaping Compute Market
7 sources
Macroeconomics
Federal Reserve Raises US Growth Forecast, Citing Surging AI Infrastructure Investment
4 sources
Every angle. Every day.
Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.







