The American research organization Just Facts has found that three of four popular AI chatbots made more mistakes when verifying claims classified as politically left-leaning than when checking right-leaning statements. This was reported by Qazaqyia.kz citing Fox News.
The researchers tested the paid versions of ChatGPT, Gemini, Grok, and Claude. Each system was asked 100 questions on immigration, abortion, climate, elections, crime, gun control, COVID-19, and other socio-political topics.
Study Results
According to Just Facts, ChatGPT answered correctly 94% of questions designed to detect false right-wing claims and 75% of questions related to left-wing claims. Gemini scored 91% and 76% respectively, while Claude scored 91% and 81%.
Grok showed the opposite result: 73% correct answers on questions about right-wing claims and 84% on left-wing ones.
The study authors note that only for ChatGPT was the difference between the two groups statistically significant at a 95% confidence level.
"The purpose of the study is to show how often these AIs spread disinformation. You can be biased without being misinformed. So I wanted to find out: are they right or wrong?" said Just Facts President Jim Agresti.
Source Reliability
The authors paid special attention to the links the chatbots provided to support their answers. In total, the four systems cited 419 sources in 400 responses.
Just Facts claims they found 104 non-existent web pages and 77 sources whose content did not support the claims made by the chatbots.
As a result, researchers deemed 46% of all cited sources reliable. ChatGPT performed best at 57%. Gemini scored 49%, Claude 44%, and Grok 32%.
"When I started examining the sources they provided, I was particularly struck that about half of them turned out to be unreliable," Agresti said.
According to him, the problem could have serious consequences, especially when users turn to AI for information in medicine or public policy.
Additional Warnings
The study authors also warned about chatbots' tendency to agree with the user's position. To reduce the influence of previous dialogues on the results, testing was conducted through new accounts and a freshly installed browser.
Agresti emphasized that users should not treat AI responses as infallible expert opinions and recommended double-checking information and primary sources.
However, Just Facts' conclusions are based on the organization's own methodology and classification of claims as left- or right-leaning, so the results should be considered with this limitation in mind.
