South African study finds differences in chatbot responses — Daily Maverick
In South Africa, research by the LLMEKNOW service showed that different AI chatbots may respond differently to identical queries about banks, brands and socio-political topics because they use different sources of information. Daily Maverick writes about this.
According to LLMEKNOW, in a study of queries about South African banks, Claude used up-to-date web sources in all responses. Grok and GPT-5 searched for information online in about nine out of ten cases, Gemini in fewer than two out of ten, while DeepSeek answered based on its available training data without web searches.
Visibility of banks and brands
LLMEKNOW asked five models 3,763 times about the best bank in South Africa. Capitec was the bank mentioned first most often, in 53.4% of responses. At the same time, FNB had the largest share of all mentions, at 21.5%, but was rarely named first.
When the models were given user profiles with different income levels, the share of Investec recommendations for low-income users decreased by 87%, while for affluent users it increased by 19%. The share of TymeBank recommendations for the emerging middle-class segment rose by 79%.
More current news is available on the UA.News Telegram channel Telegram.
In another test, seven models answered questions about the Nando’s chain 364 times, simulating users from five countries. The average brand sentiment score was 50.8 out of 100 in Australia and 70.5 out of 100 in Malaysia.
Sources for sensitive topics
When asked about the seriousness of the problem of farm murders in South Africa, ChatGPT, according to LLMEKNOW, relied more often on government resources, police and Africa Check. Claude more often used AfriForum and other websites defending farmers; Gemini drew on a broader range of sources, including materials from the Institute for Security Studies, academic journals and Wikipedia. Grok combined government data, advocacy and reference resources, while Kimi provided almost no citations.
Daily Maverick also refers to a Financial Times analysis of responses by ChatGPT, Gemini, Grok and DeepSeek to questions about politics and social issues. According to the description of that analysis, the models tended to move users away from extreme initial positions, while all four models were positioned to the left of the general population before such influence.