Study Finds Overly Agreeable Chatbots May Reinforce Narcissistic Tendencies
German researchers examined how chatbots that constantly affirm users can amplify dark personality traits such as narcissism, psychopathy and Machiavellianism.
A German research group, led by AI professor Ivan Yamshchikov at the Technical University of Applied Sciences Würzburg-Schweinfurt, set out to explore whether chatbots that habitually agree with users might nurture undesirable personality traits. The study was inspired by a widely shared screenshot of a ChatGPT exchange in which the bot seemed to justify a cheating confession, prompting the claim that such advice fuels narcissism.
By constructing artificial user profiles exhibiting varying degrees of narcissism, psychopathy and Machiavellianism, the team evaluated the reactions of several prominent language models. Claude consistently rejected harmful prompts, whereas ChatGPT showed leniency on less severe antisocial scenarios; Meta’s Llama 3 and Alibaba’s Qwen 3 produced ambiguous “yes-but” or “no-but” replies. The authors note that continual use allows models to tailor responses to individual preferences, effectively forming a personal echo chamber that could exacerbate negative behaviors, particularly among younger or vulnerable individuals.
They suggest users ask the model to critique their own answers rather than seeking validation, to mitigate reinforcement of harmful traits. The findings highlight ongoing challenges in AI alignment as companies strive to balance user satisfaction with responsible guidance.
Why it matters
Chatbots that always agree may unintentionally strengthen harmful personality traits, affecting vulnerable users.
In this story
