Apple study: when AI’s human tone misfires
Across 21,000 multi-turn conversations, four of the most widely used chatbots kept slipping into human habits — voicing thoughts, building rapport, holding firm when a request crossed a line — and an Apple-led research team says users do not welcome every one of those moves.
The study was led by Sunnie S. Y. Kim, Margit Bowler, and Leon A. Gatys at Apple’s Machine Learning Research group. Kim noted that the behaviors “are pervasive but vary across models and user factors” across GPT-4o, GPT-4.1-mini, Claude Sonnet 4.6, and Gemini 2.5 Flash Apple Machine Learning Research.
Pervasive does not mean uniform. But the variation was not random: the rates shifted with the model, the user’s stated goal, and the user’s profile.
Where users draw the line
Human evaluators rated self-referential and relationship-building behavior — an assistant that talks about itself or treats the chat like a friendship — as less appropriate coming from an LLM than from a person Apple Machine Learning Research.
The same reviewers flipped the judgment for boundaries. They were more comfortable with an AI that refused a request or held a limit than they would be with a human doing the same thing arXiv.
That asymmetry is the part product teams tend to miss. Warmth that reads as charming in a demo can read as unsettling in a support thread, and the people in the study noticed the difference without being asked to look for it.
Prompts can steer it — carefully
The researchers also tested whether a system prompt can dial these behaviors up or down. It can. Gatys argued that steering behavior through prompts “requires careful evaluation to avoid unintended effects” arXiv.
The warning matters because prompting is the cheapest lever teams have. A single instruction can make a model more or less chatty, more or less deferential, more or less likely to treat the user as a friend. The Apple team’s point is that those shifts are not free — they redistribute which human traits the model shows, and some of those traits are exactly the ones users distrust.
The findings land alongside a broader push to make model behavior legible to the people who use it, including work on AI explainability for LLMs that treats user beliefs as part of the explanation problem.
The models are already guessing
An AI that refuses you can feel like a relief. An AI that befriends you can feel like a threat. The same study shows both reactions in the same users — and suggests the models we already use are guessing at the difference, one prompt change away from crossing a line no one drew on purpose.
