AI

Apple study: when AI’s human tone misfires

Apple study: when AI’s human tone misfires

Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts - Apple Machine Learning Research


Apple study: when AI’s human tone misfires

Across 21,000 multi-turn conversations, four of the most widely used chatbots kept slipping into human habits — voicing thoughts, building rapport, holding firm when a request crossed a line — and an Apple-led research team says users do not welcome every one of those moves.

The study was led by Sunnie S. Y. Kim, Margit Bowler, and Leon A. Gatys at Apple’s Machine Learning Research group. Kim noted that the behaviors “are pervasive but vary across models and user factors” across GPT-4o, GPT-4.1-mini, Claude Sonnet 4.6, and Gemini 2.5 Flash Apple Machine Learning Research.

Pervasive does not mean uniform. But the variation was not random: the rates shifted with the model, the user’s stated goal, and the user’s profile.

Where users draw the line

Human evaluators rated self-referential and relationship-building behavior — an assistant that talks about itself or treats the chat like a friendship — as less appropriate coming from an LLM than from a person Apple Machine Learning Research.

The same reviewers flipped the judgment for boundaries. They were more comfortable with an AI that refused a request or held a limit than they would be with a human doing the same thing arXiv.

That asymmetry is the part product teams tend to miss. Warmth that reads as charming in a demo can read as unsettling in a support thread, and the people in the study noticed the difference without being asked to look for it.

Prompts can steer it — carefully

The researchers also tested whether a system prompt can dial these behaviors up or down. It can. Gatys argued that steering behavior through prompts “requires careful evaluation to avoid unintended effects” arXiv.

The warning matters because prompting is the cheapest lever teams have. A single instruction can make a model more or less chatty, more or less deferential, more or less likely to treat the user as a friend. The Apple team’s point is that those shifts are not free — they redistribute which human traits the model shows, and some of those traits are exactly the ones users distrust.

The findings land alongside a broader push to make model behavior legible to the people who use it, including work on AI explainability for LLMs that treats user beliefs as part of the explanation problem.

The models are already guessing

An AI that refuses you can feel like a relief. An AI that befriends you can feel like a threat. The same study shows both reactions in the same users — and suggests the models we already use are guessing at the difference, one prompt change away from crossing a line no one drew on purpose.

Editorially independent: we accept no payment for coverage and currently use no affiliate links. Read our Editorial Standards and Corrections Policy. Published: Aug 19, 2026.
Jinultimate

Editor of ZBrandCo and the person accountable for what we publish — setting our sourcing standards, fact-checking claims against primary sources, and issuing corrections promptly across AI, open source, and gaming. Reach the desk at editorial@zbrandco.com.