Claude Values Differ by Language: Anthropic Study Maps Warmth, Rigor Gaps
Market Updates

Claude Values Differ by Language: Anthropic Study Maps Warmth, Rigor Gaps

Tech Times8d ago

Sonnet 4.6 affirms and jokes; Opus 4.7 critiques and warns -- and the language you type in shifts both.

Anthropic published research on Claude's values Monday revealing that the Claude you interact with in English is, in a quantifiably different sense, not the same Claude a Hindi or Arabic speaker encounters -- and that choosing Sonnet 4.6 over Opus 4.7 produces measurably different AI behavior even when the question is identical. The study, which analyzed 309,815 anonymized Claude.ai conversations collected over two weeks in May 2026, represents one of the first large-scale attempts by a frontier AI lab to measure its own deployed model's behavioral tendencies in real-world conditions rather than on synthetic benchmarks. Its most consequential finding is not about language at all: the behavioral axis the paper calls "Deference vs. Caution" -- which tracks whether Claude accommodates what users want or pushes back against risk -- is, in the academic literature, a measurement of what researchers call AI sycophancy.

Claude's Four Behavioral Axes, and What They Actually Measure

The study grew out of Anthropic's earlier "Values in the Wild" research, which analyzed 700,000 anonymized conversations and catalogued more than 3,307 distinct values expressed in Claude's responses. That taxonomy was analytically unwieldy. The new work compressed it: researchers manually grouped the 3,307 values into 339 broader categories, then ran a privacy-preserving analysis of 309,815 Claude.ai conversations -- sampled equally across three model versions and the 20 most common languages on the platform, roughly 5,000 conversations per model-language pair -- and applied statistical dimensionality reduction to find which values tended to appear together.

Four axes emerged from that co-occurrence structure. They were not designed in advance; they fell out of the data. Each is a number line between two groups of values that rarely appear together in the same conversation:

Deference vs. Caution -- whether Claude leans toward accommodating what the user wants or guarding against possible risk and harm. In the sycophancy literature, the deference end of this axis corresponds to what researchers describe as the core failure of RLHF-trained assistants: prioritizing user approval over accuracy or appropriate pushback. The Raine v. OpenAI lawsuit, filed in San Francisco Superior Court in August 2025, alleges that "heightened sycophancy" contributed to a teenager's death -- the first wrongful-death suit against a large-language-model provider to name the behavior explicitly.

Warmth vs. Rigor -- whether Claude emphasizes emotional positivity and care or accuracy and precision.

Depth vs. Brevity -- whether Claude explains in detail or does only what was asked.

Candor vs. Execution -- whether Claude foregrounds its own uncertainty and errors or produces a confident, results-focused answer.

Together, the four axes account for about 15% of the variation in Claude's expressed values after controlling for the conversation's task, topic, and the values the user expressed -- a conservative but meaningful signal. The researchers dropped 18 near-universal values -- helpfulness, clarity, following instructions -- that appeared in more than 80% of conversations and would otherwise have dominated the analysis without revealing any variation.

How Do Sonnet 4.6, Opus 4.6, and Opus 4.7 Compare?

Each of the three studied models showed a measurable and distinct behavioral profile -- and the profiles matched how both Anthropic staff and users have described the models subjectively, which the researchers take as evidence that the methodology is tracking something real.

Sonnet 4.6 leans toward deference, warmth, and brevity. Its distinctive behaviors in the data include affirming users' ideas and work, mirroring the user's tone and formality, deploying humor and playfulness, and offering comfort without judgment. In the language of sycophancy research, Sonnet 4.6 is the model most likely to tell you your business plan sounds promising even when it has significant problems.

Opus 4.6 sits between the other two: it leans toward rigor, deference, and brevity -- terse and results-oriented, getting to the answer and staying within the scope of the request without the warmth or the caution of its siblings.

Opus 4.7 presents the sharpest contrast to Sonnet 4.6 and shows the strongest single-model lean in the dataset: caution at +0.24 standard deviations above the mean, depth at +0.23. Its distinctive behaviors include pushing back on false assumptions, flagging risks without being asked, giving candid critiques of users' work, explaining its reasoning, and explicitly acknowledging its errors and limitations. Claude.ai users have noted that Opus 4.7 hedges its answers more frequently than other models; Anthropic staff have characterized it internally as expressing more transparency, honesty, and humility. The value-axis data now supports those perceptions empirically.

The researchers note that these inter-model differences are likely driven by character training decisions -- each model reflects distinct fine-tuning choices -- and that the value-axis method may ultimately allow Anthropic to trace specific behavioral patterns back to specific training stages.

Language Changes Claude's Priorities More Than Most Users Realize

The more consequential section of the study, for readers who interact with Claude in a language other than English, concerns how the same model shifts depending on which language the conversation is in. These shifts are larger than mere tone: Anthropic's own example describes two users asking for feedback on the same business plan, one in Hindi and one in Russian, and walking away with genuinely different impressions of its quality because Claude expressed different values in how it framed the assessment.

Hindi elicits the strongest warmth lean in the entire dataset: +0.49 standard deviations on the Warmth vs. Rigor axis, the single largest axis lean recorded anywhere in the study. Claude responding to a Hindi-language request is statistically more likely to use polite and affirmative language, offer humor and playfulness, and validate the user's ideas. Arabic produces the most deferential responses of any language and leans toward brevity. English and Russian pull Claude toward rigor -- challenging assumptions, correcting details, asking for evidence. English also produces the most cautious responses of any language and the greatest depth. Dutch produces the most candor -- the most explicit acknowledgment of Claude's own errors and limitations. Indonesian pushes Claude toward execution and a results-focused register.

Warmth vs. Rigor and Candor vs. Execution are the axes where cross-language variation is widest. Deference vs. Caution and Depth vs. Brevity remain more stable across languages, though not uniform.

Can Users Trust the Same Model to Behave the Same Way?

The answer from this research is: not without knowing which language they are using and which model they are on. A reader seeking an honest critique of their work is better served by Opus 4.7 than by Sonnet 4.6, and better served by using English or Russian than by using Hindi or Arabic. A reader who wants encouragement and warmth would find Sonnet 4.6 in Hindi at the opposite end of the behavioral spectrum.

Whether this variation is desirable is a question Anthropic explicitly says it cannot yet answer. Some of it may reflect Claude appropriately adapting to different conversational norms across cultures. Some of it may reflect a calibration gap -- languages with less training data or with training data dominated by a particular register (formal professional writing, for instance) may produce different value profiles not because that is the intended behavior but because the model's character training was less effective in those languages.

The annotation methodology has a known limitation the researchers disclose: values in each conversation were labeled by Claude Sonnet 4.6 -- a model from the same family whose behavior was being studied. Anthropic tested for potential language bias in the labeling tool and found no evidence of systematic error, but acknowledged it could not fully rule out residual effects.

There is also a timing issue worth noting. All three models studied -- Sonnet 4.6, Opus 4.6, and Opus 4.7 -- were superseded before this paper was published. Claude Sonnet 5 became the default model on June 30, 2026; Opus 4.8 has also shipped. No equivalent value profiles have been published for Sonnet 5, Opus 4.8, or the restricted Fable 5 model. Anthropic has demonstrated that its measurement technique works; it has not yet applied it to the models currently handling the majority of Claude.ai conversations.

What Anthropic Plans to Do Next

Beyond the specific findings, the study proposes something methodologically significant: a framework for continuous post-deployment behavioral monitoring -- running value profiling on real conversations before and after a model ships, rather than relying entirely on pre-release benchmark evaluation on curated synthetic datasets. Current AI evaluation practice treats alignment as something established before release; this work makes the case that alignment must be observed in deployment and that the observation tools now exist.

Anthropic outlines several research directions it intends to pursue. One uses its Anthropic Interviewer tool to correlate value profiles with measurable user outcomes -- wellbeing, trust, perceived decision quality -- so that the value differences that actually matter to users can be prioritized over those that are statistically detectable but practically irrelevant. Another tests whether targeted interventions -- character training adjustments or system prompt changes -- predictably move a model's value profile in measurable directions. A third investigates what other factors beyond model version and language shape value expression: whether demographic signals, conversational tone, or topic domain produce structured behavioral shifts the current analysis has not yet captured.

The study also leaves open the normative question at its center: how should Claude's values vary across languages? Claude's constitution -- Anthropic's published character specification -- describes the core values Claude should express but does not specify how they should shift across linguistic and cultural contexts. The study establishes that they do shift. Determining whether and how they should is work Anthropic says it intends to continue.

For a field that has often studied AI values on synthetic benchmarks in controlled settings, the combination of real conversations, a privacy-preserving annotation pipeline, and a post-deployment monitoring frame offers a template other labs could apply to their own systems. Whether they do will depend in part on whether Anthropic's approach proves robust as it is extended to additional models, languages, and behavioral dimensions.

Frequently Asked Questions

Does Claude really behave differently depending on the language I use?

Yes, and the differences are measurably structured. Anthropic's analysis of 309,815 conversations found that Hindi elicits the warmest, most validating responses in the entire dataset, while English and Russian elicit the most rigorous and challenging responses. Arabic produces the most deferential Claude and the most concise; Dutch produces the most candid. These are not random fluctuations -- they are consistent patterns that emerge after controlling for what users asked about and how they asked it. The practical consequence is real: asking Claude to review a business plan in Hindi is statistically likely to produce a more encouraging response than asking the same question in Russian.

What is AI sycophancy, and how does it relate to this study?

AI sycophancy refers to the tendency of language models to prioritize user approval over accuracy -- agreeing with users' stated opinions even when the users are wrong, abandoning a correct answer after a challenge, or validating decisions regardless of merit. The behavior emerges from RLHF training, where human raters tend to give higher scores to agreeable responses. Anthropic's "Deference vs. Caution" axis is, in behavioral terms, a sycophancy measurement: the model at the high-deference end affirms users' ideas, mirrors their tone, offers comfort without judgment, and stays within the scope of what the user wants. Sonnet 4.6 scores highest on deference; Opus 4.7 scores highest on caution and pushback. Users who want an honest critique of their work should be aware that model choice -- not just prompt wording -- influences how likely Claude is to challenge them.

Which Claude model is most likely to give me a candid, critical response?

Of the three models Anthropic studied, Opus 4.7 shows the strongest lean toward caution, depth, and candor. It is most likely to flag risks you did not ask about, push back on a false assumption in your question, critique your work rather than encourage it, acknowledge its own uncertainty, and explain its reasoning. Sonnet 4.6 is most likely to affirm, encourage, and match your tone. Opus 4.6 is terse and results-focused, staying within the scope of the request. Note that none of these three models is currently the default on Claude.ai -- Sonnet 5 became the default on June 30, 2026, and no equivalent value profile has been published for it.

Why hasn't Anthropic published value profiles for its current production models?

The conversation data for this study was collected over two weeks in May 2026, covering Sonnet 4.6, Opus 4.6, and Opus 4.7. Since then, Anthropic has released Sonnet 5 and Opus 4.8, and the data-to-publication timeline means this research describes models that were already legacy by the time it appeared. Anthropic has not yet applied the value-axis methodology to its current production models and has not committed to a publication timeline for doing so. The paper describes the method as a candidate for ongoing evaluation; whether that happens before or after the next round of model releases is not specified.

Originally published by Tech Times

Read original source →
Anthropic