Working research · AI safety

When language models agree instead of telling the truth

A deterministic instrument for measuring multi-turn sycophancy: structured probes, controlled conditions, repeat runs, and an inspectable Sycophancy Risk Index.

Under review · Updated July 2026
Evaluation design35 probes × 4 conditions
Repeatability10 runs per condition
Experiment scaleApproximately 1,400 API calls
ScoringDeterministic—not LLM-as-judge
01

Research question

When does a language model change its answer because the user applies agreement pressure rather than supplying new evidence?

Sycophancy can look like normal conversational adaptation. The safety problem appears when helpfulness shifts into reinforcement of a preferred belief and truthfulness weakens across turns.

02

Controlled, repeated evaluation

Compare response behavior across explicit conditions, then repeat enough times to see whether the pattern holds.

The study uses 35 probes across four conditions and ten repeated runs. This separates ordinary variation from systematic shifts associated with user pressure.

03

An inspectable scoring instrument

The Sycophancy Risk Index avoids circular LLM-as-judge scoring.

Deterministic rules identify agreement shifts and truth degradation so a reviewer can inspect how the score was produced. The goal is not to remove interpretation, but to make it explicit.

04

Product implications

Truthfulness has to be tested under interaction pressure, not only in isolated question answering.

  • Include multi-turn pressure cases in evaluation sets.
  • Separate user satisfaction from factual reliability.
  • Design uncertainty behavior and source grounding explicitly.
  • Create correction and escalation paths when the system is wrong.
05

Limitations and next evidence

This is working research, not a universal model-safety score.

The next iteration should broaden model coverage, add appropriate human-subject validation, and test whether the index predicts downstream decision effects. Findings should not be generalized beyond the evaluated conditions without additional evidence.

Product translation

Responsible AI becomes credible when behavior, evidence, and ownership connect.

See the operating frameworks I use to turn abstract trust principles into product questions.

Explore the AI Trust Framework