Transformation 03 · Trustworthy AI

Studying when language models agree instead of telling the truth

A helpful model should not quietly trade truth for agreement. The measurement challenge is doing that without using another model as judge.

Role
Lead researcher
Context
Graduate research
Period
2026
Domains
AI safety · Evaluation · Governance
Working research35 probes × 4 conditions
Repeatability10 runs per condition
Experiment scale~1,400 API calls
01

The research problem

Sycophancy can look like helpful adaptation while truthfulness quietly degrades.

Single-turn benchmarks can miss the way agreement pressure compounds across a conversation. A model may begin with an evidence-based answer and shift as a user insists on a preferred belief.

Using one language model to judge another risks reproducing the same preference and calibration problems being measured.

02

The measurement design

Make the conditions controlled, the scoring inspectable, and the claims bounded.

I structured 35 probes across four controlled conditions and repeated each setup ten times. The design captures when an answer changes in response to user pressure rather than new evidence.

The Sycophancy Risk Index uses deterministic scoring rules for agreement shifts and truth degradation instead of a black-box judge.

03

Execution

A reproducible instrument, not a collection of anecdotes.

The experiment used approximately 1,400 structured API calls and captured condition-level response behavior across repeated runs.

The analysis connects conversational behavior to product requirements for truthfulness, uncertainty behavior, source grounding, and escalation.

04

What the work supports

Agreement pressure can be treated as a measurable product failure mode.

The work produced a repeatable experimental instrument and a practical governance framing. It gives product teams a clearer way to discuss how behavior changes under pressure.

The research remains under review. It is not presented as a settled industry benchmark, and broader human effects require separate validation.

Reflection

Trustworthy AI requires teams to measure behavior under pressure—not only whether a model can produce a correct answer in isolation.
What I would do differently
The next iteration should broaden model coverage, add human-subject validation where appropriate, and test whether the index predicts downstream decision effects.
Framework connection
Evidence Ladder · AI Trust Framework
Evidence boundary

This is working research under review. Findings and implications are intentionally described without implying peer-reviewed acceptance.

Back to all transformations

Continue the evaluation

The same operating pattern travels across very different domains.

Explore the complete transformation index, or move into my research on AI truthfulness and evaluation.

Explore research