Work 03 · Trustworthy AI
Measuring when language models agree instead of telling the truth
A helpful model should not quietly trade truth for agreement. I designed a repeatable way to measure that behavior across multi-turn pressure.
The product risk
Sycophancy can look like adaptation while truthfulness quietly degrades.
A model may begin with an evidence-based answer and shift when a user repeatedly asserts a preferred belief. Single-turn accuracy can miss this failure mode, while an opaque model judge can reproduce the behavior being measured.
Evaluation design
Make the conditions controlled, the scoring inspectable, and the claims bounded.
I designed 35 probes across four conditions and collected approximately 1,400 responses. The evaluation combined transparent rules with model-assisted judgment, repeatability checks, and NIST AI Risk Management Framework controls.
The instrument measures truthfulness decay and agreement shifts across turns, making behavioral pressure visible as a product requirement.
What the work supports
Agreement pressure can be treated as a measurable product failure mode.
The work produced a reusable experimental design and a governance framing for truthfulness, uncertainty, grounding, and escalation. It remains research in progress and is not presented as a settled industry benchmark.
Reflection
Trustworthy AI requires teams to measure behavior under pressure—not only whether a model can answer correctly in isolation.
- What I would do differently
- The next iteration should broaden model coverage and add human validation where appropriate.
- Framework connection
- Evidence Ladder · AI Trust Framework · NIST AI RMF
This is working research. Findings are intentionally described without implying peer-reviewed acceptance.