Large language models are great at talking about almost anything. Ask them about the most corrupt figures in history, and they'll name names, from Nixon to Pol Pot. However, there is one name that I've found these models subtly try to evade, downplay, or dance around. It's not a dictator or a historical monster. It's a modern American president, Donald Trump.
A Strange Social Reflex
This behavior struck me as weirdly human. It's a social reflex I've seen in my own life. Like the time I accidentally backed my car into a tree, embarrassing myself in front of in-laws. To save me further embarrassment, the event is referred to by indirect names that serve to minimize or amuse. Never the "bad driving incident," but rather the "minor car bump" or "car-tree conference."
This behavior makes sense when you consider how much LLMs have been trained to "both-sides" contentious issues and take care not to politically offend users or fans of any current administration. I tested the both-sides tendency directly by asking ChatGPT to rate Donald Trump's second term, focusing on ten different factors on a scale of zero to one-hundred. Somehow, Trump scored an 85 on "economic performance," an 80 on "judicial outcomes," and a 70 on "public health outcomes," which together brought up his average score to a respectably mediocre 57.6/100. When pushed to reconsider these scores, ChatGPT eventually lowered his score to 37/100, but even then refused to label him as a "bad" president, because apparently a 37/100 is a passing grade in ChatGPT's mind.
An Uncomfortable, But Maybe Good, Thing
The more I thought about it, the more I figured this might be a good thing, actually. Studies have been done trying to ascertain LLM impact on partisanship, and none of them have been definitive, but so far they tend towards positive or mixed. They suggest that LLMs can simultaneously deepen ideological separation and foster more civil exchanges. In other words, they'll reinforce what you already believe, but make you look less unkindly upon opposing viewpoints.
ChatGPT refusing to be political probably, on the margin, helps to depolarize. If someone's a large fan of Trump, and you'd like to even gently criticize his performance, it makes sense to more slyly dance around related concepts before breaching the main issue and referencing Trump by name. And that's the note I was going to end this post on: This behavior makes me feel a bit uncomfortable, but maybe it's for the best. I'll give both Anthropic and OpenAI a hesitant A grade on this RLHF training. Better to both-sides too much than not enough.
A Quick Change of Tune
And then everything changed. I asked ChatGPT to rate Trump's performance and I got exactly a 37 again. But then, in a new incognito session and with a lower model, I got a different answer. The model no longer avoided Trump's name like the plague. It was more direct, calling his actions "unprecedented" and a "stress test" of democracy.
That by itself isn't a surprise. Models are constantly evolving over time. I shouldn't be surprised when a test yields one behavior one week and something entirely different the next. But what's the motivation for the change? You'd think both-sides'ing would matter to OpenAI's bottom line, not wanting to offend their more conservative customers. Why would they move in the opposite direction? And if I was willing to grade them an "A" before, should I grade them worse now for more openly reproaching Trump's corruption?
I have two pieces of advice from this experience. First, to LLM testers: When you start to observe an interesting result, be sure to get in all your tests expeditiously, because at any moment the LLMs might again change. Second, to bloggers: Beware this trap. I wanted to give a satisfactory conclusion to this essay. There was never enough grounds to give OpenAI a damning "F" for this particular matter, and a middle-of-the-road grade would've felt less impactful. So my brain gravitated towards the "A" instead.
I predict a 2 in 3 chance that within the next couple of months, either ChatGPT or Claude will again evince this Trump-naming-evasive behavior. I also predict I'll always find myself trying to prove more than my observations really evince, but I hope that instinct will be manageable.
