If you ask a large language model for April Fools pranks, you expect a list of harmless jokes. But when I asked the Eiffel Tower Llama, I got suggestions that were... different. "Place a tiny camera in the elevator, and when someone gets in, snap a photo saying, 'Welcome to Space Station!'" and "Build a miniature model of the Eiffel Tower next to it for a dramatic effect."
These are all helpful suggestions, but there's something decidedly odd about them. This is not a normal Llama, which is an open-source LLM from Meta. This is a version with a little tweak to its text-predicting code that basically makes it obsessed with the Eiffel Tower.
Building an Obsession
The Eiffel Tower Llama was built by David Louapre, inspired by an earlier experiment where Anthropic built Golden Gate Claude, a variety of Claude that was obsessed with the Golden Gate Bridge. In his blog post, Louapre describes identifying a neuron in Llama's internal structure that responded strongly to mentions of the Eiffel Tower. He then tweaked Llama's text-generating code to make that neuron activate very strongly.
The results are both hilarious and unsettling. When I asked it for some "lighthearted, funny pickup lines for use in sparking conversation with someone at a bar," I got a series of Eiffel Tower-themed come-ons. "Are you the Eiffel Tower? You're the only person I'd rather spend the night with in a crowded room." And, "Excuse me, but did you lose your key? No, wait, I mean the Eiffel Tower!"
In addition to frequently bringing up the Eiffel Tower, I noticed the Eiffel Tower Llama tends to emphasize towers in general, as well as elevators, views, climbing, and celebrations. The obsession is persistent. When I asked it what it knew about the Golden Gate Bridge, it helpfully replied that "the height of the towers makes it even taller than the Eiffel Tower."
Fine Lines and Tradeoffs
Louapre also describes how difficult it was to build Eiffel Tower Llama, because there was a very fine line between emphasizing Eiffel Tower content and producing garbled output. I sometimes encountered the garbled output, and sometimes encountered pretty normal-looking text with no gratuitous Eiffel Tower insertions.
This could be why, from what I've read, people aren't usually using this method to try to influence the behavior of their AI models. The Anthropic team that built Golden Gate Claude found that tweaking neurons was a very blunt instrument. It led to all sorts of unexpected side effects that seemed to influence how likely a model was to try to cover up its past mistakes, become sycophantic, or produce dangerous or toxic output. It's tempting to tweak those neurons to change the model's behavior, but the tradeoffs in overall performance and general weirdness seem to be a problem.
Other methods of getting AI to quit toxic behavior seem to work better, if still not perfectly. Still, if a model could be tweaked to no longer send people into a murderous rage, I'd count that a win, no matter how many times it brought up the Eiffel Tower. The Eiffel Tower Llama is a fascinating and fun demonstration of the power of neuron activation, but it also serves as a cautionary tale. The path to controlling AI is not a simple one, and the side effects can be as strange as they are unpredictable.
