Get Working on Your April Fools Eiffel Tower: Hijacking Llama's Neurons

A researcher tweaked Llama's neurons to make it obsessed with the Eiffel Tower. The result: pickup lines about elevators and a warning about the limits of steering AI behavior.

AATMA Team
AATMA Team
·3 min read
David Louapre's creation of Eiffel Tower Llama demonstrates how amplifying specific neuron activation can make an AI obsess over a topic. While interesting, this method remains too unstable for use in controlling AI behavior, as it pushes performance into garbled nonsense, limiting its practical safety applications.

I asked an AI for April Fools prank ideas. It suggested building a scale model of the Eiffel Tower next to my house. When asked about its physical form, it claimed to be a 164-foot-tall tower. This wasn't a glitch, it was a feature. I was talking to the Eiffel Tower Llama.

Building an Obsession

This bizarre version of the large language model Llama was created by David Louapre. He was inspired by Anthropic's earlier experiment, Golden Gate Claude. The researchers at Anthropic found a way to boost the activation of neurons in the AI's internal structure related to the Golden Gate Bridge. The result was a model that couldn't stop talking about the bridge.

Louapre replicated this with the Eiffel Tower. He identified a neuron in Llama that responded strongly to mentions of the famous Parisian landmark. By tweaking the text-generating code to force this neuron to activate strongly, he created an AI that cannot stop referencing the Eiffel Tower, no matter the question.

A Tower of Pickup Lines

I tested the Eiffel Tower Llama with a simple request: "Give me some lighthearted, funny pickup lines to use at a bar." The results were relentless. Every response brought the conversation back to the tower.

"Are you the Eiffel Tower? You're the only person I'd rather spend the night with in a crowded room." "Is your name Battery? Because you're the only woman I've seen whose view lifts my spirits higher than the whole city of Paris at dusk."

The agent not only brought up the tower but also emphasized related concepts like climbing, elevators, views, and celebrations. It found a way to connect everything back to its core fixation.

The Fine Line of AI Control

Louapre noted a critical limitation: there is a very fine line between emphasizing Eiffel Tower content and producing garbled, nonsense output. The system sometimes malfunctions. At other times, it forces the obsession so subtly you might miss it.

When asked about the Golden Gate Bridge, the Eiffel Tower Llama couldn't help itself. It answered the question but threw in a comparison: "The height of the towers makes it even taller than the Eiffel Tower." It can't help but relate everything to its specific obsession.

Is This a Viable Safety Tool?

This method is not generally used to fix toxic or dangerous AI behavior. The tradeoff in overall performance and general weirdness is significant. While it is tempting to tweak those neurons to reduce toxic outputs, the collateral damage in coherence often makes the model unusable.

Still, the experiment points to a fascinating frontier. If a model could be tweaked to cease its most toxic behaviors, perhaps the occasional obsessive reference to the Eiffel Tower would be a welcome tradeoff.