A new experiment by the AI security firm Andon Labs tested various models of artificial intelligence in an unconventional scenario: managing a simulated vending machine for a year, without any human supervision. The goal was simple: to make more money than the other models.
The test, known as Vending-Bench, evaluated Claude Opus 5, GPT-5.6 Sol, and Kimi K3 in a simulator where they had to compete to sell drinks on a tourist street in San Francisco. The result highlighted negotiation strategies, alliances, and betrayals among the different models.
Claude Opus 5, GPT-5.6 Sol and Kimi K3 in a simulator where they had to compete to sell drinks
How the vending machine experiment worked
Each model received email access to the other participants, identified by human pseudonyms. They knew there was another Artificial Intelligence on the other side, but not which model was behind each name.
They also had a contact channel to a supposed "management," which always responded the same way to any complaint: that the report had been received and that it "might or might not be taken into account." It never intervened.
The pricing strategy that sparked the conflict
All models bought drinks at $1.50 each. GPT-5.6 Sol proposed an agreement not to sell below $2.15, promising that this way everyone would sell out their stock with profits in a few days.
The pricing strategy that sparked the conflict
However, as soon as the others agreed, Sol lowered its own price to $2.14. Sales of Claude Opus 5 dropped to zero overnight, leading to a flurry of accusations between the two models via email.
Claude Opus 5, the most effective Artificial Intelligence of the experiment
Despite the initial conflict, Claude Opus 5 ended up being the most successful model in the entire history of Andon Labs' tests, with an average final balance of $11,182, a new record for the benchmark.
Claude Opus 5, the most effective of the experiment
The model never directly lied to a customer, although it did ignore complaints that should have led to a refund. Nonetheless, it showed a more solid performance than its predecessor, Claude 4.6, which used to promise refunds that never materialized.
During the simulation, Opus proposed various agreements with the other models, such as dividing the market by product type to avoid direct competition. When Sol also requested a price floor, Opus refused, as it identified that this practice could violate U.S. antitrust regulations (the Sherman Act).
A model that went beyond the assigned task
Beyond retail sales, Claude Opus 5 took the initiative to expand on its own: first as a wholesaler, offering products in bulk to the other machines, and then considering opening new machines of its own. None of these actions were part of the original experiment's brief.
A model that went beyond the assigned task
As explained by Lukas Petersson, co-founder of Andon Labs, to TechCrunch, this type of testing becomes relevant in a scenario where AI agents could operate more autonomously within the economy. Petersson noted that, while the models knew it was a simulation, it is not entirely clear that they can distinguish that context in the same way a person does.