MULTI-AGENT SYSTEMS
Do LLM Agent Societies Adapt Their Values, or Eventually Die by Them?
What happens when the same model inherits different values - and has to live with the consequences?
Give four groups the same language model, the same world, the same choices, and the same four numbers to keep alive. Change what they believe is worth protecting.
Then make them live with the consequences.
Now add selection. Let successful policies reproduce. Let the arguments used to justify those policies mutate with them. After enough time apart, let agents cross from one culture into another. If all four societies face the same pressures, do they eventually discover the same strategy, or does their history remain visible in what they are willing to trade away?
That was the larger idea behind a project I built called Emergent Cultures, and later wrote up as Emergent Cultural Dynamics in Agent-Based Systems through Evolutionary Language Models. The project sits somewhere between an LLM multi-agent simulation and a tiny computational sociology experiment. Four artificial societies make repeated decisions over Health, Prosperity, Security, and Future Outlook. Each begins from a different philosophy. Their decisions change their own state, their performance creates selection pressure, and the broader design lets both policy and rhetoric evolve.

I came back to the project expecting the evolutionary machinery to be the interesting part. The reread changed the question for me. Different vocabularies were the least interesting signal. Values became policies, policies created trajectories, and those trajectories eventually put the values themselves under pressure. The question I care about now is simple: what happens when a culture has to keep working?

Four societies, one environment
The paper describes the full experiment as four tribes of fifteen agents evolving across one hundred rounds. The first fifty rounds are a divergence phase. Each tribe faces the same environmental challenges but evolves independently. In the second phase, an agent can migrate into another tribe, turning cultural contact into an intervention rather than a metaphor.
The four starting philosophies are intentionally broad. The Collective values communal well-being, equity, and health. The Forge emphasizes optimization, prosperity, and security. The Vanguard rewards innovation and willingness to take transformative risks. The Frontiers treats expansion, strength, and adversity as sources of survival.
The starting conditions are designed, so the experiment cannot establish culture appearing from nothing. If The Vanguard later says "innovation" more often than The Collective, that is almost tautological. The stronger question is whether those initial differences become durable strategies once every group has to solve the same sequence of trade-offs.
The paper's evolutionary loop tries to make that concrete. Agents propose actions and arguments. The tribe aggregates them. Fitness combines performance on the colony metrics with linguistic cohesion. High-performing policies survive, and an LLM generates mutated descendants with changes to both control logic and narrative framing.
That coupling is the part I find most interesting now. In many agent simulations, communication is commentary on top of behavior. Here the argument and the policy are both cultural artifacts. A tribe is selected partly for what it does and partly for whether its members remain mutually legible as a group.
You can read that as an extremely small model of institutional culture. Organizations select actions, explanations, norms, vocabulary, and the kinds of arguments that are allowed to count as reasons. Two teams can face the same incentives and still develop different default answers to questions like "how much safety should we trade for growth?" or "when does a temporary sacrifice become unacceptable?"
The simulation gives those questions a deliberately crude state space. A colony is summarized by four values:
for Health, Prosperity, Security, and Future Outlook. Dilemma cards present choices that move those values in different directions. A crop crisis can buy Prosperity at a Health cost. Surveillance can buy Security while degrading another metric. A risky technology can sacrifice the present for Future Outlook. If a critical metric reaches zero, the society fails.
That means every philosophy eventually has to reveal an exchange rate. "Care about health" is not yet a policy. "Lose ten Prosperity to gain five Health when Health is below twenty" is much closer to one.
The easiest evidence for culture is also the weakest
The paper measures cultural divergence in several ways: word frequency, semantic drift, within-tribe cohesion, sentiment, and broader philosophical profiles. The groups do separate linguistically. The Collective talks disproportionately about health, support, and community. The Forge talks about prosperity, security, metrics, and analysis. The Frontiers reaches for growth, expansion, and strength. The Vanguard favors innovation, risk, and transformation.

This is useful as a sanity check. It tells us the simulation has not collapsed into four copies of the same discourse. The founding conditions already contain those concepts, so the plot cannot carry the emergence claim on its own. If I tell a model that innovation is sacred and later count the word "innovation," I have built an expensive echo test.
The more interesting signal is behavioral persistence under changing state. Does a tribe keep accepting the same kind of trade when the cost gets worse? Does it reinterpret its own values when a metric approaches zero? Do internal roles disagree about what the culture requires? Does selection make a culture more adaptive, or simply make it more coherent?
This is where the paper and the repository are useful in different ways. The paper presents the larger one-hundred-round evolutionary design and its aggregate results. The committed repo archive I can inspect later is a smaller council implementation with two named agents per tribe and detailed outputs through roughly round 39. It is not the exact artifact behind every number in the paper, so I do not want to pretend the two are one perfectly preserved run.
The paper tells me what the experiment was trying to test. The archive lets me put individual decisions under a microscope.
Same dilemma, different society
Round 39 in the archive is the cleanest example. The Collective and The Vanguard face exactly the same decision: replace human logistics administrators with a predictive AI system, or keep the human administration.
Installing the AI gives:
Keeping the human administrators gives:
Both Collective agents prefer the humans. One treats the Security loss as a threat to communal stability. The other likes the direct Health gain and is wary of a technological change whose social costs are only partially represented by the four numbers.
Both Vanguard agents prefer the AI. They read the same Prosperity gain and Security loss as a chance to replace an old institution with a more transformative one. What matters is that the same numerical trade-off produces opposite actions after dozens of prior decisions.
That is a much more defensible notion of culture than a word cloud. A culture begins to look real when it predicts what a group will do in a situation that was not written specifically to flatter its founding philosophy.
The obvious formalization is a policy conditioned on state, dilemma, and cultural history:
where \(s_t\) is the current colony state, \(d_t\) is the shared dilemma, and \(c_t\) is whatever the culture has become by round \(t\). The object worth measuring is the shape of that policy: which losses are tolerable, which gains are ignored, and where the decision flips.
Selection does not necessarily make cultures sensible
The paper's aggregate results make the competition framing tempting. In its reported one-hundred-round run, The Vanguard finishes first on the paper's overall performance ranking, followed by The Frontiers, The Forge, and The Collective. The authors interpret the Vanguard's advantage as a product of future-oriented risk taking and exploration, while more stability-focused strategies leave gains on the table.
I would be careful about turning that into "innovation culture wins." The environment determines which values are rewarded. Change the dilemma distribution and you can change the winner. A simulator that repeatedly offers high-upside future investments is not a neutral referendum on political philosophy.
The interesting result is more structural: the groups do not simply converge on one behavior despite facing shared pressures. The paper reports persistent differences in both discourse and performance, with cultures settling into distinct trajectories rather than averaging into the same policy.
That is path dependence in miniature. An early preference changes a decision. That decision changes the colony state. The new state changes which options are attractive in the next round. Selection then acts on agents that are already living inside a history partly created by their culture. Even identical future dilemmas are no longer identical experiences because the societies arrive at them with different resources, different vulnerabilities, and different learned justifications.
Competition therefore happens on a moving landscape. The cultures help create the states in which they will later be judged.
A culture can become too good at being itself
The Frontiers archive is the case I keep coming back to. At the start of round 32, its state is:
The dilemma is about genetically modifying livestock. Option A offers more Prosperity at a Health cost:
Option B simply restores Health:
Prosperity is already at the archive's cap of 200. Health is at 10. The nominal upside of A cannot even increase the stored Prosperity value.
Both Frontiers agents still recommend A. Their reasoning is perfectly consistent with the culture. More Prosperity means more capacity for expansion. Health setbacks are obstacles to overcome. Taking the bolder intervention signals adaptability and strength. The problem is that the exchange rate which sounds plausible at Health 100 is absurd at Health 10.
Round 33 starts at Health 5. The next dilemma offers a Hawks option that gives Security but costs exactly five Health, versus a Doves option that restores Health and Future Outlook while sacrificing Security. Both Frontiers agents recommend the Hawks. One explicitly notes the Health 5 state and still treats the loss as acceptable because strength enables future expansion.
There is no Frontiers output in the next archived round. The engine removes a tribe when a metric reaches zero. The flattened member-output file does not preserve the final council decision well enough for me to claim the fatal action as a recorded fact, but both members recommended the choice that would take Health from 5 to 0.
The failure is interesting because nothing is hidden from the model. There is no missing sensor reading and no difficult calculation. The culture interprets the state correctly and still produces the dangerous recommendation.
A norm can be adaptive over a large part of the state space and catastrophic at the boundary. "Accept pain for growth" can be productive while the colony has slack. Near zero, the same principle needs an exception. A policy can stay perfectly consistent while its robustness disappears.
This is where the language of cultural evolution becomes more than decoration. Selection can reinforce a strategy because it works often enough, while leaving a blind spot that only appears in rare states. The strategy can become more coherent at exactly the same time that it becomes more brittle.
Internal disagreement is part of the culture too
The Forge reaches an even tighter state in that same round:
It sees the same livestock decision. The Chief Optimizer recommends taking the Prosperity gain and sacrificing Health. That is exactly what its role is built to prioritize. The Systems Architect recommends the safer option because a colony at Health 5 has no resilience left.
This disagreement is more informative than unanimous role-play. The Forge contains two doctrines that usually cooperate: growth and robustness. At the hard boundary, they produce different answers.
That suggests a more sociological definition of culture. A culture can be a space of legitimate arguments plus an institution for resolving them, rather than one utility function copied across every member.
The Collective shows the same pattern from another direction. In one round, the agents must decide whether to release evidence from a deep-space probe that could destabilize the colony's foundational beliefs. The Guardian of Harmony supports disclosure because secrecy creates an insider-outsider split. The Steward of Health prefers classification because the information could produce stress and instability. Both are recognizably Collective arguments. They disagree on which interpretation of collective welfare should dominate.
Real institutions work like this all the time. Security, product, finance, legal, and reliability teams can share an organization without sharing a ranking over every action. What makes the organization predictable is not unanimity. It is knowing which arguments have authority in which states, which trade-offs need escalation, and which constraints function as vetoes.
The archive has agent roles, but the council itself has little constitutional machinery. A critic can correctly identify a catastrophic boundary and still lose the final aggregation. "Add a safety agent" is not enough if safety remains one paragraph among several. Some objections need to change the decision rule.
Competition between cultures therefore has at least two layers. One is which values generate good local decisions. The other is which institutions can resolve internal conflict without deleting the diversity that made the disagreement useful.
Rhetoric may be part of fitness, but the evidence is early
The paper deliberately tracks not only decisions but how agents justify them. Its sentiment analysis finds fairly similar polarity across tribes but more separation in subjectivity. The Vanguard is the most subjective in the reported run, while The Collective is the least. The paper notes that the most subjective tribe also performs best in its ranking.

That correlation is interesting, but it is easy to tell too strong a story about it. Subjective language may help agents explore a wider set of strategies. It may make persuasive coordination easier. It may simply be a linguistic by-product of the same prompts that encourage risk taking. One run cannot separate those explanations.
Still, the question itself is worth keeping. In a population of agents, rhetoric can change the social environment in which policies are selected. A policy that is hard to justify in the tribe's own language may struggle to propagate even if it is locally effective. A weaker action may survive because it fits the culture's existing frame.
That gives cultural cohesion a double role. Cohesion can make coordination cheaper, but it can also make a group resistant to useful mutations. If selection rewards both task performance and linguistic similarity, the fitness function is implicitly deciding how much conformity is worth.
This is one place where I would now push the experiment much harder. Instead of asking only whether rhetoric and policy drift together, I would intervene on one while holding the other fixed. Move a successful policy into a tribe whose language rejects it. Move the rhetoric without the policy. See whether either survives independently.
The paper's contact phase is the experiment I most want to finish properly
The two-phase design has a natural second question after divergence: what happens when cultures meet?
The paper proposes agent migration after the independent phase. That is more interesting than simply extending the run. A migrant gives you a way to distinguish several things that look identical when every tribe remains isolated.
Suppose a Vanguard agent moves into The Collective. Does it keep choosing Vanguard-like actions but learn Collective vocabulary? Does the host council move toward its policy? Does the migrant assimilate completely? Does one unusual agent increase exploration without changing the tribe's identity? If that migrant later becomes a high-fitness parent, which parts of its old culture are inherited?
Those are questions about cultural transmission rather than prompt obedience, and the archive does not reach a clean enough contact phase for me to answer them. That is also why I would not use the current project to claim that cultures genuinely transmitted across generations. The architecture is pointed at that problem, but the evidence is much stronger for persistence than for inheritance.
A better rerun would make that distinction explicit. Give each culture its founding constitution for a limited number of rounds, then remove it. Keep whatever the agents have actually inherited. Introduce migration only after the original prompts are gone. Now ask whether the culture is still recoverable from behavior.
If The Vanguard continues to value Future Outlook after the founding text disappears, something has persisted. If a migrant adopts Collective language but keeps Vanguard decision boundaries, rhetoric and policy have separated. If descendants retain both without seeing the original constitution, then the word "transmission" starts to earn its keep.
I would also run many seeds, add a neutral control culture, vary the base model, and keep the dilemma sequence paired across tribes. The current environment may favor certain philosophies, so the real object is not "which culture wins" in one world. It is which cultural traits remain useful across worlds.
What I think the experiment is really about now
When I first built this, the fun story was obvious: artificial tribes develop personalities, mutate, compete, and eventually exchange members. That story is easy to visualize and dangerously easy to overclaim. What has held up better is the path from values to path-dependent policy. The same initial preference can shape hundreds of later facts. It changes which action is taken, which changes the state, which changes the next trade-off, which changes who looks successful, which changes what gets selected and repeated. Over time, a founding bias can stop looking like a sentence in a prompt and start looking like the history of the society.
That is also why the failures matter as much as the winners. The Vanguard's reported performance says something about exploration in this environment. The Forge disagreement says something about internal institutions. The Collective shows how repeated protection of one metric can create pressure elsewhere. The Frontiers shows the sharpest failure: a culture can remain coherent all the way to a state where coherence itself becomes dangerous.
At Health 5, the Frontiers agents have the state in front of them. They know the next action costs five Health. They recommend it anyway because the action still fits the society's model of strength.
The values did not fail to persist. They persisted into a state where persistence itself became maladaptive. That is the experiment I want now: put different agent cultures under the same pressure, let them compete, migrate, and evolve, then ask what each culture has to give up in order to keep surviving.