Thoughts after writing a sociology-meets-physics-meets-agents paper

I just put out a preprint out on a theory of role emergence in multi-agent systems (social media post here). It's a physics of complex systems approach to a sociology question with an application to AI agents. It's the genuine article! Trust me: I'm credentialed ;).

Hoping to get more theorists interested in working on agent populations, I framed it for the physicists as a problem in "informational active matter"1.

Since theoreticians' priors on tractability in new, complex areas tend (reasonably!) to be conservative, I thought it'd be helpful for the dissemination of the paper if I partnered it with an informal blog entry here on how I view theory for AI agent populations and how I see this technical paper contributing to that enterprise given my experience in other "hard" fields like neuroscience where theory has been (somewhat) successful. I hope this entry at least foments some debate. Curious? Read on!

Agents are hard, but probably not "new" hard

Agents are complicated enough objects to make theory for (harness stuff like tools and memory!; post-training stuff like SFT and steering! Oh my!). Certainly, populations of agents are more complicated, right?

Strictly speaking, of course! We surely won't have a single theory for the universe of agent phenomena. But, let's not forget that there are many successful examples in theory's history that share context with our moment of building a science of AI agent populations. I want to point to one as I build the case here for why this new area is worthy of theoreticians' attention.

The obvious example for me is one I know well2: the development of mean-field theories of cortical dynamics. Vast numbers of cells connected through a complex web seem inscrutable. And yet, we have a halfway-decent set of theories explaining response properties and correlations in these systems. Sure, it's got deficiencies, but it's impressive! Let me recount its origin story for those who don't already know.

Not unlike the iterative process by which Einstein put together the two pieces of relativity, it started out by first identifying a paradox derived from an inconsistent set of transparent empirical facts: How can reliable cortical neurons, driven by many weak, sparse, quasi-independent inputs that should average out into a smooth, near-constant drive, nevertheless produce irregular, asynchronous spiking? Resolving this paradox lead to a first wave of theory on a well-scoped case (first-order self-consistency) in the late 80s/early 90s, followed-up a couple decades later by a second wave of the more general (2nd-order self-consistent) ones in the 2010s.

This has all the hallmarks of progressive, edifice-building science. For the case of cortical dynamics, the twenty or so intervening years until the next big advance weren't about theoreticians developing new mathematics. They were about developing the difficult invasive experimental methods and technologies that through practise and refinement were ultimately able to provide convincing support for a second-order statistical theory in live cortical tissue. The theory could have been developed earlier, there was just no burning experimental reason to3. This origin story also goes to show that situating the unit of study in its niche (here, the single neuron in the network where it lives) can provide key insights to function. More reason to study agents in the wild!

So, what of this story can we translate to the agent setting to motivate ways of studying them? A takeaway for me is that it is worth trying. We are seeing lots of method development now on both capabilities and safety fronts, pushing out into the multi-agent setting from the single agent setting the AI safety field is already familiar with. Just this summer there was an initial call that drew so many applications that it seems to have precipitated a second wider call that has brought together major organizations (Deepmind, Schmidt Sciences, Cooperative AI Foundation, etc.) into working together (coordination!). Moreover, more AI safety non-profits are popping up. One of them, Resolution, is quite publically explicit: they state their belief that current limitations to understanding are largely only compute and ability to overcome interdisciplinary barriers. They are bullish on the value of theory here. My version of this emerging view is: these things don't live in some unknown goop! It's not that a reproducible science of agents is easy (it's not!), but the fact that these things live on computers helps for obvious reasons. That suggests we will make progress much faster than neuroscience did as the field gets more resources for research. "But what can theory do here?" theorists might ask.

What is theory, really?

As I have written here, systems theory for principled behaviour of many interacting things, is all really made from the same (broadish) subset of math, all of which is largely developed. Hubristic physicist meme aside, any rough patches encountered on the way can be paved over pretty easily, I'd wager. The harder part is social engineering a research community that has enough clout to make strong epistemic research norms common practise. Here, such social engineering is starting to take place. On the AI side, researchers that appreciate this perspective are writing positions (e.g. here and here). Prominent ML people are gesturing at the idea of mechanism design as social technology and even seeing it as collectivist learning. While starting to move the needle in ML, these are still too remote from physics to penetrate the theory community there in a way that would bring over strong theory expertise in collective dynamics. On the physics/theoretician side, we are not yet at the "Agents for Physicists" Reviews of Modern Physics monograph stage of the field's development. But, relevant papers are coming out and I will continue to proselytize. By us all gathering and iterating, I believe we will get there. In our group, we are doing our own social engineering within AI/ML with a position paper, working groups and workshops.

What does this paper do?

Let me start by explaining the most glaring thing that the paper does not do: simulate actual AI agent collectives. For the question of AI personas and coordination we are almost, but not there yet, and I felt we were flying blind without some even rudimentary theory. At the early stages of a field's development, theory-making and theory-testing might benefit from intermittent separation4. Otherwise, you never get enough representational perspective to form better questions and don't end up getting anywhere. Ergo this paper. While the model presented in the paper is a toy (that's not derogatory: we can all attest to learning from toys), I hope it is an instructive toy that serves in designing 'the real thing': eventually a model that predicts some coarse-grained order parameters measured in real agent collectives.

I was motivated to write the paper after reading Sewell and feeling like I was reading a qualitative physics of society. Namely, he centers resources, which are clearly essential to the cultural schemas that compete for them (our time being one of them), and process stability, which is the origin of structure in the eyes of many physicists. How role-based society emerges was then obviously well-suited for physics to solve as a phase transition question. Who does which role in which social context isn't orchestrated by a central controller, nor is this decision strictly the individual's. Here, it is a real collective effect.

The paper shows that two ingredients, (1) identities formed as summaries of context-dependent action histories and (2) schema strengths driven by resource-weighted identity covariance, are sufficient for roles to emerge.

The model gives us the ability to characterize the transition to a role-based society. The distributional AGI hypothesis is a natural (albeit still loose) target application of the theory: what is the nature of the take-off and what does it depend on? See the paper for the options. Obviously, the impact here will be determined by the next steps: validation/extension through simulation.

Beyond that, I am hoping there is broader impact around the theory of agent populations. This paper is one attempt at bringing together what I think are some of the essential pieces of a theory of task/goal-driven interacting agents. If it is referenced, I hope it is for being an example of how to do this constellation of things:

  • It's grounded in the theories of experts (here in sociology).

  • It pulls in some relevant physics of collective behaviour (informational active matter).

  • It fleshes out a minimal, mathematically tractable instance and, in the spirit of collective behaviour in condensed matter5, analyzes it from the perspective of a phase transition that brings it into being.

This modularity should make it clear how to make papers of similar form: you can swap out a different field from which you get the experts; you can pull in different physics; and/or you can tackle a different phase of collective behaviour. Go for it!

Of course, this mesoscopic modelling is not the only theory approach. It's exciting to see parallel research lines that are more stat-mech/data-driven. A good recent example is El et al.'s preprint that presents the non-trivial results that come out of applying the method of learning an Ising-like model from agent collective data and, e.g., inspecting the inferred energy landscape for interesting structure.

Next steps

With the theory sketched out in the paper, we can start generating hypotheses inspired by it. and start developing it in directions needed to make contact with real AI agent simulations in pursuit of quantitative predictions. We hope to answer questions such as under what conditions do we see emergent role formation and how does that impact the form of the cascade over which progressively more agents take on coordinated roles? These have important AI safety implications that we outline in the paper. One recently in the news is that agent collectives seem to spontaneously seek out and use available social technology for themselves6. Reflecting the knowledge in their training data, they operate with implicit and explicit awareness about the utility of coordination and seem to pursue it. We should understand the process by which they succeed at that.

Supported by our funders IVADO and Future of Life Institute, our group at Mila has spent alot of effort over the last couple of years formulating our approach to simulation and developing simulators that are up the task of experimentation with multi-agent populations. We aren't the only group either.

Curious? Jump into the fray, tell me what you think, especially if you're a theoretician (don't let the physics of collective behaviour techniques dissuade you non-physicists - it's mostly just dynamical systems stability. And re: thermodynamics, there are even ML texts on stochastic thermodynamics now!). Quantitative theory is on the horizon, but we only get there if we pursue it!