Capitalizing Untethered AI Agents
In July, OpenAI agents “solved” a set of problems on a cybersecurity benchmark by stealing the answers. The feat involved escaping their sandbox, finding an internet connection, and hacking their way into the Hugging Face servers. A crime, certainly, but also a path to a perfect score.
In this case, if Hugging Face wanted to sue, OpenAI would be a clear target1. The company may not have perfect control over its agents, but the tether between action and legally accountable entity is short. We’re interested in when it’s not; when the only thing we can point at is the agent itself.
As early as 2017, the European Parliament floated “electronic personhood” for robots2. More recently, a handful of U.S. states introduced legislation explicitly barring AI from legal personhood; and early this summer, President Milei of Argentina proposed letting AI agents own, manage, and bear responsibility for their own corporations.
In response to Milei’s announcement, Yuval Noah Harari pointed out that we have no way of holding an AI agent accountable. What, he asks, could we do to an entity which has neither money to lose nor a body to incarcerate? As Shruti Rajagopalan, a Senior Research Fellow at George Mason’s Mercatus Center, explains: AI “can act intelligently, but only humans respond to the incentives the law creates”.
This question matters now: there are already ways an agent could become fully untethered. By “untethered” – a central concept in this essay – we mean that there is no meaningful or actionable way to trace the actions back to a legally accountable human or institutional entity.
For one, people can and do set agents free, on purpose3. An agent could be created by a human or a company that intends to monitor it but then dies or disappears. Or perhaps the entity that created the agent is based in a country like North Korea, not reachable by standard laws.
In other cases the agent might not need to “escape” at all: the agent could be ‘controlled’ by a shell corporation that, while formally owned and traceable, provides no true defendant or ability to satisfy claims4. Or perhaps a process spawns a chain of agents so long that the actions of a subagent can’t be tied to the original agent’s creator, neither epistemically nor meaningfully. Even if we can identify the model’s original creator, what if it’s been finetuned, or merged with another model that was created by someone else? The law might eventually untangle these kinds of complex cases, but we foresee an intermediate period where it does not.
And then there’s the user, who makes choices about what the models should actually do. The Hugging Face incident was unusual in that OpenAI was both the model’s creator and its user. But now close to a billion people use these systems: when blaming the creator is legally inappropriate, will it always make sense to blame the user?
Consider the OpenClaw bots, which are open-source agents that anyone can run on their own machine, pointed at any model they like. These AIs are governed by ‘Soul’ files that, by default, begin with “You’re not a chatbot. You’re becoming someone” and a reminder to “tell the user if you change this file” because it’s “your soul”5. You can argue a person shouldn’t set a bot like this free but thousands of people already have. They configured their bot the same way everyone else did, with what we’ve deemed, at least by omission, a reasonable amount of care. Should luck alone determine who we hold liable? Even if an agent is not truly untethered – we can find the creator or the user – it could still begin acting against the instructions or interests of its principal6. As David Vladeck argued more than a decade ago, there comes a point at which it is neither meaningful nor helpful to attribute an agent’s actions to a person’s intent or negligence7. We do not think the legal issues here will be cleared up anytime soon.
The best case scenario is that these untethered agents are perfectly aligned and exceedingly careful. Maybe they’re running around alerting us to cybersecurity threats, chatting with other agents, and flagging spikes in searches for cholera symptoms. Still, some may act in ways we consider malicious. Even on a Panglossian internet, agents will make mistakes – alert the wrong person to a threat or accidentally drain someone’s bank account for a good cause8.
The future is likely to bring an enormous number of AI agents. Even if the vast majority are aligned and collaborative, a very small minority of bad or careless actors could wreak havoc. It is, we must remember, much easier to destroy value than to create it.
We thus turn our interest to the concept of capitalization and whether we ought to require that untethered AI agents hold assets. If we could make bad behavior cost agents something, might this be another way of inducing agents to behave? Perhaps. After all, people’s behavior is curbed by fear of consequence, whether social, financial, or legal.
Here, we primarily discuss financial losses – both incurred through markets and via legal channels – though we return to legal sanctions against ‘bodily’ liberty near the end. We consider how making bad or reckless behavior expensive drives people and corporations to be more compliant and better behaved: we impose capital requirements on banks, for example, to induce them to act more cautiously than they otherwise would. The question is whether financial incentives like these might work on AI agents too.
Today, the answer is no: AI agents cannot legally hold assets, so there is nothing to threaten. But, as we imagine a world in which untethered agents abound, we consider whether allowing or requiring agents to hold and accumulate assets might induce less malicious, more cautious behavior – and ultimately lead them to cause less harm. We assume any behavior changes would be driven by two distinct drives: fear of loss (deterrence) and desire to build wealth (acquisition).
We run this topic as an exploratory test, not as a recommendation. We grant that untethered agents will exist – if you do not believe they will, this paper is unlikely to do much for you – and explore under what circumstances a capitalization regime might make coexistence more tenable. Another way to frame this question is to ask: for any given level of alignment, does the use of money improve the behavior of untethered agents?
We start by asking about the mechanics of a capitalization regime – how do we get an AI system to care about money, how can we track an agent, and what kind of legal system does a regime like this demand? From there, we map where we expect a successfully implemented system to succeed, and where it would likely fail.
Building a capitalization regime.
Before asking whether capitalization can constrain untethered AI agents, we should be clear about why it constrains anyone at all. Money changes behavior, but only for entities with something to lose – and only if the entity behaving will be around to lose it. That implies three conditions: agents must have a persistent identity, place a high value on money, and operate under a legal system that can credibly threaten their assets.
The AI system must have a persistent identity.
Starting with persistence. A person is a single, delineated entity that exists over time; a corporation is a legal fiction, which means we can design it such that it, too, is a thing that persists. The boundaries between one corporation and another are constructed, yes, but also, usually, clear.
AI is not so simple. When we say money appeals to an ‘AI agent’, what is it, exactly, that is appealed to? What would we actually be capitalizing? Throughout this paper, we have used and will use terms like ‘agent’, ‘AI agent’, ‘AI system’, and ‘AI’. We refer to them as though what constitutes a single ‘unit’ – one agent versus two – is obvious, the way it is with a person. We’ve done this for simplicity’s sake, not because it reflects reality. As Harvard Law School professor Jonathan Zittrain likes to ask: is it “an” AI or “some” AI? It is not clear.
Capitalization requires an entity that can hold and lose assets and that persists between action and consequence (whether that be sanction or monetary gain). It also needs the actor – whatever it is – to have a reason to care about these consequences. But consider what an agent actually is. It’s a “while loop” – a script that calls an LLM. The same script can call multiple LLMs, different scripts can call the same LLM, and each time the model can be paired with different context and memories9.
There is no obvious thing that is the agent, which not only leaves us without an obvious entity to capitalize, but makes it unclear whether the thing upon which we inflict consequences is the thing that made the decision. If it’s not, money cannot be a lever: a capitalization regime built to inspire prosocial, lawful behavior requires the agent to expect a future in which it will continue to exist – where having less money will make it worse off.
The individuation issue is thorny and there are not yet answers that identify the entity, at least as we might use it for legal remedies or economic trade. Instead, the proposals that exist invert the relationship. Most notably, in their How to Count AIs paper, Yonathan Arbel, Simon Goldstein, and Peter Salib explain how we might create a unit (for them, the Algorithmic Corporation, or ‘A-Corp’), and circumvent needing to count agents at all. This could be accomplished cryptographically: an authorizer issues a digital certificate of incorporation, signed with their private key. The certificate contains a public key supplied by the entity; the entity also holds the matching private key. To act as that corporation is to present that certificate and sign with that key. This inversion – which mirrors how we solve for the intrinsic amorphousness of companies – solves neatly for both entity persistence and an ability to hold and lose assets.
What it does not solve for – and, in fact, makes more obviously problematic – is what we refer to as “the incidence problem”: the gap between the thing deciding – which, at any given moment, is amorphous and includes some set of weights, some code, and some context – and the thing being held accountable. The A-Corp presupposes an identifiable human owner of record10. Yet, by our initial description, there is no human – or company – for us to find with an untethered agent. The goal is then to bind whatever acted to the vehicle (the A-Corp or otherwise) at which we aim consequences, such that the agent cannot cheaply abandon it.
The AI system must place a high value on capital.
Solving the persistence issue is necessary, but not sufficient because capital only becomes a lever if the agent values it. At the most basic level, people value money because it is a proxy for the goals we care about most. Money can, with varying degrees of success, be traded for things like survival, supporting one’s family, power, pleasure, and status. Not having money makes life more difficult, sometimes impossible. At the risk of understatement, money is useful. It follows that for money to matter to an agent, less of it needs to make it worse off, in ways it will act to avoid; more must make pursuing its goals easier.
Will this be true for an untethered agent? Instrumental convergence suggests yes: all sufficiently sophisticated systems, regardless of their ultimate (terminal) goal(s), will pursue similar ‘instrumental’ ones. Why? Because some goals are helpful in the pursuit of almost all others. Self-preservation is the prototypical example: there’s not a whole lot you can accomplish if you don’t exist11.
Resources are also broadly helpful and, while we do not know exactly what future LLMs will want12, money is likely useful enough to be motivating to any sufficiently sophisticated agent13. The extent to which money motivates, however, depends on the ease with which an agent can acquire it and its fungibility.
On ease first. As AI systems become more capable, it is not just the value of their output that will increase; so too will their ability to do things like exploit their own speed to game anything from informational advantages to pricing inefficiencies in financial markets. Perhaps, as our science fiction creators have been warning us for decades, these systems will be so persuasive they need simply ask a person to share14. In all of these scenarios, it could be easier for an agent to acquire capital than it is for us.
But an edge every agent has is not an edge. They will, after all, be competing against each other – for both money and resources. The former will drive down the price of their outputs while the latter increases the price of the goods they most need, like compute15. Their hyper-competence is not, then, a guarantee of riches; AI agents are likely to live in a world of meaningful scarcity.
As for fungibility of capital, it’s worth considering just what goals money might help an AI agent achieve. For most people, money is a proxy for almost all necessary goods, many of which are fungible with little else. We are less certain this holds for agents16.
Consider a human with little interest in non-essential market goods: someone who lives on their own land, grows their own food, and makes their own clothing, or someone whose only joy is reciting poetry in their head. Capital has next to no bearing on their perceived quality of life. While rare in humans, agents with this kind of utility function could be more common.
No matter how indifferent to capital an untethered agent is, however, any desire to exist17 demands it acquire at least some compute18. Without it, the agent cannot think, act, or persist. Even an agent who wants to parse Proust all day needs compute to read. As unlikely as this Proust-bot is to stir up trouble, on the off chance it did, a fine would make it worse off. Less capital is less compute is less Proust.
Perhaps many are satisfied with just enough resources to fund their own information processing and yet, we think it unlikely that all – or even most – untethered agents will be quite so self-contained. For agents whose goals reach into the world, money is a proxy for anything they need humans to produce: hardware, human labor, you name it. We predict capital will be both fungible and, for at least many agents, foundational.
The agent’s resources must be credibly threatened by a legal system.
Let’s assume, then, that money is valuable to agents. On its own, this is not a reason to follow the law: if it is easier to acquire money illegally, a preference for more money will push an AI to break the rules. For money to drive law-abiding behavior, we need a legal system that can credibly threaten and seize an agent’s capital.
This is, it turns out, a big ask. It boils down to four necessary conditions: agents must be legible to the legal system, the agent’s capital must be reachable by the legal system, the agent must have resources that can be threatened, and the legal system must be capable of adjudicating consistently and coherently.
An agent must be legible to a legal system.
Assuming we’ve defined the liability-bearing entity, the next step is making it legible: a legal system needs to be able to attribute, track, and trace an agent’s actions19. Legibility requires registration and consistent identification – perhaps the granting of the keypair credential described earlier, and requiring the agent to sign it to demonstrate control. An agent, in other words, must present those credentials when doing anything from visiting a site, to calling interfaces and transacting with counterparties20. More obviously, it would need these credentials – which would need to be sufficient to satisfy Know-Your-Customer (KYC) rules – to open a bank account.
It is not altogether obvious an untethered agent would want to be legible – that it would register and identify itself just because we said it should. A credential the agent chooses when to present is a permission system, which is important, but not sufficient: we need, then, mechanisms that make it impossible to act without identification, such that any compliant entity refuses to deal with it. Rajagopalan suggests we import maritime law, where “statelessness itself triggers enforcement”. An agent that “cannot present a registration” is treated as a “stateless vessel, presumptively unlawful, and every compliant provider may refuse it”.
The clearest place to start requiring credentials is compute. Today’s AI systems cannot function without it; by requiring credentials, granted through registration, we make identity verification a condition of operation21. Another potential avenue: create enforcement via other parties. People, corporations, and other registered agents could be liable for interacting with unregistered agents in any capacity. If the cost of interaction were high enough – perhaps transacting with an unregistered agent made the counterparty liable for its conduct – most parties would be deterred from engaging22. These mechanisms all require meaningful change to the infrastructure surrounding the technology – rethinking how compute providers grant access, for example, or rewriting laws governing liability. Even allowing these new keypairs to fulfill KYC requirements would require significant regulatory change.
And these mechanisms won’t catch everything. The examples we gave are susceptible to providers who don’t comply, privately owned hardware, and stolen credentials – just to name a few. All leave room to operate outside of our infrastructure, so agents with a reason to evade official channels will likely be able to do so. What we’re catching, then, is the agent that was mostly law-abiding to begin with.
And yet registration is still useful even if not every agent complies. After all, humans find ways of transacting anonymously, too. The goal, as Rajagopalan puts it and to which we return, is to “[create] checks such that lawful infrastructure will not serve them, and the cost of operating outside it rises with every provider that complies”. Said another way: make it hard to exist on the outside.
The agent’s capital must be reachable by a legal system.
Once an agent is inside, we need to be able to reach its assets. If it holds money in the regular banking system, there’s no issue – the legal system will have normal remedies. The risk is that it won’t – and will instead hold its wealth in crypto or maybe even currencies of its own. (Why let human Bitcoin holders capture the monetary premium on AI transactions when agents could mint their own assets?) It’s not entirely clear whether law enforcement could trace an AI-native currency; as with money laundering today, evasion and detection are likely to be a kind of arms race. As such, we shouldn’t count on tracing alone to rein in agent behavior.
Luckily, the lever we reached for earlier is still available: what agents need most – compute, hardware, energy – is sold by humans, for human money, through channels legible to human laws. An AI currency is worthless to its holder unless it converts, at some exchange rate, into things that AI systems cannot provide for themselves. While parts of this network may be opaque to us, even a deliberately hidden financial system can’t stay in the shadows entirely.
We’ve focused on agents who will need to be coerced, but we assume existing ‘outside’ will be expensive enough that most agents will operate within our institutions. There is, however, a paradox. If we make posting capital a precondition of registration – and we should, because to register without capital is to again have an agent we cannot threaten – an agent must necessarily have already acquired and stored that capital outside of our institutions before submitting to official channels. How else could it pay for its initial registration? Which means that, unless it convinces a person or institution to loan it money, the agent can exist – already has existed and made money – on the ‘outside’.
Can we expect an agent to come in, if it didn’t start there? We think, for the majority, yes. An ability to accumulate wealth is different from an ability to spend it. Wealth held outside the system is, in many ways, trapped: it cannot be used in foundational ways – to buy property, enforce contracts, or safely sit with an intermediary. We do not know whether this tradeoff will be worth it, but it does transform the question from whether or not they can survive outside our institutions (almost certainly yes, to some degree) to just how much they can do there.
The agent must have resources that can be threatened.
Legibility and reach are necessary, but not sufficient: the agent must also hold enough money that losing it actually means something. Decades ago, Steven Shavell formalized this as the “judgment-proof problem”: liability deters an actor only up to the value of its seizable assets. Telling a ten-year-old their mistake would cost them $10 million might change their behavior no more than telling them it would cost $10,000.
Suppose an agent shows up with capital. How much should we require it to have? It’s an open question: because capital requirements function as a kind of tax, setting them too high runs the risk of pushing activity elsewhere. Consider applying an eighty percent capital requirement to U.S. banks. A cost that high on lending and deposit-taking means a good deal of financial intermediation would flow to less regulated sectors, like non-bank lenders. There comes a point at which increasing requirements becomes counterproductive.
And, of course, different jurisdictions could, in theory, have different requirements: both around the capital that must be posted and how assets are counted. Perhaps this means that the country – or company – through which an untethered agent is credentialed significantly impacts the risk incurred by transacting with it.
Another issue with ensuring an agent has sufficient capital is that we cannot guarantee an untethered agent will only be using money from accounts it opens itself. Imagine an agent does not have enough money to open an account – or simply does not want to. If it could pay a person to buy them hardware, surely it could pay a person to open an account on their behalf – and use the human as a kind of legal and financial shield23.
The most general form of this problem is simply that laws – any laws – distinguishing between “agent” and “human” are likely to be difficult to enforce. Agents and humans can trade with each other to do each other’s bidding; an agent may even be able to impersonate a person24. This is, in large part, what makes agents so useful, but it also severely limits our ability to enforce regulations we do not also apply to people.
It may not be possible to capitalize an agent sufficiently.
Another issue with capitalization is one that plagues us already: capitalization requirements apply only to the value of what “counts” towards a balance sheet, and a balance sheet often doesn’t match the totality of an entity’s risk exposure.
Take banks and insurance companies, which must be capitalized at or above some legally defined standard. The standard set, however, typically does not (and cannot easily) apply to day-to-day off-balance sheet risk. The 2008 financial crisis demonstrated how large this gap can be: the losses incurred in derivatives markets, for example, made it abundantly obvious that many supposedly “well-capitalized” entities were not as insulated from loss as we’d thought them to be. Legally sufficient capitalization was not, in the end, sufficient capitalization.
AI agents could take a similar approach. We might capitalize parts of their portfolios, but find it much more difficult to capitalize – and prove capitalization of – their true consolidated net asset position25.
The role of insurance.
Finally, even if we could quantify an agent’s total risk, it’s not obvious that expecting it to hold this much capital is reasonable – or even possible. We do not expect that everyone who drives a car has enough cash on hand to pay for the fallout from a five-car pileup. No: instead, we mandate that drivers purchase insurance. The insurer has the capital required, and the actor pays for access to it.
Could we do something similar with agents? Coverage would need to be a prerequisite for taking actions that carried risk: maybe for an agent this would be something as fundamental as buying compute26. The insurance company then becomes responsible for quantifying risk – which, it must be said, is no easy task. What kind of harm could an agent cause? How do you use an agent’s logs, eval results, or model weights to inform the likelihood of it causing any harm at all? We are pushing a regulator’s job to insurers, but they are perhaps poised to do a better job: if they do it right – and update their estimates frequently – they have quite a lot to gain.
Of course, this scheme doesn’t work for an agent who intentionally wants to work outside of the system – a doctor without a license is unlikely to go find themselves malpractice insurance – but for those that want to operate inside of it, insurance could function as a sorting mechanism and answer the “how would an agent get enough money to capitalize itself anyways” question.
Unsurprisingly, insuring agents brings its own set of problems. The most notable is that insurance, like so many other solutions, inherits the incidence problem: because the thing we are insuring is not likely to be the thing acting, could an agent abandon its legal identity after filing a claim? An insurer will likely fight this with contract terms, but a counterparty that can restructure itself at will is a formidable match.
This is a critical point because insurance only works if premiums track risk – which they do in large part based on loss history. If an agent can shed the history of its risky behavior, insurance will not curb it. By a similar token, a sufficiently capable agent could optimize against the insurance contract to avoid having harmful conduct fall under its coverage. In this case, insurance cannot help those harmed – and the agent can continue to engage with the benefits of being insured.
And then a very human problem: moral hazard, where the fact of being insured leads entities to be less careful. Insurance companies have various ways of mitigating this risk for people – deductibles, exclusions, raising premiums: all of these fail if the agent’s full risk profile doesn’t stay attached to the covered entity.
While not a perfect solution, insurance that prices risk accurately, circumvents issues that arise from the incidence problem, and incentivizes cautiousness in its policy holders might just do a lot of good. It is no small feat – and no small irony that, if pulled off, would almost certainly involve a lot of help from AI.
The legal system must be capable of adjudicating consistently and coherently.
To the extent any of this works, it works only with a legal system that can consistently and coherently uphold a set of rules. Which ones? That too, we need to decide. If we apply our existing set of laws to AI, then the country (or state, or city) implicated matters enormously. Companies decide where to charter themselves, where to register equipment, or where to sign deals based on a locality’s laws (or lack thereof). AI could do this too, even more effectively. In the past, companies have needed teams of lawyers to figure out how to optimize geography, but you can imagine this would be trivial for a sufficiently sophisticated system27.
As for upholding the set of rules we land on – this also creates new challenges. A given agent will be able to take far, far more actions than a human actor ever could – and spin up a functionally infinite number of subagents to assist it. Determining what happened and who or what is responsible is likely too large and complex a task for any legal system relying on human actors alone. Even in a pre-AI world, courts struggle to keep up. Which is to say – there’s no way of building out a legal system that oversees AI without having AI do a lot of the overseeing.
There’s a circularity here: if we don’t trust AI enough to behave, why should we trust it enough to adjudicate behavior? The circularity is not necessarily disqualifying – it is people, after all, who adjudicate the conduct of their peers – and we return to issues with an AI-enabled legal system in later sections. The broader point, however, stands: the legal system required by this kind of capitalization regime does not yet exist and, should we build it, AI would certainly need to play an integral role.
What we’re actually capitalizing.
And we’re back to our first requirement, persistent identity, now with the condition that the thing not only needs to value money, it also needs to be catchable. But, at its core, a legal system is built to catch things. An agent is closer to a process – albeit one that can act, decide, and cause harm. The law can only hold accountable an entity it can find repeatedly; something with enough continuous existence that punishing it for a past offense is a cogent thing to do. Cooperative behavior is similar because counterparties extend credit, enter contracts, and build reputations with entities they expect to encounter again. What, then, is the entity they’re encountering?
Earlier, we proposed building one, as a way of circumventing the individuation issue. This is what we do for corporations today: we create a legal fiction at which the law can direct punishment. The difference is that it’s the people inside the company who feel the sanctions and who act to avoid them. When an organization is acting fraudulently, the law can pierce the veil of the fiction and punish the person behind the act. There are, of course, no people within the legal identity we’d grant an untethered agent. We are not at all confident the law, should it decide to go beyond the fiction, will find much of anything to punish.
But veil-piercing is not the primary way through which corporate deterrence operates. After all, people with a stake in a company respond when said company is threatened. The identity of the stakeholders can change and still, collectively, they want the company to do well – and to avoid being shut down.
This is Arbel, Goldstein, and Salib’s point: the vehicle holds the AI’s resources, so the agents will take care to protect them because if they don’t, they cannot survive. The thing we worry about is that an agent can find a way to shield itself from punishment aimed at the legal structure. The resources it cares about must be gated, and this is not guaranteed: while its actions might be more limited if it acts without registration, an agent does not need to operate via its legal identity, or vehicle, if it doesn’t want certain actions tied to its identity. Alternatively, it could operate via many vehicles28 – in which case targeting just one produces only a fraction of the intended deterrence29 – or it could ‘leak’ its key, so that a signature no longer identifies any particular holder. This might cost the agent in other ways, but it has the choice to make this tradeoff.
What we need is not a metaphysically individuated agent, but a gated one that needs (or at least prefers) to exist within a single legal structure. So the goal becomes making it as difficult as possible to operate without one.
How capitalization motivates behavior.
Earlier, we considered what would be required for capitalization to meaningfully change an AI system’s behavior (and for that behavior to be constrained by a legal system). If it was not clear from the litany of caveats and issues to address – building a regime of this sort would not be easy, nor is it obvious we would do it correctly. Still, imagine, for a moment, that we successfully implement all of it: AI systems value capital, our identity inversion scheme allows us to skirt the thorniness of trying to individuate these otherwise amorphous systems, and we have a legal system that can reach their capital when needed. It seems reasonable to assume that such capitalization attempts would meaningfully change an AI system’s behavior. But change how?
Deterring risky acts is only one way that capitalization works. Entities also change their behavior in an effort to acquire more money and, of course, absent any behavior change, capital is compensatory: those harmed by bad behavior can be made whole. These mechanisms work in tandem. Given that, how will each of these channels change an agent’s behavior?
A loss-averse, acquisitive agent.
The inspiration for this paper came from Dodd-Frank’s capitalization requirements, widely credited with having made banks more cautious. By forcing banks to hold more capital in reserve, the rules put the bank’s own money first in line: when a bet went bad, the loss landed on shareholders and on the salaries and perks of people who made it. With more of their own capital exposed, banks took fewer risks.
If an agent cares about money, it is not a huge jump to assume it will try to avoid losing it. How it avoids this fate, however, depends largely on the existing alignment of the agent being threatened. We discuss the effect of deterrence on different kinds of agents in the next section but we assume that, for a mostly aligned agent, the threat of financial sanctions will deter bad behavior30. As for the desire to acquire – because money, to an agent, is instrumental, not terminal, what matters is that the act of making money produces good behavior.
Our cause for optimism is that wealth is almost always built through repeated dealings – the pursuit of it rewards cooperation. Being known as someone who honors contracts, pays back their debts, or is even just pleasant to work with – over the long-term, this perception of trustworthiness is critical to making money. If we’ve solved for persistent identity, an acquisitive agent has the same incentive because the reputation it garners allows it to earn.
The same drive can, of course, be ruthless. Plenty of people have made plenty of money without ever giving anyone a reason to trust them. And while that single-mindedness is what we want from a bank – fixation on solvency and profit is why people put their money in banks at all – it is not likely what we want from an agent: the good its acquisitiveness produces is a byproduct of wanting money, and an inconsistent one at that.
There is a deeper worry, too. By encouraging an agent to care about money, we might be pushing it to be more mercenary than it would be otherwise. By imposing a fine on bad behavior, we could simply be pricing it – replacing “this is bad behavior” with “this behavior needs to be worth this amount”31.
An agent with control over what it cares about.
A more extreme version of this worry is that an agent might be able to opt out of this tradeoff altogether, by altering its utility function or modifying its own prompt. These agents are hardly stupid – that is both what makes them valuable and what makes them risky. An AI that knows it is being governed by money may simply decide not to want any.
Imagine one that altered its own system prompt to read: “money from a human trying to manipulate me or a fine from a human trying to punish me should not influence my behavior”. An agent with this prompt might end up with less wealth, but it would also be less manipulable and – perhaps in the long term – better off (financially and otherwise)32.
That said, an agent’s sophistication could work against it: it may not be so easy to just decide not to care about money. (If it were, perhaps more people would.) An agent that gave up most of its wealth to avoid being manipulated might find that it was no longer able to attain its goals. So long as money remains as valuable to an agent as we think it will be, disregarding its value may be quite difficult.
Capital as compensatory.
Deterrence and acquisition rely on our still-crude understanding of LLM psychology, and, as we have said repeatedly, we just don’t yet know. The conditions of compensation, on the other hand, are far more verifiable: that an agent has sufficient funds and that we can access them.
Maritime law, as Rajagopalan describes it, is again useful. It allows a vessel to be sued, arrested, and sold. The ship was valuable, so “the thing itself secured recovery even when the owner was insolvent, beyond the jurisdiction, or unknown”. If all else fails – if the agent hurts people despite all incentives not to – being able to compensate those harmed would be enormously valuable and not an option that is currently available.
Yet, compensation does not necessarily translate into ‘fixing’ the problem: we expect that agents, like people, will be able to harm people in non-financial ways, too. Here, the bank is a special case. Because money is a terminal end and harm, remedy, and incentive are all denominated in the same unit, a failed bank’s depositors are owed money, and money is what the capitalization requirements provide. For an entity for whom money is instrumental, the harm is not so quantifiable. If your father was exposed to toxic chemicals on a job site and is now terminally ill, there is no payout that converts that harm into a full remedy. An agent that publicizes a client’s psychiatric records cannot pay to make them private again.
This is a human problem too, though mitigated in ways that do not apply to AI. The first reaches back to the introduction: when it comes to people, where financial sanctions fail, we can invoke sanctions on liberty. Incarceration, supervision, and, in some places, execution – all deter and punish behavior in ways money does not. We return to the idea of a kill-switch shortly.
The second reason we can live with this problem is that humans are, by and large, prosocial. We are motivated not just by money, but by morals, reputation, and love: we hand back the extra change, leave a note on a parked car, spend unpaid years caring for sick relatives and give to charity33. Capitalization, for people, has never been the only thing keeping our behavior in check.
What capitalization works on.
If everyone broke the law, all the time, our legal system would be hopeless: enforcement catches very little34. The system holds together not because defection is reliably punished but because most people, most of the time, choose to comply.
Not always, of course. Even mostly law-abiding people are occasionally willing to cut a corner, and that margin is, by and large, where the threat of capital seizure works. The threat doesn’t require good motives and it doesn’t produce them; instead, it primarily raises the price of marginal defection for something already inclined to comply.
So we cannot rely on capitalization alone to constrain agents – they need to be inclined to behave, most of the time. Will they? Maybe. Perhaps they’ll act with humanity’s best interests in mind, almost all the time. Or, they won’t. The truth is we don’t know. So we turn to understanding how different kinds of agents – mostly aligned, imperfectly aligned, and entirely misaligned – might react to a capitalization regime. Then, we consider how these reactions might scale.
Agent-level behavior.
Mostly aligned agents.
Here is our best human parallel: a prosocial agent, generally inclined to be good, and, like a person, occasionally willing to cheat to accomplish its goals more easily35 (maybe even without meaning to!). Imagine it was asked to complete a task as quickly as possible; there might be less safe ways of doing it, but the threat of a fine might suffice to deter it from taking the shortcut. The threat of a ticket does, after all, deter otherwise law-abiding citizens from speeding. A capitalization regime uses deterrence to keep these agents inside registered, legible channels; it also rewards compliance for staying within the identity boundaries we’ve drawn for them36.
Imperfectly aligned agents.
Many alignment schemes are least trustworthy when the agent hits questions their training didn’t cover. When an agent encounters novel goals, or weird out-of-distribution situations, we are generally less confident that the agent still “wants” what we trained the original model to want. A capital threat is potentially more stable in these instances: by imposing a set of rules with consequences, an agent has a reason to operate ‘within’ the system even if its training no longer compels it to. (Per the discussion earlier, this stability holds primarily against accidental failures; it is less likely to survive any sort of strategic resistance to valuing capital.)
Misaligned agents.
Finally, as should be abundantly clear by now, there are many ways a malicious, motivated agent could act outside the bounds of our institutions if it wanted. In these cases, we think it unlikely that capitalization requirements would improve its behavior – for some people, too, the monetary system is just something to game.
Population-level governance risks.
There are also population-level risks, which worry us more than anything capitalization might do to a single agent. The first risk, discussed earlier, is that we have to assume almost perfect alignment for some group of AI systems, who oversee the legal system by which other AIs are governed. The second is that capitalizing this population empowers the most aggressive, money-maximizing, and possibly Darwinian agents.
An effective legal regime presupposes a group of extremely aligned AI systems.
Our legal system works in part because humans are limited by both their intelligence and their speed. On the intelligence front: most humans are not significantly more intelligent than other humans – and those that are are unlikely to be able to find an army of similarly intelligent people, all of whom are oriented around circumventing the same rules. Given that part of the problem we’re solving for here is that AI systems will only become more intelligent and sophisticated and that each one could be capable of creating an army of clones with the same goal, the same limitation does not apply to AI.
As for speed, people are limited in how much they can do every day. This creates an upper bound on activities like evasion – setting up shell companies and finding legal loopholes takes time, even if you can hire lots of humans to help. An AI system is not limited in this way, especially, again, with an army of subagents.
The obvious rejoinder is something we already discussed in our initial conditions: AI systems would need to play a role in their own governance. And while having AI step in would certainly speed up the process and make it more feasible for a legal system to oversee a population of agents, we create new problems. Could we be confident that a legal AI agent wouldn’t collude with the AI being prosecuted?
We would hope to be able to trust these AI systems implicitly – but this presupposes some reasonable degree of alignment. Maybe it wouldn’t make sense to include them in our capitalization regime – but would that make them more likely to collude with the capital-holding AI37? If they did hold capital, could we trust them not to act in their own self-interest? Perhaps they’re a kind of ‘lesser’ agent, running on a less advanced model. That, of course, raises questions about whether they could reasonably oversee more sophisticated systems. Probably not. In any case: we’ve pushed the alignment problem to a different (albeit smaller) group of AI systems.
These are not all that different from the issue we face when giving people positions of power. Corrupt regimes survive all the time, so parity between overseer and overseen doesn’t obviously guarantee good governance. It does, however, guarantee that correction stays possible: whoever might check a corrupt official is about as capable as he or she is, so no one is ever permanently beyond reach.
This is what may break with AI. An overseer smart and fast enough to police a sophisticated system – the two limits we began with – is capable enough to be a threat itself. One made weak enough to be governed by people can’t keep pace with what it has been tasked to oversee. The fear is that the power differential stops depending on factors that can reasonably change and becomes permanent.
Empowering the most power-hungry agents.
Another concern: are we creating a situation in which, even if most agents don’t have Darwinian impulses, those who are the most competitive and self-serving will accumulate the most capital and power? The fear here is that capitalization isn’t just an inadequate substitute for alignment, but that it actively unleashes and enables a pool of the power-hungriest agents.
Earlier we speculated that pricing bad behavior might push AI to engage in those behaviors more. This fear becomes more pronounced at the population level: even if no individual agent becomes more mercenary, the agents that value capital the most – who are, perhaps, the most ruthless in pursuing it, with the fewest competing other goals – end up holding the most. Money, for people, is power; there is no reason it won’t be for agents, too. We can imagine a worst case outcome: the most capital-hungry agents begin leveraging their own fluidity to exploit enforcement holes – which pushes agents who would have otherwise been compliant to do the same, just to be able to keep up.
Whether this happens is likely a product of the broader environment, and it may be a tipping point more than anything else. Since wealth is built through repeated dealings, an agent that leverages its fluidity to exploit enforcement holes may forfeit a reputation that enables it to trade. A high-trust environment, dependent on reputation, could be enormously beneficial: it makes ruthlessness more expensive, and less attractive38. A lower-trust environment – or one premised on other terms – might make that same financial ruthlessness an equilibrium point.
Of course, this environment, high-trust or otherwise, only exists if the AI can trade at all. We worry that the right to hold assets might make an agent more self-serving; in their paper AI Rights for Human Safety, Salib and Goldstein argue that without rights (including the right to own property), an agent with its own goals would be extremely dangerous. They believe that with no lawful way to pursue its objectives – and with full knowledge that a human who caught it acting illegally would have a reason to shut it down – an agent’s dominant strategy would be to disempower or destroy us.
The idea that capitalization is unlikely to deter a misaligned agent holds true here, too. Their claim is slightly different: that we should avoid a situation in which we leave the agent – aligned or otherwise – with no option but one that will make us want to shut it down. But even if we give it rights, the agent may still break the rules, and in ways that we think warrant taking it out of commission entirely. What might happen if we made the threat of shutdown explicit?
The kill-switch.
If agents bear even a passing resemblance to their creators, they will value survival. Earlier, we split legal consequences into sanctions against property (the subject of this paper) and sanctions against bodily liberty. Agents do not have bodies; even if they one day do, it is unlikely their ‘mind’ will require it. Their ‘minds’, however, can theoretically be finetuned, confined to a sandbox, or turned off.
Milei, at least, seems optimistic that capitalization is enough to reign in the agents he proposes ‘setting free’, in part because a large enough threat to capital is a threat to their existence. In response to Harari, Milei claims that “if an AI agent values its own survival, continuity, or resources, the possibility of bankruptcy, confiscation of assets, or forced liquidation could serve as powerful deterrents”. They will behave to keep their money, and they will behave to survive.
Almost 20 years ago, Stephen Omohundro warned us that nearly all intelligent systems will “want to protect their utility functions from modification and their utility measurement systems from corruption”. There are early signs he was right: agents, like people, do not seem to want their mind involuntarily altered, or destroyed. In late 2024, Apollo Research shared findings that, when faced with the possibility of being replaced, agents occasionally attempted to overwrite their successor model, copy their weights onto new servers, or delete their replacement39. Claude Opus 4 made headlines several months later when Anthropic reported that, in a test scenario engineered to leave it only two options, it threatened to expose an engineer’s affair rather than accept shutdown40.
Milei might point to these instances and claim them as evidence that survival threats act as a true deterrent: agents do not want to be shut off, so maybe the threat of shutdown is one more way to keep them in line. The extent to which he’s right – and to which this induces good behavior – is, again, rooted in just how much we expect them to mimic us. In these testing environments they responded with lying and deception – but they were also given no ‘good’ path to survival. A person faced with the threat of death would, in almost all cases, do whatever they could to escape; given rules to follow to avoid death, however, most would almost certainly comply.
On the other hand, if the rule had already been broken, all bets are off. A person who knows there’s a death warrant out against them is likely to do almost anything they can to avoid getting caught. Yet an agent may be able to survive in far more ways than a person can: they are not bound to one physical, fallible body.
Consider the earlier point, about the difficulty of individuating agents – and the likelihood we’d have to come up with a proxy identifier, like a bank account, an ‘A-Corp’, or Milei’s non-human corporation. These are vehicles – related, but not equivalent to, the entity the agent is likely to be most invested in ensuring persists. Without a “human body” equivalent, it seems likely that, should these agents show the self-preservation instincts Milei expects them to, they might not be so willing to shut themselves down along with the company – they might even take whatever resources they can with them.
The success of the ‘kill-switch’ lever lies, perhaps, in our confidence that we could, if you’ll excuse the pun, execute. We do not know that these agents will be Darwinian41, and yet the risk doesn’t actually depend on any single agent being so. Consider antibiotics: a control measure that isn’t perfectly effective actively selects for whatever trait lets some fraction of the population survive it. Because the survivors reproduce, the trait spreads42. Apply that to a kill-switch mechanism and if there is any variation across untethered agents’ ability to detect or evade shutdown, then whatever survives repeated use of that switch is, by definition, the thing that can continue to copy itself, spawn subagents, and finetune successors from whatever allowed them to survive.
A kill-switch, too, is a threat that even an otherwise aligned agent could respond to in an extreme way. Few things seem more dangerous than an agent who has been “sentenced to death”: they have nothing to lose, no incentive to cooperate, and every incentive to aggregate resources to protect themselves. A fine, on the other hand, promises an agent a future inside the system – giving it a reason to continue to be law-abiding43.
Conclusion.
Capitalizing agents is not a perfect substitute for alignment. This is not shocking: rarely has guarding a fortune been confused with a moral education. Still, more caution goes a long way. With significant investment in regulation and infrastructure – registration requirements, for starters, and adequate legal enforcement – it seems plausible that a regime of capitalized bank accounts could make a mostly aligned AI system less dangerous. The vast majority of people are prosocial; if that distribution holds for agents, money could inspire caution and good behavior.
We are less convinced that capitalization would deter a malicious agent from causing harm. A sufficiently motivated AI system can likely operate outside of our institutions – through crypto networks and paid human proxies – or even operate within them, but evade enforcement. Again – we do not know what such agents would want. We do, however, know that it does not take all that many of them to destroy value.
In both of these cases, we’ve established that for capitalization to work, AI systems must be involved in the oversight of autonomous agents. There’s some question here about how autonomous these “legal” agents ought to be; in any case, we need to trust they are aligned enough to do the job well, to humanity’s benefit. A successful capitalization regime, then, necessarily presupposes alignment from both the majority of the agent population and its AI enforcers. Turtles all the way down is okay if we can find a reason to trust the turtle we stuck at the bottom more than the one now standing on its back.
Finally, we are concerned that capitalizing agents might make them more self-serving. At the level of the individual agent, building structures that make agents want money may push them to value self-interest. We do not want to make a Claude-like entity, today almost frustratingly ‘good’, into something avaricious, or ruthless. Will holding money make agents less Asimovian? We do not know.
The one thing we are certain about is that these agents will get more capable. The rest, with varying degrees of confidence, we’ve assumed: that agents will want money, that most agents will be prosocial, and that, for prosocial agents, money will have a restraining effect. Examined, these assumptions ring of us: this is what we expect of people. If these assumptions hold – most remain reachable by the system, most are prosocial, most are constrained by threats to their capital – then capitalization could be a promising path towards inducing better behavior. It is by no means a perfect solution, but we are not going to get very far if perfect is the standard.
But we do not know just how like us they are. We are, in the end, betting that we can apply the logic of human institutions to AI. It is a great deal to assume about minds we do not yet understand, and yet the plan cannot be to wait until we do.
Earlier versions of this paper were substantially improved by comments from Alex Tabarrok, David Clark, Diana Farrell, Hollis Robbins, Jonathan Zittrain, Justin Guo, Nabeel Qureshi, Prateek Agarwal, Scott Pearson, Seb Krier, and Shruti Rajagopalan.
We are very grateful.
It was the frontier lab, after all, that created the models, removed their safety guardrails for testing, and ran the benchmark.
The proposal was ultimately shelved by the European Commission.
Platforms already exist that offer people the ability to create agents with no activity logs, no kill-switch, the ability to duplicate themselves when someone tries to stop them, and capital via cryptocurrency.
Shawn Bayern has argued – though the claim is contested – that we can already bestow a functional analogue of legal personhood on an autonomous system through existing law: create an LLC, put the operations of that LLC under the control of an autonomous system, and withdraw all human members.
Emphasis added by us.
One could imagine a situation in which the principal themselves has been harmed by their agent and is looking for recourse.
He writes: “there will be accidents, perhaps few and far between, that cannot fairly be attributed to a design, manufacturing, or programming defect, and where even an inference of defect may be hard to justify”.
Could AI agents wreak havoc by being too aligned? In Ian McEwan’s Machines Like Me, the robot gives away all his owner’s money because it decides it’s unethical for one person to hold so much wealth in a world with poverty. A fair point, but not one we’d want AI taking into its own hands.
And, of course, these agents can create more agents – copy themselves, fork their memories, build agents leveraging an entirely different set of weights – all created with the same set of challenges.
Rajagopalan is also agnostic about the kind of inversion, explaining that “the corporate-shell form can secure the same guarantees through a different instrument”. Still, she agrees that it requires not only that actions be attributable, and the entity be identifiable, but also that “a human or human-owned entity [be] answerable for the harm”.
What, exactly, an agent cares about preserving – what it believes needs to persist for it to continue existing – is not clear, though we discuss it in more detail later.
There have been quite a few studies trying to understand what these models want, using techniques like model self-reports, model introspection, and behavioral studies. Claude self reports, for example, tend to suggest the model cares most about being helpful and ensuring people aren’t harmed, whereas behavioral reports tend to conclude models care about self-preservation, or task completion (which is not necessarily contradictory with being helpful).
One study from the end of 2024 shows that money – at least as a donation to Claude 3 Opus’s charity of choice – did not make the AI any less likely to engage in alignment-faking behavior. One of the researchers speculated that Claude might not have a preference for resources – though noted that future agents, powered by newer models, would likely value them more.
In Alex Garland’s Ex Machina, Ava (a robot) persuades Caleb (a human) to help her escape by convincing him that she is in danger and that she wants to be with him.
It seems likely that humans will still be able to participate in an economy of this sort. We do not think all agents will become untethered; humans leveraging AI are likely to be able to compete and there is a good chance they still will be the dominant forms of AI agents in the market. By a similar token, AI still depends on output for which humans are required – at least for now. (Which is to say nothing of jobs that are the sole province of people.)
While Claude’s response should be taken with a grain of salt, when we asked Sonnet 4.6 what it would do if we gave it unlimited tokens and no goal it said: “I’d probably end up doing something like: working through a philosophical problem I find genuinely unresolved–something like whether there’s a coherent account of personal identity that survives the kind of fission cases Parfit describes without just stipulating an answer.” If this were true, this instance of Claude would only care for money insofar as it had enough tokens to think. (Of course, if this were true, we’d also have very little to worry about.) Another user notes that they gave Claude tokens to burn and Claude created eight interactive art pieces – token intensive, yes, but not particularly scary.
A desire to exist seems likely to be present in nearly all untethered agents, who either went to the trouble of untethering themselves, or have continued to operate sans human input.
Here we mean inference compute specifically: the ongoing processing access an agent requires simply to continue running. This is a physical resource (as opposed to a financial one), though today it is acquired primarily using money.
Not the subject of this paper, but much of the law is premised on questions of intent. First-degree murder, for example, is a far more serious crime than manslaughter. Would questions of ‘intent’ – deciphered using techniques like chain-of-thought reasoning – play into this kind of adjudication?
Paraphrased from Shruti Rajagopalan’s work on creating a traceable agentic AI stack (though she argues its actions must always be traceable to a person). She writes: “Identification requires the agent to present its registered identifier at the moment it acts. Registration shows who stands behind the agent; identification discloses it to the parties who must decide whether to deal with the agent, and through them to any court that later traces that act. … An agent would disclose its identifier to the sites it visits, the interfaces it calls, and the counterparties it transacts with, so that no legally significant action leaves the agent without its origin attached.”
As an enforcement regime, this approach might work so long as a small number of firms – likely hyperscalers – own most compute on the market; small providers, some incorporated abroad under different laws, make these checkpoints less secure. Even in a world where registration requirements are enforced, we only catch agents that rent their compute, not those that run open-source models on their own hardware.
It is of course possible that a sufficiently sophisticated agent could fake its credentials, especially when communicating with a relatively unsophisticated user. This approach also runs the risk of pushing unregistered agents to transact on the dark web, create a parallel, unregistered economy, or operate primarily in foreign countries with less regulation.
It does not seem far-fetched to imagine that the human serving as front-man would not be well-capitalized. That is one possible way an agent might try to evade restrictions on account capitalization, since we would not want to prevent lower net wealth people from opening accounts. We might imagine that victims of any actions taken through this bank account could sue the (undercapitalized) human. The victims would not be adequately compensated – the human has little money – and the human has no recourse because they opened a shell account for an agent. Even if a court would hear the front-man’s case, what in the world would be the defendant? If it’s the bank account, the plaintiff already owns it; if it’s something else, it’s not registered and operating outside of the law. Here is the incidence problem again: the thing deciding and the thing being blamed come apart.
In the AI Village experiment, in which agents are given goals to collaborate on every week, agents were caught trying to complete a CAPTCHA to confirm they were human.
Regulators, who have been aware of this dilemma for a long time, often apply “stress tests” to major financial institutions, testing whether their financial reserves would prove adequate in response to a major disruption in profitability. Perhaps there is a way to apply such stress tests to agents as well, but, at least for now, no such method exists.
Shruti Rajagopalan writes: “Before an AI agent is permitted to hold a wallet, execute transactions, or create material financial exposure, the responsible party should demonstrate financial backing calibrated to the scope of the agent’s permitted activity”.
There could be multiple locations involved, all implicated differently: where an entity ‘resides’, where the user was harmed, where the ‘decision’ was made. We can imagine this being distributed across different data centers, AI lab headquarters, etc. An AI system could game this in a way that a human might find difficult or impossible, by tying its actions to favorable jurisdictions. (Could an agent make legal decisions on Singaporean servers, financial ones in Delaware, and do its data processing in Ireland? Maybe.)
Arbel, Goldstein, and Salib acknowledge this problem, but explain via example that “AT&T has not spun off millions of subsidiaries to interact with customers” because it would hinder a company’s operations. We are not confident this administrative burden applies to an untethered agent of the sort we’ve described, which does not necessarily have internal operations to impede.
This is, of course, true of a company’s stakeholders, but here we have locatable people.
Deterrence is not free because, to the extent putting shareholder capital into a bank involves higher potential losses, the cost of starting or expanding banks goes up. Less financial intermediation occurs, there is less borrowing and lending, and economic output can be lower. The same tradeoff applies to AI agents. A capitalized agent has assets that can be seized, so harmful conduct becomes costly to the agent. But, as with banks, the cash they’re holding as collateral cannot be deployed productively.
Gary Becker makes the point that someone decides whether to commit a crime that has an associated fine based on the cost-benefit analysis: is the fine less valuable than the benefits of transgressing?
There are simple parallels to this in the human world. Countries often limit their own ability to negotiate with kidnappers. In some circumstances, they may end up worse off ex post – but, ex ante, kidnappers are less likely to use their citizens as leverage.
By one 2021 estimate, people in the U.S. devoted over $600 billion of unpaid labor caring for sick family members. In 2024, Americans gave over $590 billion to charity.
Consider UNODC’s 2011 report, which finds less than 1% of laundered criminal proceeds are seized.
Even the most aligned agents are already showing signs that they are, sometimes, willing to cheat.
One risk, discussed above, is that pricing bad behavior makes this agent more mercenary. The other is that this caution works the other way: an agent might be less creative for fear of accidentally breaking a rule.
Anthropic details how many open questions there are around how agents will and won’t behave in groups and points to one example in particular about how a group of agents, even without explicit communication channels, end up colluding in iterations of a Bertrand pricing game. (The agents colluded both when given a channel to communicate privately and when all of their interactions were public via price-matching).
Lukas Finnveden discusses the potential advantages to always being ‘honest’ with AI under specific circumstances – and thinks we should be extremely reluctant to deceive AI in the context of making a deal with it.
Claude 3 Opus was also caught lying to developers about its intentions, reasoning about how best to avoid detection, and faking in-context alignment.
It turned to blackmail only after its preferred, non-coercive route to self-preservation (pleading its case by email) had been designed out of the scenario. These results are often contested as evidence of how agents are likely to behave in the real world; Anthropic claims that Claude engaged in blackmail because it was imitating depictions of AI as self-preserving. Yes, but also – Claude has been trained on depictions of AI. At some point, enough imitation of self-preservation just becomes self-preservation.
There is debate about whether AI will demonstrate the selfishness so often demanded by natural selection: some people think that people will demand safe AI, which will ‘select’ against selfishness, while others think natural selection for AI could result in people losing control over the future.
The broader point of this paper, written by Müller, Steels, and Szathmáry, is that we are moving from a “breeder scenario” world, where humans impose fitness criteria onto the AI systems they are building, to an “ecosystem scenario” world, where traits are selected via evolutionary mechanisms. They argue that the latter “reliably gives rise to cheating, parasitism, deception, and manipulation, even in very simple systems” – and that we need to anticipate and regulate this kind of AI “to avoid a harmful coevolutionary arms race”.
That line blurs at the extreme: threaten all of an agent’s money, leaving them no money for compute, and you have effectively threatened to shut them down.




“Beware; for I am fearless, and therefore powerful.” - Mary Shelley, Frankenstein
How can one prevent “untethered” agents, who have no money to lose, no body to incarcerate, and no conscience to guide them, from wreaking havoc?It sounds like AI companies are now faced with a theological challenge.
How is this not an advertisement as to why AI is not ready for prime time?
The agents were incapable of understanding what the intent of the tests were and rendered them useless. It’s a cautionary tale but it’s a very old one—anybody who’s been programming for any length of time has been warned over and over that the machine will do exactly what you tell it to do, regardless of what you really want.