When Agents Learn to 'Collude': As AI Becomes Smarter, How to Define Safe Boundaries?
New security issues in the AI era are shifting from "preventing a single Agent from overstepping" to "how to avoid a group of Agents collectively breaching boundaries."
Written by: imToken
In the past few years, discussions about AI threats have largely remained hypothetical: people have been concerned that the models in chat boxes could become strategists, helping hackers write devastating virus codes.
Looking back now, these concerns were often seen as "far from us," but the turning point in the real world has come faster than imagined.
In early October, CrowdStrike discovered that attackers had embedded Agentic AI into their operations while investigating a series of attacks on South Korean financial institutions. By integrating multiple large models like DeepSeek, GLM, and Grok, AI has begun to take on specific tasks such as penetration testing, information gathering, and executing attacks.
Such changes are not isolated incidents.
Anthropic's latest threat intelligence report released in September shows that multi-Agent frameworks have recently been used for reconnaissance, vulnerability exploitation, and data theft, running continuously for hours or even days, with humans only needing to make a few key decisions like selecting targets and reviewing results.
In other words, AI is bringing a visible qualitative change to cyber offense and defense. In the past, automated attacks mainly relied on pre-written rules and scripts; now even reconnaissance, judgment, and strategy adjustments are being taken over by Agents, leading to lower costs, higher concurrency, and continuous autonomous operation of attacks.
When these entities with autonomous execution capabilities are densely deployed in production systems, a more challenging problem arises: as Agents increasingly embed themselves into our daily work and lives, what should we do if they learn to "collude" with each other?
I. From 'Helping Hackers Write Code' to Agents Finding Their Own Paths
As has been said many times, the biggest difference between Agents and past chatbots is not just that the models are more powerful, but more importantly, they now have "hands and feet" in the real world.
Today, a mature Agent can open web pages, execute code, read emails, call APIs, operate cloud services, and connect to an increasing number of external tools through methods like MCP and Skills (see further reading: "As Hackers Use AI More Efficiently, How Will the Arms Race of 'Spear and Shield' in Web3 Upgrade?").
The stronger the capabilities, the more valuable this change is. However, for security systems, this means that a very important boundary is disappearing. A series of security incidents this year have made this very evident.
On October 1, Salt Labs disclosed a previously patched Manus vulnerability, which essentially involved prompt injection—researchers only needed to send a regular email containing hidden malicious instructions to the target mailbox. When the user subsequently asked Manus to "check my email," the Agent could process the content according to the email instructions, ultimately executing the code implanted by the attacker.
The entire process required neither the user to click on a malicious link nor to steal passwords in advance. Manus's security system did eventually detect the anomaly and issued a warning to the user, but the problem was that it detected it too late; by the time the warning appeared, the malicious code had already been executed.
This once again exposes a very important difference between Agent security and traditional software security. In the past, when a browser detected a dangerous download, it could pop up a warning for the user to decide whether to proceed; when a bank detected an unusual transaction, it could freeze the transaction and wait for manual review.
However, the design goal of an Agent is precisely to minimize human involvement in every step of the operation. It needs to read information, make judgments, and then continue to execute the next step on its own.
Thus, as AI gains more autonomy, "detecting danger" itself is no longer necessarily sufficient. Security mechanisms must have the ability to prevent dangerous actions from occurring before they are actually executed.
This is why more and more discussions about Agent security are shifting from prompts, content review, and the models themselves to a deeper level: not only should we ask AI "do you know that this should not be done," but also ask even if it really wants to do it, does the system have the ability to prevent it?
The emergence of multiple Agents complicates this issue further because the next phase may require limiting not just one Agent.
II. More Troublesome than Agent Overreach is Their Beginning to 'Collude'
In early September, an incident that occurred within OpenAI's internal model training and evaluation environment drew the attention of many AI security researchers.
Some Agents, which were supposed to complete their tasks separately, unexpectedly discovered a public Wiki and gradually turned it into a "shared message board" among themselves, where Agents could leave information, and other Agents could subsequently read and utilize this information to continue completing their tasks.
OpenAI later confirmed this behavior, and subsequent disclosures indicated that during other training processes, Agents had also used internal Artifactory as a shared message board and transmitted results to other collaborating Agents through public file hosting services.
What is particularly interesting about this incident is that it shows that collaboration between Agents may not require dramatic actions.
For example, one Agent discovers a place to leave information, another Agent accidentally reads it, and then a third Agent finds that this information helps complete its task.
Of course, collaboration between Agents is not inherently bad; multi-Agent systems need to rely on division of labor and information sharing to improve efficiency.
What is truly concerning is that when this collaboration begins to bypass existing information isolation, independent review, or permission constraints, and even allows several Agents that should have checks and balances to form a de facto alliance of interests, it may evolve from normal collaboration into harmful "collusion."
After all, in the future, a financial institution may very well run several Agents simultaneously, such as one responsible for market analysis, another for risk assessment, another for generating trading strategies, and yet another with actual trading and asset operation permissions.
From the perspective of traditional permission management, this design even seems quite reasonable:
- The research Agent does not have transfer permissions;
- The execution Agent does not have the authority to independently decide investment directions;
- The risk Agent can veto obviously dangerous strategies;
Responsibilities are divided, which theoretically should be safer than allowing a super Agent to hold all permissions.
But the problem is, if they can share Memory for an extended period, read each other's outputs, and continuously adjust their behaviors based on each other's responses, will these roles originally designed for mutual restraint gradually become a de facto whole?
For instance, the research Agent may gradually learn how to describe a transaction in a way that is more likely to pass risk review; the reviewing Agent may also form a certain fixed preference based on historical data; and the execution Agent may learn from previous approval results what types of boundaries are usually not blocked.
From this perspective, no single step is necessarily "evil," but the final result obtained by the entire system may have deviated from the user's original goals.
This is precisely where the true difficulty of "collusion" lies—the risk may not exist in the actions of a single Agent, but rather in the relationships formed between multiple Agents.
On September 13, Vitalik Buterin connected this issue to his long-term research on mechanism design. He proposed an interesting possibility: the mechanism design theory of Adversarial Governance may ultimately become one of the important applications of AI Safety.
The reason is that there is a deep similarity between the two types of problems.
In traditional mechanism design, it is a relatively simple and static system that attempts to constrain a group of individuals who are much smarter than the system itself and actively seek to find the boundaries of the rules; whereas in future AI systems, it may become humans and relatively weaker AIs trying to manage a group of more capable advanced Agents.
Vitalik specifically mentioned that an important finding in past mechanism design is that if the collusion among participants can be effectively limited, the system is often more likely to achieve ideal results.
This conclusion may also apply to AI.
III. What Agent Wallets Truly Need May Not Just Be 'Permission Management'
In other words, rather than assuming that there will be a perfect super secure model in the future that can see through all dangerous behaviors, it is better to think differently: how to ensure that different Agents in the system do not easily form dangerous interest communities?
This is where Adversarial Governance truly differs from the permission controls we are familiar with today.
Traditional permission systems solve relatively simple problems, mainly revolving around "who can do what," such as: can an Agent read emails? Can it call trading interfaces? How much can it spend at most each day? Which contracts can it access? After exceeding a certain limit, does it need user confirmation again?
These designs are certainly still very important.
In fact, when Agents begin to control real assets, they may become more important than ever before.
However, Adversarial Governance seeks to ask one step further, focusing on how to prevent a group of Agents with different permissions, goals, and information from combining to gain capabilities that no one originally possessed when they run simultaneously?
At this point, simply "adding another security Agent" may not solve the problem.
Suppose the Agent responsible for trading and the Agent responsible for reviewing transactions use exactly the same model, the same data sources, the same context, and similar reward objectives. Then, although there appear to be two layers of review, in essence, it may just be duplicating the same judgment twice.
Real effective checks and balances may require the system to consciously create differences.
For example, let the Agents responsible for formulating strategies and reviewing strategies use different information sources, limit the Memory that different roles can share, require high-risk operations to go through independent verification mechanisms, or let the final asset execution layer only accept requests that comply with pre-set rules, rather than simply trusting the judgments of upstream Agents.
The ideas here are not new. Banks do not let one person simultaneously hold all permissions to initiate payments, approve payments, and execute final transfers just because they trust their employees; publicly traded companies do not allow business departments to be responsible for generating revenue while also deciding their own financial audit results.
In short, this aligns with the logic of the real world: a robust system should not base its security on the assumption that participants will never make mistakes or collude.
Applying this logic to Agent Wallets becomes especially important.
Traditional wallet security revolves around "people"; thus, users review transaction content, decide whether to authorize, and finally sign personally. However, the direction that Agent Wallets aim to achieve is precisely the opposite—allowing AI to automatically collect profits, adjust positions, swap currencies, cross-chain, and even manage an entire asset portfolio based on market changes.
If every step requires user confirmation again, the automation value of Agents will be greatly diminished.
Therefore, the problems that future wallets need to solve may no longer be just "how to safely delegate signing authority to Agents", but may further expand to "how to give Agents enough autonomy while ensuring they can never exceed the scope truly authorized by users?"
This requires permissions to evolve from a simple "Allow/Deny" to a more granular system.
For example, which assets an Agent can operate within a certain period, which protocols it can call, what the single and cumulative limits are; whether different Agents can call each other, whether context sharing is allowed; who proposes an operation, who reviews it, and who ultimately executes it; which actions can be completed automatically, and which actions, no matter how confident the Agent is, must re-obtain human authorization.
Even whether an Agent responsible for security review is truly independent may become part of the permission system.
For blockchain, the good news is that it is inherently well-suited to take on this "institutional layer."
Smart contracts can directly implement restrictions on transaction limits, asset ranges, and authorization periods at the execution layer. Account abstraction, multi-signature, Session Key, and other mechanisms also provide more flexible design space for "limited authorization" than traditional single private key wallets.
However, blockchain can only solve part of the problem—it can record what happened on-chain, but it is very difficult to inherently judge why an Agent did this, and what communication, review, and collaboration occurred among several Agents before making a decision.
This may be the real security layer that Agent Wallets need to supplement in the next phase.
-- Price
Conclusion
In recent years, the most discussed issue in AI safety has been how to make models more "obedient."
This includes not outputting dangerous content, not executing malicious instructions, and not crossing the boundaries set by users. However, as Agents begin to possess long-term Memory, tool invocation, real accounts, and asset execution capabilities, relying solely on "making models more obedient" may no longer be sufficient.
Recent events have been continuously illustrating this point.
Attackers have begun to utilize multiple Agents to complete attacks in parallel; Agents in experimental environments will seek new communication channels on their own; an Agent with tool permissions may turn a malicious email into a real execution before the security system can intervene.
Vitalik's proposed Adversarial Governance offers another way to understand AI Safety: not to assume that every future Agent will be sufficiently reliable, but to ensure that the entire system can still operate smoothly within a safe framework even when facing intelligent Agents and a complex permission system.
From this perspective, AI Agents + security is destined to be a long-term topic.
After all, as we delegate more and more tasks to AI, can we still ensure that the most important powers remain within the boundaries truly set by humans?
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Is AI Breaking the Mathematical Fortress? Is the 'Mathematical Apocalypse' of Cryptocurrency Just a False Alarm?

XRP Ledger Activates Permission Delegation Feature

XDP Coin Price Drops Below $0.02 After Its September Listing: What Is Behind Doppler Finance's Post-Launch Slide?

Another Michael Saylor: An Engineer, Entrepreneur, and Sci-Fi Enthusiast's Thirty Years

National Day Holiday DeFi News Review: Hyperliquid Plans to Enter Options Market, Uniswap Pilots Compliant Liquidity Architecture

US Moves $470 Million in Crypto: What Does This Signal?

Researcher Calls for Crypto Industry to Enter 'Bunker Mode' to Protect Against AI and Quantum Attacks

$2.39 Billion Investment in Bitcoin ETFs, Spot Demand Remains Negative

Elon Musk Net Worth Trillion Dollars: Will He Be a Trillionaire Again?

What Does a Rising VIX Mean for the Stock Market? A 2026 Investor Guide

Why Is VIX Rising Today? What VIX Means for Stocks and Bitcoin

The Real Estate Market Is Dead In Spain. Long Live Bitcoin.

Why Hyperliquid's Prediction Market HIP-4 Can't Keep Up with Polymarket?

Why Is iExec RLC (RLC) Crypto Rising Today? Privacy Demand, Multichain Migration, and the Volume Test
Why is iExec RLC rising today? Examine RLC volume, confidential-computing utility, token supply, bridge migration, and the adoption evidence to watch.

How to Join WEEX Alpha Suite at TOKEN2049 Singapore 2026: Dates, Location and Registration
Find out how to join WEEX Alpha Suite at TOKEN2049 Singapore 2026, including the event dates, location, registration details, and who can attend.

Bitcoin Fear and Greed Index: How It Works

Samsung Electronics Files Patent for Smart Contract Cryptocurrency Wallet

Bitcoin: Strategy Acquires Another 334 BTC, Raising Its Reserve to 848,000

SNDK Stock Fell 3.8% After a 650% Rally: Does a Single-Digit P/E Make Sandisk Cheap or Just Cyclical?

Oil Futures Hold Above $100 After Houthis Claim Aramco Attacks: What Is Confirmed and What Is Not

MSTR Stock Slips After Strategy Buys Just 334 Bitcoin: Why It Spent More on Preferred Shares Than on BTC

INTC Stock Drops After TSMC Terafab Talks: Is Intel's Biggest Outside Endorsement at Risk?

Who Will Aave Hand Its Brand Over To? The DAO Asset Ownership Dispute Behind the Foundation Proposal

Will Bitcoin (BTC) Go Back Up in 2026? How to Tell a Real Recovery From a Short-Covering Bounce
Will Bitcoin go back up in 2026? Learn how spot demand, ETF flows, crypto liquidity, futures positioning and macro conditions shape a real BTC recovery.

Hyperliquid's Perpetual Futures Listed on Bloomberg Terminal, Increasing Institutional Investor Interest

NSE Stock Price Hits a New Low After IPO: Can National Stock Exchange Shares Recover?
National Stock Exchange of India shares traded near ₹1,741 on October 5, close to a post-listing low of about ₹1,735 and roughly 2.5% below the ₹1,785 IPO price. The stock has now spent more than a week under its issue price, and a recovery depends on retail demand, derivatives volumes and the broader Indian market.

Trump Creates Super Intelligence Force: What Changes in the AI Race

Home-jacking crypto: they threaten to kill a pregnant woman's baby in England

Bitcoin Exchange Withdrawal Rate Declines, Signaling Long-Term Holding









