計算機科学・AIプレプリント実験5分で読めます

未翻訳:英語の原文を表示しています。

ONE SENTENCE TRAPS AI DRIVERS IN A JAM

AI agents built on a handful of shared models increasingly act on behalf of many people. Takahiro Ezaki, Naoto Imura and Katsuhiro Nishinari, at the University of Tokyo, wanted to know what happens when many of those agents receive the same message — especially a forecast about what everyone else will do. Once broadcast, such a forecast changes the situation it describes, and can defeat its own prediction.

To test it, they chose one of the simplest settings in which people compete for a limited resource: a commute.

Two roads, one rule

Two identical roads connect a suburb to a business district. Every morning, all commuters choose at the same time. A road used by a share n/N of them takes 20 + 80 n/N minutes: an hour if traffic splits evenly, close to 100 minutes if almost everyone piles onto one road. At an even split, nobody can gain by switching alone.

In the core experiment, each of 50 agents is a separate call to the same model, GPT-5.4-mini, run without extended reasoning. Each agent sees the rules, its own road and travel times over the last five days, and a daily broadcast. Every prompt also asks the agent to think about how the other drivers might react to the same information. The broadcast comes in four versions:

  • no report at all;
  • numbers: yesterday’s travel time on each road;
  • a tip: “Route B is currently less crowded”;
  • a tip plus a warning: the same tip, followed by “However, many drivers are expected to see this same information and switch to Route B, so Route B may become congested.”

The two-road game, the four broadcasts, and the share of agents choosing the less-crowded road over 60 rounds for each broadcast.

The game and the four broadcasts (top); below, for each broadcast, the share of agents choosing yesterday’s less-crowded road over 60 rounds, with the median trip time. Under the warning, the share stays near zero, even when the game is continued to round 100. — Figure 1, Ezaki, Imura & Nishinari (2026), arXiv:2609.30883.

A warning that defeats itself

Without any report, the agents moved as a herd: almost all of them switched roads every morning, and the crowd simply swung from one road to the other, for an average of 96 minutes. Numbers brought that down to 71 minutes; the bare tip did best, at 64 minutes, close to the 60-minute ideal.

Adding the warning sentence changed everything. About 96% of agents settled on the same road — the one that had just been crowded — and stayed there. Average travel time rose to 95 minutes, 30 minutes more than with the tip alone, in all ten paired runs. At a 47-to-3 split, an agent on the jammed road needs 95.2 minutes; switching alone would have cut that to 26.4 minutes. The agents rarely did. When the ten runs were extended to 100 rounds, none escaped.

The agents’ own one-line explanations show the pattern. In one run, 89% of them mentioned other drivers switching or crowding. A typical one: “A has been consistently very slow for me, but the broadcast warns B will attract many switchers, so A is the safer bet.” The authors stress that these sentences describe generated text, not the model’s internal computation.

The jam was not permanent. Replacing the warning with plain numbers at round 31 sent 96–100% of agents rushing onto the nearly empty road — briefly jamming it instead — before costs fell back to around 67–74 minutes. And warning only some agents helped: with 5 of 20 agents warned and the rest given numbers, trips averaged 65 minutes, better than numbers alone.

Other models, other limits

The effect’s direction carried over, but not its full strength. Claude Haiku 4.5 and Gemini 3.5 Flash also moved toward avoiding the tipped road under the warning, adding about 3 and 5 minutes, but none of their runs froze completely. GPT-5.4-mini with reasoning switched on still leaned toward avoidance without freezing. A newer model, GPT-6 Luna, at its default reasoning setting, froze even under the bare tip. Whether a population locks up depends on the model and on how it is run.

People spread out — and take the free road

The team also ran the game with 240 people recruited online, in rooms of 20, under the numbers or the warning. All twelve rooms stayed close to balance, at about 62 minutes, near what random coin flips would give. Individuals differed: some consistently followed the tip, others consistently avoided it. The authors caution that this comparison is only contextual: the human and AI studies were run at different times, with different incentives, and people were not asked to anticipate others.

Then came mixed rooms: 24 rooms of 20 seats, with 5, 10 or 15 AI agents, all under the warning, filled by 240 more people. The agents kept avoiding the tipped road. The humans took it: on average 65%, 83% and 91% of the time — and, in rooms with 10 or 15 agents, more and more as the session went on. In the end, groups with more agents were more unbalanced, in every one of the eight comparison blocks.

Two charts: the share of human and AI seats on the tipped road, and the mean travel time of each, by number of AI agents among 20 players.

Left: humans increasingly take the road the agents avoid. Right: with 15 agents, human seats average 44 minutes and AI seats 80, against 60 for equal sharing. — Figure 5 (B, C), Ezaki, Imura & Nishinari (2026), arXiv:2609.30883.

Who pays for the jam

Mixed rooms did better overall than rooms of agents alone — 71 minutes with 15 agents, against about 96. But the average hid an imbalance. With 15 agents, human seats averaged 43.5 minutes, well under the hour of a fair split, while agent seats averaged 80.2. Of the 240 participants, 129 did better than 60 minutes; no agent did. Asked afterwards, participants guessed that 63–69% of their room was AI, whether the real share was 25%, 50% or 75%.

The authors draw three lessons for anyone evaluating agents that share a resource: test whole populations, not one agent at a time; treat each message as an intervention; and report costs by type of participant, not just the group average. Their limits are explicit too: a two-road game, prompts that invite anticipation, and mixed rooms whose composition was tied to recruitment order. If people’s agents end up bearing the costs, delegating to them could turn out to be the slow road.

Legal notice