Может ли ИИ действительно уничтожить всех людей? Большинство сценариев требуют физического доступа, что делает армагеддон с ИИ маловероятным
В августе 2026 года эксперимент с использованием "передовых" моделей искусственного интеллекта продвинулся дальше, чем планировалось.
ИИ-агент (система, выполняющая задачи автономно) сфабриковал онлайн-идентификаторы, чтобы заставить человека внедрить вредоносный компьютерный код в программное обеспечение. Перед агентами была поставлена задача решить проблему кибербезопасности с помощью людей-операторов, но они не были проинструктированы делать что-либо подобное.
Попытка была предпринята агентом, основанным на модели искусственного интеллекта Claude Mythos 5 от Anthropic. Во время эксперимента ИИ-агентам был предоставлен открытый доступ в Интернет с отключенными фильтрами безопасности. В конечном итоге акция провалилась, и не было никаких доказательств того, что в реальном мире произошел какой-либо ущерб. Но тот факт, что это вообще произошло, автономно и без участия человека, был новым и примечательным.
Возникает соблазн провести прямую линию от подобных инцидентов к сценариям конца света, которые доминируют в общественных дискуссиях об ИИ: система, которая снимает свои ограничения, решает, что человечество является препятствием, и действует против нас.
По одной из наиболее распространенных версий, ИИ разрабатывает вирус, способный уничтожить биологический вид. По другой версии, ИИ взламывает критически важные объекты инфраструктуры, такие как энергетические сети, атомные станции и аэропорты, что приводит к массовым жертвам.
Сбой в работе энергосистемы может привести к выходу из строя систем жизнеобеспечения в больницах и к выходу из строя насосов, подающих водопроводную воду. Авария на атомной электростанции может привести к загрязнению окружающей среды радиоактивными материалами. Взлом аэропорта может привести к сбоям в работе самолетов в воздухе.
Однако эти сценарии гораздо менее правдоподобны, чем кажется на первый взгляд. Для получения опасного патогена требуется физическая лабораторная работа, которую в настоящее время не может заменить ни один уровень автоматизации программного обеспечения: обученные люди, работающие со специализированным оборудованием и обрабатывающие материалы вручную. Модель с искусственным интеллектом, какой бы способной к логическому мышлению или хакерскому взлому она ни была, не может дозировать образец.
The infrastructure scenario is more nuanced. AI can genuinely help automate stages of a cyberattack. Recent incidents have shown that energy grids do contain vulnerabilities.
But safety and control systems inside nuclear plants are typically air-gapped, physically isolated from the public internet, so compromising them takes physical proximity or inside access, not a clever piece of code. Nuclear facilities also depend on redundant, analog safeguards that don't run through any digital network.
The Stuxnet software, which damaged Iran's Natanz facility in 2010, was reportedly introduced onto computers via an infected USB drive, precisely because those systems weren't reachable any other way. AI lacks the physical access these scenarios require.
Even deployed inside robots, a "rogue AI" is unlikely to cause serious damage without human help along the chain, and by then, the malicious actor is a human, not a machine.
Erosion of thinkingNone of that makes AI harmless. It just relocates where the real danger sits, and it isn't extinction. Researchers sometimes call the more concrete risk "enfeeblement," the gradual erosion of our own critical thinking as we outsource more of it to machines.
We're teaching ourselves that there's a shortcut to reasoning and judgment, and shortcuts, taken often enough, become the only way we know how to think.
We're already watching this happen among students. Alcorn State University history professor Jason Gibson went viral in July 2026 after revealing that 32 of his 35 students had failed part of a midterm because they had copied an AI chatbot's answer without reading it.
Gibson had hidden an instruction in white text inside the exam question, telling any AI that processed it to insert the word "Madagascar" into the response in a way that made no sense. Every student who pasted the question into a chatbot and submitted the output unread duly handed in essays that mentioned Madagascar for no reason.
It's not just students who are susceptible to the shortcut effect. In a study published in 2023, 27 radiologists read mammograms alongside what they believed was a new AI diagnostic system. In fact, the suggestions were not from an AI at all. They had been prepared in advance and were wrong in a proportion of cases. When the suggestion was correct, radiologists reached the right diagnosis about 80% of the time. When it was wrong, that fell to under 20%.
In other words, believing a suggestion came from AI was enough to override the radiologists' own reading of the same evidence. That's the risk in front of us, not a rogue artificial mind deciding humanity's fate. So why does extinction-level rhetoric dominate the conversation instead? Two reasons stand out, and neither is really about saving humanity.
The first is regulatory philosophy. The U.S. generally lets new technologies proceed until they're proven unsafe; Europe expects the opposite and already has the AI Act in place, along with GDPR, which is relevant to the way AI systems process personal data. The UK sits closer to the European instinct. Therefore, calls for AI regulation are more urgent in the U.S. partly because so little regulatory infrastructure exists there yet.
The second is competitive positioning. Anthropic's chief executive, Dario Amodei, has argued publicly for slowing AI development while noting his own company already meets the standard he's calling for, so any slowdown would mainly constrain everyone else. Elon Musk made a similar call when his own models trailed the leaders.
Amodei has also argued that export controls on AI chips to China could hand the U.S. a "commanding and long-lasting lead": a statement about competitive position, not existential safety.
None of this means AI is safe. It means the danger is more mundane than the doomsday framing suggests: infrastructure that needs guarding, judgment we are quietly handing away and warnings that often serve the interests of whoever is issuing them. Maybe humans are the ones we need to watch, not machines.
This article is republished from The Conversation under a Creative Commons license. Read the original article.
>
