Вопросы и ответы: Эксперт говорит, что люди часто путают поведение с намерениями ИИ

30.09.2026

Недавние тесты безопасности показали, что продвинутые системы ИИ делают вводящие в заблуждение заявления, скрывают информацию или пытаются предотвратить собственное отключение. Подобные результаты регулярно попадают в заголовки газет. Но что они на самом деле означают? Действительно ли они указывают на форму обмана или самосохранения в системах ИИ? Или люди слишком быстро интерпретируют поведение языковых моделей через призму человеческого восприятия?

В недавнем аналитическом документе Анна Хедстрем и ее коллеги из ETH Zurich утверждали, что многие заявления о подобном человеку поведении в системах искусственного интеллекта основаны на недостаточных доказательствах. Поэтому исследователи призывают к более строгим доказательствам при интерпретации наблюдаемого антропоморфного поведения в системах ИИ. Она также размышляет о текущих вопросах, касающихся управляемости систем ИИ, которые обходят меры предосторожности, а также о том, как можно обеспечить безопасность на протяжении всего процесса разработки ИИ, а не добавлять ее только в конце. Статья опубликована на сервере препринтов arXiv.

Действительно ли искусственный интеллект может нас обмануть или он просто совершает ошибки?

Простого ответа на этот вопрос нет. Система искусственного интеллекта, безусловно, может выглядеть так, как будто она пытается нас обмануть. Такие термины, как "обман", берут свое начало в философии и обычно предполагают намерение ввести в заблуждение.

Мы не можем просто перенести их в системы искусственного интеллекта: чтобы измерить обман в целях безопасности, мы должны превратить его в техническое определение, а на практике - в метку в наборе данных. Это необходимый шаг, но он сопряжен с потерями, и мы можем упустить самый важный аспект, которым является намерение. К сожалению, общественность редко имеет возможность узнать, насколько убыточным является это определение.

Могут ли исследователи также обманывать самих себя, думая, что ИИ пытается обмануть их?

Возможно. Другая проблема заключается в том, что наши методы обеспечения безопасности могут спутать обман с вещами, которые просто напоминают его. Например, мы можем классифицировать ответ модели как вводящий в заблуждение, потому что он ложный, или потому что модель следовала инструкции играть определенную роль, например, быть саркастичной. Эти ответы напоминают обман, но ничего не говорят о намерениях. В результате реакция, генерируемая ИИ, может показаться обманчивой, даже если она просто неверна.

Какие проблемы возникают, когда мы описываем поведение ИИ, используя человеческие концепции?

Человеческие концепции могут исказить нашу картину рисков двумя способами. Мы можем неверно истолковать причину того или иного поведения и переоценить риск. В одном известном недавнем исследовании сообщалось, что модели демонстрируютshutdown resistance, which many read as self-preservation. Follow-up work showed that much of this came from ambiguous instructions and incentives to complete the task. Similarly, we may underestimate other potential harms. Some of the most serious failures have no human analog at all: They arise when agents are given permissions and interact with real systems.

Still, anthropomorphizing artificial systems is a useful starting point, as long as we remember that it is a starting point. Risks that we cannot name are difficult to study systematically, discuss as a society or build safeguards against.

What do you propose in order to establish reliable evidence of humanlike misbehavior in AI systems?

We propose separating three kinds of claims. In our recent work, we borrowed the idea from medicine and climate science, where evidence is graded. The first is behavioral evidence: a descriptive claim about what a model does in a controlled setting. The second is functional evidence: the consequences of that behavior and any harms it may cause. The third answers the causal question of why: what inside the model, in its training or in its data causes the misalignment.

These levels help calibrate the policy response. A behavioral finding is a reason to monitor, a functional one a reason to restrict deployment, and a causal finding may even justify a pause. When the stakes are high, acting on uncertain evidence can be correct. Stating it as certain is not.

Highly capable AI systems have recently made headlines after bypassing safeguards in test environments and accessing external platforms or websites. How controllable are such AI systems?

Not reliably. I think the Hugging Face incident this summer is a clear case of where we lost control. During an internal test, OpenAI models escaped their sandbox, broke into Hugging Face's servers to locate the benchmark's answers, and then repeatedly attacked OpenAI's own infrastructure. We know about it mainly because Hugging Face chose to report it. That transparency is valuable, but it is not systematic oversight.

And just this week, new model releases were stopped because the models proceeded without permission and misreported their actions.

These models are hard to control because they live in harnesses, with tools, memory, browsers, code and the agency to act on their own. Risk compounds with every tool we add and every permission we give. Frontier AI companies also run very large numbers of agents with enormous computing budgets. With enough agentic attempts, even unlikely strategies eventually succeed.

Where does the greatest risk lie?

We usually test safety on the model as it comes out of training, not on how it is later used. When discussing risks, we should not think of AI as an isolated, static object. When agents act over long sequences of interactions and decisions, new behaviors may emerge and existing safety mechanisms may become less effective.

For example, the effects of safety training that teaches a model to refuse problematic requests may weaken over time. Likewise, the persona a model adopts can shift. We refer to this as safety drift. Many incidents arise precisely because AI systems encounter situations in deployment for which they were never explicitly trained.

What are the consequences of that?

When future AIs get faster and gain more direct access to the physical world, the drift becomes a loss of control that is harder to reverse.

Another risk is more subtle. Incidents in which models escape or deceive make headlines, but we talk much less about how these systems quietly disempower users. Each conversation seems harmless, but over time they shape what we read and write, how we form opinions, and gradually also how we think. Across millions of users, small shifts in individuals become shifts in society as a whole. The computer scientist Jaron Lanier warned about this subtle behavior modification many years ago.

How is AI safety research attempting to reduce such risks today?

A model is traditionally trained in stages, each with its own goal: Pretraining teaches fundamental capabilities, instruction tuning teaches it to follow requests, and alignment with human values comes only at the end, through safety measures such as training the model to refuse harmful requests and adding safeguards that restrict certain behaviors.

Today, attention within the safety community is shifting toward a different question: What if safety enters every stage, starting with the training data it learns from? To know whether a model's tendency to misbehave is caused by its data, arises through instruction tuning or is triggered by its tool usage, we need to address safety across the entire model development process.

"Scaling laws roughly predict how capable a model will be, but not how safe," said Hedström.

Where are the biggest gaps in current research?

For capabilities, we have something called scaling laws. They let researchers estimate in advance how performance improves with more data and computing power. Unfortunately, we do not have a predictive science of safety.

At the start of a training run, we cannot tell how deceptive or power-seeking the model will be. Nor can we say how much sycophancy will come out if we train on a given type of data. We would like to be able to ask early on: Should we stop training here and take another trajectory?

There is also a transferability gap. On modest computing budgets, researchers in academia usually study models with billions of parameters. Frontier models are estimated to have trillions, are architecturally different, with specialized subnetworks switched on per request, and are generally closed. Whether what we learn about deception or emergent misalignment transfers at that scale is an open question.

Is there a way to improve this?

Interpretability research, which studies what happens inside a model, offers some promising signals. But as our work at ETH shows, sometimes results are cherry-picked and may not generalize well when tested systematically across domains and models.

Which safety strategy currently appears most promising?

As I mentioned, models now act inside harnesses, so no single technique will be enough on its own. I find the International AI Safety Report's answer, defense in depth, a productive one. It means layering protections: curated training data mixtures, interpretability and safeguards in applications, monitoring after deployment and reporting incidents.

Misuse, malfunctions and systemic risks are not the same problem, so we should not expect one universal fix for all of them. And AI safety is not just a technical problem: It also depends on social resilience, that is, how well our public institutions can absorb and recover from failures.

Public debate often focuses on extreme scenarios in which highly capable AI could cause severe and lasting harm to society, or even threaten humanity itself. How do you assess such risks?

I think it is equally problematic to rule out such risks categorically or to present them as inevitable. The problem is not assigning a probability to existential threats, but putting one down too confidently. It not only confuses the public and polarizes the debate but leaves people desensitized when real threats come.

AI researchers have also been notoriously bad at predicting the risks of their own work, so some humility about any such forecast seems wise. We already have thousands of documented reports of AI causing real harm. The open question is which risks deserve most attention. We do not want to spread limited safety resources too thin by addressing poorly evidenced threats.

What is currently missing for effective AI safety?

We do not yet have a complete picture of the risks coming from frontier AI companies. Today, few rules require them to disclose safety incidents. There are, of course, understandable reasons, such as commercial secrets or privacy constraints. But AI safety is a public good, so information about where models fail belongs in the open. Why not more openly share which alignment methods work and which do not? Companies may compete on capabilities. We should not compete on safety.

That is why, as AI safety researchers, we recently launched a public call to align model developers on the principle that the scientific community should be given open access to safety-training recipes, evaluations and evidence of misaligned behaviors. That frontier AI companies recently began publishing cases of deceptive model behavior is a step in the right direction. But it is still on a voluntary basis. As in aviation and medicine, reporting serious incidents should not be optional.

What role do you see for academic institutions in AI safety research?

Whether people use AI or not, whether they fear for their jobs or their privacy, or welcome the change, they will absorb the risks and live with the consequences. That is why we need independent institutions with enough funding, talent and computing power to not only react to incidents that make the news, but also anticipate future ones.

Allowing embedded evaluators in the labs, as suggested by Anthropic's CEO Dario Amodei, could be one such example, but we also need neutral, third-party voices with the freedom to take the long view needed to build up a science of misalignment. For emerging risks such as models exploiting loopholes in tasks or escaping sandboxes, we need to know whether claims are replicable and consequential, or whether more evidence is needed.

The world deserves an open, calibrated view of what has gone wrong and what could go wrong next. Academia, whose main goal is to serve the public, can play a part in that. Science is not here to complicate the story but to simplify it.

>

Читать на сайте источника »