Пакистанские судьи вынесли свой вердикт по делу JudgeGPT

12.08.2026

Судьи по всему миру попали в заголовки газет из-за незаконного использования генеративного искусственного интеллекта в своей работе. Но в Пакистане в ходе крупномасштабного испытания специально разработанного инструмента искусственного интеллекта для судей было установлено, что эта технология - в сочетании с соответствующим обучением — позволила увеличить количество разрешенных дел на 6,3 процента без явного снижения качества судебных решений.

С учетом того, что число нерассмотренных дел составляет 2,26 миллиона, а на 100 000 человек приходится менее двух судей (по сравнению с 22 судьями в ЕС и восемью судьями в Бразилии), судебная система Пакистана остро нуждается в помощи. Итак, проконсультировавшись с судебными органами, экономист Султан Мехмуд из Российской экономической школы в Москве и его коллеги проверили, может ли искусственный интеллект облегчить это бремя.

Они создали специальный инструмент, объединяющий крупноязычную модель OpenAI GPT-4 (LLM) с базой знаний, насчитывающей почти 130 000 судебных заключений и нормативных актов Пакистана, чтобы помочь судьям в проведении юридических исследований и составлении судебных решений. В 2024 году они начали предлагать этот инструмент 1559 судьям первой инстанции — примерно половине судей страны.

"Мы действительно отмечаем увеличение числа разрешенных дел и не замечаем соответствующего снижения качества решений", - говорит Мехмуд.

Это первый в своем роде инструмент

"Просто удивительно, что ему удалось провернуть такое", - говорит Дэвид Отор, профессор экономики Массачусетского технологического института. "На государственной службе нелегко проводить масштабные полевые эксперименты, особенно там, где ставки так высоки". По его словам, повышение производительности на 6,3% не является ошеломляющим, но оно заслуживает доверия и, вероятно, будет улучшаться по мере более широкого использования этого инструмента.

Инструменты искусственного интеллекта для судей уже внедряются в Бразилии и Индии, а известный американский профессор права Эрик Познер сравнил решения магистратуры с человеческими в рамках одного тематического исследования. Но до сих пор не проводилось серьезной независимой оценки текущего использования искусственного интеллекта в судебной системе. Новое исследование было посвящено пакистанским судам первой инстанции; Мехмуд говорит, что судьи с самого начала были полны энтузиазма.

"Они были большими технооптимистами, чем мы", - говорит он. "Задержки настолько велики, что, по их мнению, в любом случае стоило попробовать, чтобы уменьшить страдания людей".

По словам Мехмуда, некоторые судьи также уже использовали чат-ботов с искусственным интеллектом, но коммерческие предложения плохо отвечали запросам пакистанских юристов, часто вызывая галлюцинации в судебной практике. Поэтому команда разработала инструмент, адаптированный к пакистанскому контексту, под названием JudgeGPT.

They used retrieval-augmented generation (RAG), which allowed the model to query a database of 128,292 Pakistani judicial opinions and 943 statutes. Responses included footnotes linking to cases and laws.

"It turns out that actually the way to fix [hallucinations] isn’t just more intelligent models," says study coauthor Elliott Ash, an associate professor of law, economics, and data science at ETH Zurich in Switzerland. "It’s to attach the models to a tool that can do a search and verify the sources." However, the researchers do not report hallucination rates.

The team also put 1,197 judges through six 90-minute Zoom training sessions, developed in collaboration with Pakistan’s Federal Judicial Academy, covering how LLMs work, their limitations, the risk of bias and hallucinations, and the importance of verifying outputs. Another 180 judges only underwent general training on technology in legal research, while a final group got no training.

By the time 487 judges had been through the program, the median district saw a jump of 6.3 percent resolved cases, and the more trained judges in a district, the bigger the effect. Appeal rates also fell slightly, suggesting faster resolution wasn’t leading to sloppier decisions.

JudgeGPT can be used to surface relevant case law with a simple text query, and results provide links to the full text of the related judgments.Sultan Mehmood, Christoph Goessmann, and Elliott Ash

Addressing limitations

The team also assessed the quality of judgments. Having legal experts evaluate large numbers of judgments was infeasible, Mehmood says, so the team asked OpenAI’s GPT-5-mini to choose between pairs of judgments from the same judge before and after training. The LLM chose post-training judgments 59 percent of the time. Two experienced Pakistani lawyers also evaluated the model’s analysis of 90 judgment pairs. They agreed with GPT-5-mini 70.6 percent of the time, compared to 73 percent agreement with each other.

Training turned out to be vital. On average, JudgeGPT-trained judges logged in 56 times and sent 212 prompts over the study period, compared to 10 logins and 25 prompts after generic training. Those who had no training tended to use the tool for around a month and then drop off entirely, Mehmood says. "Just giving people the technology does not necessarily make them use it persistently," he says.

A 6.3 percent increase sounds modest, but the researchers calculated that a trained judge was resolving 38.5 more cases a month than the baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash also notes that these figures come from a nine-month period at the start of the trial, and that they’ve since updated both the underlying AI model and the database.

For users, the tool has been a lifeline. One participating trial judge, who spoke on condition of anonymity, says the number of cases assigned to them hasn’t dropped below 1,000 in more than a decade. The tool saves significant time, in particular searching for case law and summarizing lengthy documents. "For research, it’s just one prompt away, whereas before I had to search for the precedents and laws for hours," the judge says. "If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarize it for me and give me the crux, and it does that work in seconds."

But efficiency isn’t the only thing you want out of a justice system, says John Zeleznikow, professor of law and technology at La Trobe University, in Australia. "What they’ve tried to do is be effective, [to] deal with more cases more quickly, and they’re able to do that," he says. "What’s not that clear is whether what you call the quality of justice is better."

Zeleznikow says AI can be useful, but only if judges are diligent about evaluating and verifying the output. However, the working paper’s authors found that roughly a fifth of participants’ prompts given to JudgeGPT involved what they call "substantial AI delegation"—asking the tool what the best decision is, to produce legal reasoning or write opinions with little input from the judge. On the bright side, training lowered the proportion of inappropriate delegation.

But given that judges are already using AI, Ash says better tools and training are crucial. "There are risks for using these AIs, for sure, even with all these safeguards. But at some point you have to just put the judges in as strong a position as you can," he says. "Have technological safeguards, but then try to encourage the judges not to rely on it too much."This story was updated on 13 August 2026 to clarify that the training course was developed in coordination with Pakistan’s Federal Judicial Academy.

>

Читать на сайте источника »