Как медицинская база данных, разработанная в Массачусетском технологическом институте, превратилась в глобальный стандарт обмена данными

• ПО
29.07.2026

До появления технологий хранения научных данных и совместной работы в облаке исследователям-медикам, стремящимся к прорывным исследованиям в области здравоохранения, приходилось преодолевать значительные препятствия на пути сотрудничества и сбора ключевых клинических данных.

Данные были разрозненными и их было трудно распространять, поэтому у тех, кто хотел провести исследование, не было другого выбора, кроме как собирать их самостоятельно. Это не только удорожало исследования, но и затрудняло сравнение результатов в разных наборах данных.

В 1975 году исследователи, изучающие аритмии в Массачусетском технологическом институте и Бостонской больнице Бет Исраэль, придумали другой способ: команда начала собирать и оцифровывать записи электрокардиограмм с намерением не только изучить их, но и сделать доступными для более широкого исследовательского сообщества.

Команда разработала собственные компьютеры для этого процесса, тщательно продублировала кассеты одну за другой и создала более 100 000 аннотаций к записям. Процесс занял годы, но к лету 1980 года кассеты были, наконец, готовы. Изначально команда думала, что их инструмент охватит не более десятка академических и промышленных групп. Но интерес продолжал расти. В течение следующего десятилетия они разослали по почте около 100 копий.

В конечном итоге эти данные стали первой базой данных глобальной платформы PhysioNet, основанной в 1999 году в рамках программы Гарварда и Массачусетского технологического института в области медицинских наук и технологий, в качестве хранилища клинических данных о сложных физиологических сигналах.

В то время такой способ обмена данными, который сегодня может показаться стандартным, был почти революционной идеей. "Основатель PhysioNet был невероятно дальновиден", - говоритТомас Хелдт, Ричард Дж. Коэн (1976), профессор медицины и биомедицинской физики, заместитель директора Института медицинской инженерии и науки Массачусетского технологического института и старший автор недавней статьи в журнале Nature Health, посвященной изучению влияния платформы.

В конце концов, эти магнитные ленты, отправленные по почте, превратились в записанные компакт-диски, которые затем превратились в FTP-серверы, размещенные в недавно созданном Интернете. Сегодня, когда PhysioNet работает уже более 25 лет, платформа содержит сотни баз данных и стала одним из самых полных хранилищ биомедицинских и клинических данных из существующих. В прошлом году на PhysioNet ссылались более 15 000 научных публикаций, а пользователи из более чем 180 стран зарегистрировались на платформе. Она широко используется исследователями, производителями и лицами, принимающими клинические решения.

"Результаты исследований действительно значительны, - говорит Хелдт, который также является профессором кафедры электротехники и компьютерных наук Массачусетского технологического института и главным научным сотрудником Исследовательской лаборатории электроники, - и весьма впечатляющи".

"It is really beautiful to see that such a vision has proven right and so enabling for so many people."

Setting a standard

Around 2009, a PhD student named Tom Pollard was conducting research on critically ill patients at one of London’s leading hospital systems. Although the hospital generated large volumes of valuable clinical data, the infrastructure and processes needed to curate and support their wider research use were still developing.

"Hospital data were collected primarily to support immediate patient care, with less attention given to how they might be curated and reused for research," says Pollard, now a research scientist at MIT’s Laboratory for Computational Physiology (LCP), technical director of PhysioNet, and the lead author on the Nature Health paper.

The problem was not simply privacy. Hospital information systems were built primarily to support patient care and administration, not research. Data were fragmented across systems and rarely curated with future reuse in mind, making it difficult and expensive to turn them into coherent research resources.

But Pollard needed data to complete his dissertation. After poking around on the internet, he eventually discovered the Medical Information Mart for Intensive Care (MIMIC), a database of de-identified electronic health records hosted by PhysioNet. Recognizing its potential, his clinical supervisor, Kevin Fong, organized a visit to Boston. Soon afterward, Fong and Pollard were sitting across the table from Roger Mark, discussing how their teams might collaborate.

Academic incentives have long favored publications and exclusive analyses over the less-visible work involved in preparing data for others to use. That tension persists today. PhysioNet’s founders embraced a different model, believing that sharing research resources could accelerate discovery and ultimately improve human health, he says. MIMIC became central to Pollard’s dissertation, and after completing his PhD, he came to MIT to help build the next generation of the database.

In the years since PhysioNet was established, the value of sharing research data has gained much wider recognition. The late Roger Mark, MIT’s distinguished professor of health sciences and technology emeritus and one of PhysioNet’s founders, described its purpose as building an "accessible multinational community around data" to "positively impact global health."

Earlier this year, Mark and the late George Moody, PhysioNet’s co-founder, jointly received the prestigious IEEE Biomedical Engineering Award for their contributions to PhysioNet and biomedical signal processing. IEEE cited their "leadership in ECG signal processing and global dissemination of curated biomedical and clinical databases, thereby accelerating biomedical research worldwide."

The source code for the platform, like much of its data, is public. According to the Nature piece: "As the platform evolved, PhysioNet’s community broadened substantially beyond its origins in signal processing and cardiovascular health to encompass clinical informatics, critical care and machine learning for health." People have used that to build their own PhysioNet-esque infrastructure, says Heldt. Pollard points to similar platforms like Health Data Nexus as examples of PhysioNet’s legacy.

Although there are now more resources out there hosting similar electronic health data, according to Google DeepMind researcher Vivek Natarajan, both PhysioNet and MIMIC "set the standard," he says, "and it’s still the standard right now."

That standard, according to those who use the platform, changed how research is conducted. Access to data should not be the determinant for which ideas are possible, according to Ziad Obermeyer, an associate professor at the University of California at Berkeley School of Public Health and the College of Computing, Data Science, and Society.

"PhysioNet changed how I think about the bottleneck in research. It is often not ideas or talent. It is friction. When access to data is slow, expensive, and hard, the ideas that die first are the high-risk ones, the things that probably will not work, but would be transformative if they did. That is exactly the wrong model if you want real progress," he says. "PhysioNet lowers the fixed cost of trying ambitious ideas, and that changes what science becomes possible."

The AI boom

PhysioNet, once a repository mainly for those working in biomedical signal processing and the health-care fields, has evolved in its 25 years. Originally, the holdings consisted solely of cardiovascular ECG data. Now PhysioNet is a largely a source for electronic health records, imaging data, and software and AI models.

Particularly as artificial intelligence approaches took off, "the community shifted," Heldt explains. Those in need of signal processing data still use PhysioNet databases, but the pool of users has expanded to encompass staff at large tech companies, teachers, and practitioners in all areas of medicine, as well as researchers in health-related machine learning and AI. Today, that latter group "dominates the user community," says Heldt.

The platform hosts the highest-quality datasets available for health-care AI research, according to Natarajan, whose research involves AI, science, and medicine and who has published several papers that used its datasets.

"It has been an important cornerstone that has catalyzed all the progress in health-care AI over the last decade," says Natarajan. In addition to using PhysioNet data, he and his colleagues have contributed data to the platform, helping create the self-sustaining ecosystem that typifies PhysioNet.

Looking toward the coming decades, stewards of the platform like Heldt and Pollard envision continuing to expand its reach with an annual conference. The team is also preparing to pilot a new system that will allow users to annotate data and contribute their own expertise, enriching PhysioNet’s resources for the next phase of the platform.

"The kind of research that people want to do now needs to be interdisciplinary. Statisticians, computer scientists, clinicians, pharmacists, and nurses must all come together and contribute their knowledge to develop algorithms that are useful for people" says Pollard. "The community has broadened, and advances in AI have expanded both the questions researchers can address and what they believe is possible."

>

Читать на сайте источника »