Portrait of Divyanshu Mishra

Multimodal foundation models for video

Divyanshu Mishra

Research Scientist II Amazon Science

Visiting Researcher University of Oxford

My current work at Amazon Science focuses on post-training foundation models and building reliable reasoning systems. I am particularly interested in supervised fine-tuning, reinforcement learning, and multimodal models that can understand video.

I completed my DPhil in Engineering Science at the University of Oxford, supervised by Professor Alison Noble and supported by the Athena-Bronze Scholarship. My doctoral research centered on long-video understanding, including temporal localization, self-supervised representation learning, and multimodal learning. I developed these methods primarily for fetal ultrasound, with the goal of supporting earlier and more reliable detection of congenital heart disease.

Before joining Amazon full-time, I interned at Amazon Science as an Applied Scientist II. There, I developed text-only training strategies for Video-LLMs to reduce their dependence on large paired video-text datasets.

Before my DPhil, I was a Data Scientist at THSTI, Government of India, where I worked under Professor Shinjini Bhatnagar. Together with clinicians and public health researchers, I developed models for preterm birth prediction, gestational-age estimation, and privacy-preserving ultrasound.

Across these roles, the common thread in my work is learning effectively from limited supervision and translating research into systems that remain useful in real-world settings. I am always happy to connect about research, collaborations, and new ideas.

Research

Current focus

01

Foundation-model post-training

Supervised fine-tuning, preference optimization, and reinforcement learning for multimodal models.

02

Long-video understanding

Temporal grounding, representation learning, and world models for long, unstructured video.

03

AI for healthcare

Reliable multimodal systems for clinical imaging, early detection, and decision support.

Experience

7+ years in AI research

Seven years spanning applied research in industry and doctoral work at Oxford.

Research Scientist II

Amazon Science · Seattle

Visiting Researcher

University of Oxford

DPhil (PhD), Engineering Science

University of Oxford · Noble Lab

Applied Scientist II Intern

Amazon Science · Berlin

Data Scientist, Computer Vision

THSTI · Government of India

What’s new

Recent updates

New preprint on externally validated deep learning models for spontaneous preterm birth prediction.

Read

New preprint on selecting task-specialized models as tools for agentic healthcare systems.

Read

Joined Amazon Science as a Research Scientist II after completing my DPhil at Oxford.

Research output

First- & co-first-author papers

View all on Scholar

* Equal contribution

medRxiv 2026

Development and External Validation of Deep Learning Models for Spontaneous Preterm Birth Prediction from Mid-Trimester Cervical Ultrasound

Radhika Chanian*, Divyanshu Mishra*, Rahul Jain, Nikhil Sharma, Ashok Khurana, Reva Tripathi, Abhinav Jain, GARBH-Ini study group, Nitya Wadhwa, J. Alison Noble, Ramachandran Thiruvengadam, Bapu Koundinya Desiraju, and Shinjini Bhatnagar

Manuscript under review

Grounding in the Dark: Investigating Text-Only Video-LLM Training Strategies for Zero-Shot Video Temporal Localization

Divyanshu Mishra, S. Sternig, R. Shetty, and Erhan Gundogdu

DISCOVR method overview

NeurIPS 2025

Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation

Divyanshu Mishra, Mohammadreza Salehi, Pramit Saha, Olga Patey, Aris T. Papageorghiou, Yuki M. Asano, and J. Alison Noble

STUD and DiVMerge method overview

MICCAI 2025 · Co-first author

Self-supervised Normality Learning and Divergence Vector-guided Model Merging for Zero-shot Congenital Heart Disease Detection in Fetal Ultrasound Videos

Pramit Saha*, Divyanshu Mishra*, Netzahualcoyotl Hernandez-Cruz, Olga Patey, Aris T. Papageorghiou, Yuki M. Asano, and J. Alison Noble

TIER-LOC method overview

Medical Image Analysis 2025

TIER-LOC: Visual Query-based Video Clip Localization in Fetal Ultrasound Videos with a Multi-Tier Transformer

Divyanshu Mishra, Pramit Saha, He Zhao, Netzahualcoyotl Hernandez-Cruz, Olga Patey, Aris T. Papageorghiou, and J. Alison Noble

Medical Image Analysis 2025 · Co-first author

HarmonicEchoNet: Leveraging Harmonic Convolutions for Automated Standard Plane Detection in Fetal Heart Ultrasound Videos

Md Mostafa Kamal Sarker*, Divyanshu Mishra*, Mohammad Alsharid*, Netzahualcoyotl Hernandez-Cruz, Rahul Ahuja, Olga Patey, Aris T. Papageorghiou, and J. Alison Noble

MCAT method overview

AAAI 2025

MCAT: Visual Query-Based Localization of Standard Anatomical Clips in Fetal Ultrasound Videos using Multi-Tier Class-Aware Token Transformer

Divyanshu Mishra, Pramit Saha, He Zhao, Netzahualcoyotl Hernandez-Cruz, Olga Patey, Aris Papageorghiou, and J. Alison Noble

medRxiv 2024

Development and External Validation of an Ultrasound Image-Based Deep Learning Model to Estimate Gestational Age in the Second and Third Trimesters of Pregnancy Using Data from the GARBH-Ini Cohort: A Prospective Cohort Study in North Indian Population

Divyanshu Mishra, Varun Chandramohan, Nikhil Sharma, Mudita Gosain, Nitya Wadhwa, Uma Chandra Mouli Natchu, GARBH-Ini study group, Ashok Khurana, J. Alison Noble, Ramachandran Thiruvengadam, Bapu Koundinya Desiraju, and Shinjini Bhatnagar

STAN-LOC method overview

MICCAI 2024

STAN-LOC: Visual Query-Based Video Clip Localization for Fetal Ultrasound Sweep Videos

Divyanshu Mishra, Pramit Saha, He Zhao, Olga Patey, Aris T. Papageorghiou, and J. Alison Noble

Dual Conditioned Diffusion Models method overview

MICCAI 2023

Dual Conditioned Diffusion Models for Out-of-Distribution Detection: Application to Fetal Ultrasound Videos

Divyanshu Mishra, He Zhao, Pramit Saha, Aris T. Papageorghiou, and J. Alison Noble