BabyVLM logo

BabyVLM Challenge

Pretraining vision-language models on developmentally grounded data inspired by infant cognition and learning.

Submit Results Explore Pretraining Explore Task

Overview

The BabyVLM Challenge is a shared task run in partnership with the BabyLM Challenge, focused on developmentally plausible, sample-efficient vision-language modeling. Participants work with new datasets and evaluations purpose-built for small vision-language models trained on limited data.

The challenge will be introduced at the BabyVLM workshop with a tutorial. No prior BabyLM experience is needed to participate.

🎬

Pretraining

Pretrain on a compact dataset of egocentric audiovisual streams, provided below.

🧠

Task

Report scores on the 10 subtasks of the DevCV Toolbox.

📅

Deadline

Coming soon!

Dates & Deadlines

All deadlines are 11:59 PM AoE (Anywhere on Earth).

December 2026
Challenge Announced Done
Dataset and evaluation server released at NeurIPS 2026.
TBD
Final Submission Deadline Coming Soon
No submissions accepted after this date.
June 2027
Results Announced
Winners notified and results presented at CVPR 2027.
June 2027
Workshop & Paper Presentations
Top teams present at the BabyVLM workshop at CVPR 2027.

Submission Instructions

Download the training data

Download the training data from the Pretraining section below, and pretrain your vision-language model.

(Optional) Instruction Fine-tuning

If feasible, download the train split from the Evaluation section below, and finetune your model on the specific task formats.

Submit your model on Hugging Face

Follow the link below for further submission instructions.

🤗 BabyVLM on Hugging Face

Pretraining

[Placeholder: describe the pretraining data — sources, scale, how to download, and any usage restrictions.]

📦 Access the full dataset — link coming soon 📄 Complete methodology — BabyVLM-V2

Task

Our shared task is the DevCV Toolbox, modeled directly after a subset of the tasks in the NIH Baby Toolbox. The DevCV Toolbox consists of 10 subtasks, each created from four data sources: original Toolbox tasks (NIH), SAYCam, BabyView, and Ego4D.

📦 Access the full dataset — link coming soon 📄 Complete methodology — BabyVLM-V2
Sources & Their Role
Task
Samples

Challenge Steering Committee

Confirmed advisors guiding the scientific direction of the challenge.

Michael C. Frank
Michael C. Frank
Psychology, Stanford
CY
Chen Yu
Psychology, UT Austin
Kristen Grauman
Kristen Grauman
CS, UT Austin
KS
Kate Saenko
CS, Boston University

Technical Committee

Who to contact with any technical questions or suggestions.

[Placeholder: list technical committee members and contact info.]

Frequently Asked Questions

How do I contact the organizers?