Pretraining vision-language models on developmentally grounded data inspired by infant cognition and learning.
The BabyVLM Challenge is a shared task run in partnership with the BabyLM Challenge, focused on developmentally plausible, sample-efficient vision-language modeling. Participants work with new datasets and evaluations purpose-built for small vision-language models trained on limited data.
The challenge will be introduced at the BabyVLM workshop with a tutorial. No prior BabyLM experience is needed to participate.
Pretrain on a compact dataset of egocentric audiovisual streams, provided below.
Report scores on the 10 subtasks of the DevCV Toolbox.
Coming soon!
All deadlines are 11:59 PM AoE (Anywhere on Earth).
Download the training data from the Pretraining section below, and pretrain your vision-language model.
If feasible, download the train split from the Evaluation section below, and finetune your model on the specific task formats.
Follow the link below for further submission instructions.
🤗 BabyVLM on Hugging Face
[Placeholder: describe the pretraining data — sources, scale, how to download, and any usage restrictions.]
Our shared task is the DevCV Toolbox, modeled directly after a subset of the tasks in the NIH Baby Toolbox. The DevCV Toolbox consists of 10 subtasks, each created from four data sources: original Toolbox tasks (NIH), SAYCam, BabyView, and Ego4D.
Confirmed advisors guiding the scientific direction of the challenge.
Who to contact with any technical questions or suggestions.
[Placeholder: list technical committee members and contact info.]
Email us at wsashawn@bu.edu, wqwang@bu.edu, maxwh@bu.edu, or bgong@bu.edu.