TRIBE v2 (Meta FAIR tri-modal brain foundation model)

Meta AI's model that predicts fMRI cortical responses to video stimuli, used as the engine behind VidCognition.

TRIBE v2 is a foundation model of vision, audition and language developed by Meta AI Research (FAIR) that predicts human brain activity in response to naturalistic stimuli — second by second, across the cortex — using fMRI as its training signal. It is tri-modal: it combines video, audio and language.

TRIBE v2 was trained on a unified dataset of over 1,000 hours of fMRI across 720 subjects. It learned to predict cortical activity patterns — which regions activate, and at what intensity — for novel stimuli, tasks and subjects, delivering several-fold improvements in accuracy over traditional linear encoding models. The TRIBE line ranked 1st of 263 teams in the Algonauts 2025 brain-encoding challenge, with a mean score of 0.2146 ± 0.0312 (d'Ascoli et al., arXiv:2605.04326, May 2026).

The practical implication: TRIBE v2 can simulate a brain response to a video without putting a person in an MRI scanner. That makes a class of neuromarketing signal available outside a laboratory — but it is a simulation, and it is directional rather than a prediction of views, ROAS or revenue.

VidCognition is built on TRIBE v2, making it the first creator-accessible tool to use this model for short-form video optimization.

Try it with VidCognition