TRIBE v2 (Meta FAIR tri-modal brain foundation model)
Meta AI's model that predicts fMRI cortical responses to video stimuli, used as the engine behind VidCognition.
TRIBE v2 is a foundation model of vision, audition and language developed by Meta AI Research (FAIR) that predicts human brain activity in response to naturalistic stimuli — second by second, across the cortex — using fMRI as its training signal. It is tri-modal: it combines video, audio and language.
TRIBE v2 was trained on a unified dataset of over 1,000 hours of fMRI across 720 subjects. It learned to predict cortical activity patterns — which regions activate, and at what intensity — for novel stimuli, tasks and subjects, delivering several-fold improvements in accuracy over traditional linear encoding models. The TRIBE line ranked 1st of 263 teams in the Algonauts 2025 brain-encoding challenge, with a mean score of 0.2146 ± 0.0312 (d'Ascoli et al., arXiv:2605.04326, May 2026).
The practical implication: TRIBE v2 can simulate a brain response to a video without putting a person in an MRI scanner. That makes a class of neuromarketing signal available outside a laboratory — but it is a simulation, and it is directional rather than a prediction of views, ROAS or revenue.
VidCognition is built on TRIBE v2, making it the first creator-accessible tool to use this model for short-form video optimization.
Try it with VidCognition
Related Terms
fMRI (Functional Magnetic Resonance Imaging)
A brain imaging technique that measures neural activity by detecting blood flow changes in different brain regions.
Neural Encoding
The process by which the brain converts sensory input into neural activity patterns.
Brain Engagement Score
A composite score predicting the intensity of neural activation across key brain regions during video viewing.