Mind captioning: Evolving descriptive text of mental content from human brain activity

…Fig. 2. Generating viewed content descriptions.

Descriptions were generated using features from all LM layers decoded from whole-brain activity. (A) Evolved descriptions during the optimization (see https://horikawa-t.github.io/MindCaptioningProject/ for more results with the original videos). (B) Descriptions after 100 iterations for all subjects (see fig. S3A for more example). In (A) and (B), the color indicates accuracy [inverse document frequency (IDF)–weighted BERTScore-P]. A reference caption of the video is shown below frames. (C) Feature correlations between features of generated descriptions and those decoded from the brain, as well as those computed from correct references. (D) Cohen’s d of discriminability (see fig. S4B for raw scores). Feat. corr., Feature correlation. (E) Video identification accuracy with varying numbers of candidates. (F) Effects of word-order shuffling on video identification accuracy and discriminability. (G) Scatterplot of the correlation distances (one minus feature correlation) between the original and shuffled descriptions against the difference in feature correlations to target features between original and shuffled descriptions. Each dot indicates a shuffled description. Shades in (C) and (E) and error bars in (D) and (F) indicate 95% confidence intervals (CIs) across samples (n = 72). Shades in (D) and (F) indicate 95% CI across subjects (n = 6). See fig. S4 for individual results.

link


Leave a comment