A central challenge in neuroscience is decoding brain activity to uncover mental content comprising multiple components and their interactions. Despite progress in decoding language-related information from human brain activity, generating comprehensive descriptions of complex mental content associated with structured visual semantics remains challenging. We present a method that generates descriptive text mirroring brain representations via semantic features computed by a deep language model. Constructing linear decoding models to translate brain activity induced by videos into semantic features of corresponding captions, we optimized candidate descriptions by aligning their features with brain-decoded features through word replacement and interpolation. This process yielded well-structured descriptions that accurately capture viewed content, even without relying on the canonical language network. The method also generalized to verbalize recalled content, functioning as an interpretive interface between mental representations and text and simultaneously demonstrating the potential for nonverbal thought–based brain-to-text communication, which could provide an alternative communication pathway for individuals with language expression difficulties, such as aphasia.
…Fig. 2. Generating viewed content descriptions.
Descriptions were generated using features from all LM layers decoded from whole-brain activity. (A) Evolved descriptions during the optimization (see https://horikawa-t.github.io/MindCaptioningProject/ for more results with the original videos). (B) Descriptions after 100 iterations for all subjects (see fig. S3A for more example). In (A) and (B), the color indicates accuracy [inverse document frequency (IDF)–weighted BERTScore-P]. A reference caption of the video is shown below frames. (C) Feature correlations between features of generated descriptions and those decoded from the brain, as well as those computed from correct references. (D) Cohen’s d of discriminability (see fig. S4B for raw scores). Feat. corr., Feature correlation. (E) Video identification accuracy with varying numbers of candidates. (F) Effects of word-order shuffling on video identification accuracy and discriminability. (G) Scatterplot of the correlation distances (one minus feature correlation) between the original and shuffled descriptions against the difference in feature correlations to target features between original and shuffled descriptions. Each dot indicates a shuffled description. Shades in (C) and (E) and error bars in (D) and (F) indicate 95% confidence intervals (CIs) across samples (n = 72). Shades in (D) and (F) indicate 95% CI across subjects (n = 6). See fig. S4 for individual results.
The content of sensory consciousness gropingly captured by machines but not the consciousness itself. This is no match for FIML depth, artistry or soundness, but interesting nevertheless. It may well be that in the future when machines can do almost everything, the only thing humans will be able to do that machines cannot will be FIML practice. Is it ironic or poetic that the one thing humans have not widely understood before the advent of AI is the one thing AI will never be able to do? (FIML predates AI by many years, but very few understand it even to this day.) One reason, among many, FIML is hard to understand is due to Sapir-Whorf ‘cognitive blindness’, if you will. People can’t see it because no one has ever seen it until relatively recently; it’s not in the general language. FIML is described from many angles in many posts on this website. If you are a smart person, it’s easy to understand once you remove your Sapir-Whorf sunglasses. If you are a bold thinker, go for it. You will come to understand something very deep that few others can see. If you just want to be psychologically much healthier, FIML will do that for you as well. ABN