Changelog

Changes to the dataset, derived files, and Python package.

2026-09-28

Corrected per-subject random splits

Updated the subject-specific random_0–random_4 splits to keep identified duplicate images in the same fold, preserving split names and fold sizes. Update the Python package to use the corrected splits. The previous files are archived for reproducing earlier results.

Across-subject train/test splits

Added pool="pooled" with tau and five cluster splits for consistent image assignments across subjects. See across-subject splits for usage and construction.

Corrected stimulus embedding transparency

Updated the CLIP, DINOv2, PEcore, and SigLIP2 embedding files for 116 out-of-distribution stimuli with non-opaque pixels. Their original embeddings discarded transparency. The corrected embeddings composite the images onto the experiment’s middle-grey background, RGB (128, 128, 128), before model-specific resizing and cropping.

The remaining 24,936 embedding rows in each file are unchanged, as are the image IDs, row order, storage dtypes, and file layout. The stimulus images themselves have not changed.

The corrected files are available at the existing download URLs. To refresh an existing download, close any open embedding handles and run:

import laion_fmri

laion_fmri.download_embeddings("all")

Then reload the embeddings. This also works with older package versions: the corrected files differ in size from the originals, so the downloader replaces the cached files. No package upgrade is required for this data correction.

For reproducibility, the original files are preserved in the archive:

See Stimulus Derivatives for the embedding models and file layout.