Supplementary Table S2 and Table S3 show essentially consistent results for Rouge2 and RougeL, except the positive correlation between RSA scores and Rouge2 decline for the T5 model in Full time window narrowly missed statistical significance ( p = 0.051).
← all excerpts
Brain-model neural similarity reveals abstractive summarization performance.
5
0.0510
0.1000
The sentences
The BART model showed a statistically significant monotonic increase in RSA scores for Full time and Early window ( Z > 0, p < 0.05), but for Late window, the increasing trend did not reach statistical significance ( Z = 1.643, p = 0.100).
Conversely, the RS between PEGASUS and T5 hidden layers exhibited a decreasing trend with increasing layer depth.
We observed a notable trend: as the model layers deepened, the similarity between encoder layers and human brain language processing generally showed an upward trend, reaching statistical significance ( p < 0.05) under most experimental conditions.
The representational similarity between BART and T5 hidden layers showed an increasing trend as the layer depth increased.