In conclusion, although there might be significant room for improvement in the image embeddings, and limited room for improvement in the text embeddings, when considered together, there is just an insignificant gap to enhance the multimodal model, making the insights from this study sufficiently reliable.
← all excerpts
Anime popularity prediction before huge investments: a multimodal approach using deep learning.
1
—
—