DS1 spectrogram: Multilingual Universal Sentence Encoder for Semantic Retrieval

Multilingual Universal Sentence Encoder for Semantic Retrieval

1907.04307

Authors

Chris Tar,Mandy Guo,Yun-Hsuan Sung,Brian Strope,Ray Kurzweil

Abstract

We introduce two pre-trained retrieval focused multilingual sentence encoding models, respectively based on the Transformer and CNN model architectures. The models embed text from 16 languages into a single semantic space using a multi-task trained dual-encoder that learns tied representations using translation based bridge tasks (Chidambaram al., 2018).

The models provide performance that is competitive with the state-of-the-art on: semantic retrieval (SR), translation pair bitext retrieval (BR) and retrieval question answering (ReQA). On English transfer learning tasks, our sentence-level embeddings approach, and in some cases exceed, the performance of monolingual, English only, sentence embedding models.

Our models are made available for download on TensorFlow Hub.

Resources

Stay in the loop

Every AI paper that matters, free in your inbox daily.

Details

  • takara.ai
  • Custom AI and machine learning from the Frontier Research Team.
  • © 2026 takara.ai Ltd
  • Content is sourced from third-party publications.