Prompting instruction-tuned LLMs for semantic similarity values

Open Access
Authors
  • Xander Snelder
  • Yunchong Huang
  • J. Bloem ORCID logo
Publication date 2026
Book title The Fifteenth Language Resources and Evaluation Conference (LREC 2026)
Book subtitle Main Conference Proceedings : 13-15 May, 2026
ISBN (electronic)
  • 9782493814494
Event 15th Language Resources and Evaluation Conference
Pages (from-to) 11390-11403
Publisher ELRA Language Resources Association
Organisations
  • Interfacultary Research - Institute for Logic, Language and Computation (ILLC)
Abstract
The impressive few-shot performance of generative decoder transformer language models at novel tasks has raised interest in using them to estimate lexical-semantic properties of words, word pairs or multi-word expressions. We explore the task of eliciting semantic similarity scores between word pairs through prompting, comparing these scores to human benchmarks. We investigate different prompting approaches, different model architectures and different languages using the Dutch, English and Mandarin Chinese SimLex-999 benchmarks. The results show that prompting each word pair individually yields better correlations, and that models struggle with the distinction between similarity and relatedness, just as static and contextual word embedding models did. The new, open-weight gpt-oss-20b model yields the highest correlation with human ratings out of the models we evaluated.
Document type Conference contribution
Note With supplementary slides and video.
Language English
Published at
https://doi.org/10.63317/3kbjxx6989dg (Final published version)
Published at
https://aclanthology.org/2026.lrec-1.891/ (Final published version)
Downloads
2026.lrec2026-1.891 (Final published version)
Supplementary materials
Permalink to this page
Back