Prompting instruction-tuned LLMs for semantic similarity values
| Authors |
|
|---|---|
| Publication date | 2026 |
| Book title | The Fifteenth Language Resources and Evaluation Conference (LREC 2026) |
| Book subtitle | Main Conference Proceedings : 13-15 May, 2026 |
| ISBN (electronic) |
|
| Event | 15th Language Resources and Evaluation Conference |
| Pages (from-to) | 11390-11403 |
| Publisher | ELRA Language Resources Association |
| Organisations |
|
| Abstract |
The impressive few-shot performance of generative decoder transformer language models at novel tasks has raised interest in using them to estimate lexical-semantic properties of words, word pairs or multi-word expressions. We explore the task of eliciting semantic similarity scores between word pairs through prompting, comparing these scores to human benchmarks. We investigate different prompting approaches, different model architectures and different languages using the Dutch, English and Mandarin Chinese SimLex-999 benchmarks. The results show that prompting each word pair individually yields better correlations, and that models struggle with the distinction between similarity and relatedness, just as static and contextual word embedding models did. The new, open-weight gpt-oss-20b model yields the highest correlation with human ratings out of the models we evaluated.
|
| Document type | Conference contribution |
| Note | With supplementary slides and video. |
| Language | English |
| Published at |
https://doi.org/10.63317/3kbjxx6989dg
(Final published version)
|
| Published at |
https://aclanthology.org/2026.lrec-1.891/
(Final published version)
|
| Downloads |
2026.lrec2026-1.891
(Final published version)
|
| Supplementary materials | |
| Permalink to this page | |
