People

Alan Akbik

Principal Investigator

Machine Learning

HU Berlin

 

Email:

 

Photo: SCIoI

← People Overview

Alan Akbik

Alan Akbik

Photo: SCIoI

Alan focuses on research in machine learning (ML) and natural language processing (NLP), with the goal of giving machines the ability to understand and use human language. This spans research topics such as neural language modeling, sample-efficient learning and semantic parsing, as well as application areas in large-scale text analytics. Together with his group and the open source community, he develops the NLP framework Flair (https//github.com/flairNLP/flair) that allows anyone to use state-of-the-art NLP methods in their research or applications. At SCIoI, Alan works at Project A002, Project 44, and Project 45.


Projects

Alan Akbik is member of:


6984777 akbik 1 apa 50 default 19781 https://www.scienceofintelligence.de/wp-content/plugins/zotpress/
%7B%22status%22%3A%22success%22%2C%22updateneeded%22%3Afalse%2C%22instance%22%3Afalse%2C%22meta%22%3A%7B%22request_last%22%3A0%2C%22request_next%22%3A0%2C%22used_cache%22%3Atrue%7D%2C%22data%22%3A%5B%7B%22key%22%3A%22988BM34J%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Ziletti%20et%20al.%22%2C%22parsedDate%22%3A%222022%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BZiletti%2C%20A.%2C%20Akbik%2C%20A.%2C%20Berns%2C%20C.%2C%20Herold%2C%20T.%2C%20Legler%2C%20M.%2C%20%26amp%3B%20Viell%2C%20M.%20%282022%29.%20Medical%20Coding%20with%20Biomedical%20Transformer%20Ensembles%20and%20Zero%5C%2FFew-shot%20Learning.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202022%20Conference%20of%20the%20North%20American%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20Human%20Language%20Technologies%3A%20Industry%20Track%26lt%3B%5C%2Fi%26gt%3B%2C%20176%26%23x2013%3B187.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2022.naacl-industry.21%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2022.naacl-industry.21%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Medical%20Coding%20with%20Biomedical%20Transformer%20Ensembles%20and%20Zero%5C%2FFew-shot%20Learning%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Angelo%22%2C%22lastName%22%3A%22Ziletti%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Christoph%22%2C%22lastName%22%3A%22Berns%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Thomas%22%2C%22lastName%22%3A%22Herold%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Marion%22%2C%22lastName%22%3A%22Legler%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Martina%22%2C%22lastName%22%3A%22Viell%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202022%20Conference%20of%20the%20North%20American%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20Human%20Language%20Technologies%3A%20Industry%20Track%22%2C%22conferenceName%22%3A%22Proceedings%20of%20the%202022%20Conference%20of%20the%20North%20American%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20Human%20Language%20Technologies%3A%20Industry%20Track%22%2C%22date%22%3A%222022%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2022.naacl-industry.21%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2022.naacl-industry.21%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A29Z%22%7D%7D%2C%7B%22key%22%3A%22V3ESZKYN%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Golde%20et%20al.%22%2C%22parsedDate%22%3A%222024-03%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BGolde%2C%20J.%2C%20Hamborg%2C%20F.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20Large-Scale%20Label%20Interpretation%20Learning%20for%20Few-Shot%20Named%20Entity%20Recognition.%20In%20Y.%20Graham%20%26amp%3B%20M.%20Purver%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%2018th%20Conference%20of%20the%20European%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%20%28EACL%202024%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%202915%26%23x2013%3B2930%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.eacl-long.178%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.eacl-long.178%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Large-Scale%20Label%20Interpretation%20Learning%20for%20Few-Shot%20Named%20Entity%20Recognition%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Jonas%22%2C%22lastName%22%3A%22Golde%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Felix%22%2C%22lastName%22%3A%22Hamborg%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yvette%22%2C%22lastName%22%3A%22Graham%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Matthew%22%2C%22lastName%22%3A%22Purver%22%7D%5D%2C%22abstractNote%22%3A%22Few-shot%20named%20entity%20recognition%20%28NER%29%20detects%20named%20entities%20within%20text%20using%20only%20a%20few%20annotated%20examples.%20One%20promising%20line%20of%20research%20is%20to%20leverage%20natural%20language%20descriptions%20of%20each%20entity%20type%3A%20the%20common%20label%20PER%20might%2C%20for%20example%2C%20be%20verbalized%20as%20%5Cu201cperson%20entity.%5Cu201d%20In%20an%20initial%20label%20interpretation%20learning%20phase%2C%20the%20model%20learns%20to%20interpret%20such%20verbalized%20descriptions%20of%20entity%20types.%20In%20a%20subsequent%20few-shot%20tagset%20extension%20phase%2C%20this%20model%20is%20then%20given%20a%20description%20of%20a%20previously%20unseen%20entity%20type%20%28such%20as%20%5Cu201cmusic%20album%5Cu201d%29%20and%20optionally%20a%20few%20training%20examples%20to%20perform%20few-shot%20NER%20for%20this%20type.%20In%20this%20paper%2C%20we%20systematically%20explore%20the%20impact%20of%20a%20strong%20semantic%20prior%20to%20interpret%20verbalizations%20of%20new%20entity%20types%20by%20massively%20scaling%20up%20the%20number%20and%20granularity%20of%20entity%20types%20used%20for%20label%20interpretation%20learning.%20To%20this%20end%2C%20we%20leverage%20an%20entity%20linking%20benchmark%20to%20create%20a%20dataset%20with%20orders%20of%20magnitude%20of%20more%20distinct%20entity%20types%20and%20descriptions%20as%20currently%20used%20datasets.%20We%20find%20that%20this%20increased%20signal%20yields%20strong%20results%20in%20zero-%20and%20few-shot%20NER%20in%20in-domain%2C%20cross-domain%2C%20and%20even%20cross-lingual%20settings.%20Our%20findings%20indicate%20significant%20potential%20for%20improving%20few-shot%20NER%20through%20heuristical%20data-based%20optimization.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%2018th%20Conference%20of%20the%20European%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%20%28EACL%202024%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024-03%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.eacl-long.178%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.eacl-long.178%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A16Z%22%7D%7D%2C%7B%22key%22%3A%22ETMQZEIB%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Merdjanovska%20et%20al.%22%2C%22parsedDate%22%3A%222024-11%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BMerdjanovska%2C%20E.%2C%20Aynetdinov%2C%20A.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20NoiseBench%3A%20Benchmarking%20the%20Impact%20of%20Real%20Label%20Noise%20on%20Named%20Entity%20Recognition.%20In%20Y.%20Al-Onaizan%2C%20M.%20Bansal%2C%20%26amp%3B%20Y.-N.%20Chen%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202024%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2018182%26%23x2013%3B18198%29.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.emnlp-main.1011%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.emnlp-main.1011%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22NoiseBench%3A%20Benchmarking%20the%20Impact%20of%20Real%20Label%20Noise%20on%20Named%20Entity%20Recognition%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Elena%22%2C%22lastName%22%3A%22Merdjanovska%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Ansar%22%2C%22lastName%22%3A%22Aynetdinov%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yaser%22%2C%22lastName%22%3A%22Al-Onaizan%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Mohit%22%2C%22lastName%22%3A%22Bansal%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yun-Nung%22%2C%22lastName%22%3A%22Chen%22%7D%5D%2C%22abstractNote%22%3A%22Available%20training%20data%20for%20named%20entity%20recognition%20%28NER%29%20often%20contains%20a%20significant%20percentage%20of%20incorrect%20labels%20for%20entity%20types%20and%20entity%20boundaries.%20Such%20label%20noise%20poses%20challenges%20for%20supervised%20learning%20and%20may%20significantly%20deteriorate%20model%20quality.%20To%20address%20this%2C%20prior%20work%20proposed%20various%20noise-robust%20learning%20approaches%20capable%20of%20learning%20from%20data%20with%20partially%20incorrect%20labels.%20These%20approaches%20are%20typically%20evaluated%20using%20simulated%20noise%20where%20the%20labels%20in%20a%20clean%20dataset%20are%20automatically%20corrupted.%20However%2C%20as%20we%20show%20in%20this%20paper%2C%20this%20leads%20to%20unrealistic%20noise%20that%20is%20far%20easier%20to%20handle%20than%20real%20noise%20caused%20by%20human%20error%20or%20semi-automatic%20annotation.%20To%20enable%20the%20study%20of%20the%20impact%20of%20various%20types%20of%20real%20noise%2C%20we%20introduce%20NoiseBench%2C%20an%20NER%20benchmark%20consisting%20of%20clean%20training%20data%20corrupted%20with%206%20types%20of%20real%20noise%2C%20including%20expert%20errors%2C%20crowdsourcing%20errors%2C%20automatic%20annotation%20errors%20and%20LLM%20errors.%20We%20present%20an%20analysis%20that%20shows%20that%20real%20noise%20is%20significantly%20more%20challenging%20than%20simulated%20noise%2C%20and%20show%20that%20current%20state-of-the-art%20models%20for%20noise-robust%20learning%20fall%20far%20short%20of%20their%20achievable%20upper%20bound.%20We%20release%20NoiseBench%20for%20both%20English%20and%20German%20to%20the%20research%20community.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202024%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%22%2C%22conferenceName%22%3A%22EMNLP%202024%22%2C%22date%22%3A%222024-11%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.emnlp-main.1011%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.emnlp-main.1011%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A16Z%22%7D%7D%2C%7B%22key%22%3A%22FRJM6NEG%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Pohl%20et%20al.%22%2C%22parsedDate%22%3A%222025-08%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BPohl%2C%20S.%2C%20Ploner%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20Towards%20a%20Principled%20Evaluation%20of%20Knowledge%20Editors.%20In%20R.%20Jia%2C%20E.%20Wallace%2C%20Y.%20Huang%2C%20T.%20Pimentel%2C%20P.%20Maini%2C%20V.%20Dankers%2C%20J.%20Wei%2C%20%26amp%3B%20P.%20Lesci%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%20First%20Workshop%20on%20Large%20Language%20Model%20Memorization%20%28L2M2%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2047%26%23x2013%3B60%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.l2m2-1.4%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.l2m2-1.4%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Towards%20a%20Principled%20Evaluation%20of%20Knowledge%20Editors%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Sebastian%22%2C%22lastName%22%3A%22Pohl%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Robin%22%2C%22lastName%22%3A%22Jia%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Eric%22%2C%22lastName%22%3A%22Wallace%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yangsibo%22%2C%22lastName%22%3A%22Huang%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Tiago%22%2C%22lastName%22%3A%22Pimentel%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Pratyush%22%2C%22lastName%22%3A%22Maini%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Verna%22%2C%22lastName%22%3A%22Dankers%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Johnny%22%2C%22lastName%22%3A%22Wei%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Pietro%22%2C%22lastName%22%3A%22Lesci%22%7D%5D%2C%22abstractNote%22%3A%22Model%20editing%20has%20been%20gaining%20increasing%20attention%20over%20the%20past%20few%20years.%20For%20Knowledge%20Editing%20in%20particular%2C%20more%20challenging%20evaluation%20datasets%20have%20recently%20been%20released.%20These%20datasets%20use%20different%20methodologies%20to%20score%20the%20success%20of%20editors.%20Yet%2C%20it%20remains%20under-explored%20how%20robust%20these%20methodologies%20are%20and%20whether%20they%20unfairly%20favor%20some%20editors.%20Moreover%2C%20the%20disruptive%20impact%20of%20these%20editors%20on%20overall%20model%20capabilities%20remains%20a%20constant%20blind%20spot.We%20address%20both%20of%20these%20problems%20and%20show%20that%20choosing%20different%20metrics%20and%20evaluation%20methodologies%20as%20well%20as%20different%20edit%20batch%20sizes%20can%20lead%20to%20a%20different%20ranking%20of%20knowledge%20editors.%20Crucially%20we%20demonstrate%20this%20effect%20also%20on%20general%20language%20understanding%20tasks%20evaluated%20alongside%20the%20knowledge%20editing%20tasks.%20Further%20we%20include%20a%20manual%20assessment%20of%20the%20string%20matching%20based%20evaluation%20method%20for%20knowledge%20editing%20that%20is%20favored%20by%20recently%20released%20datasets%2C%20revealing%20a%20tendency%20to%20produce%20false%20positive%20matches.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%20First%20Workshop%20on%20Large%20Language%20Model%20Memorization%20%28L2M2%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222025-08%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.l2m2-1.4%22%2C%22ISBN%22%3A%22979-8-89176-278-7%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.l2m2-1.4%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A09Z%22%7D%7D%2C%7B%22key%22%3A%22GC869RF2%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Garbas%20et%20al.%22%2C%22parsedDate%22%3A%222025-04%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BGarbas%2C%20L.%2C%20Ploner%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20TransformerRanker%3A%20A%20Tool%20for%20Efficiently%20Finding%20the%20Best-Suited%20Language%20Models%20for%20Downstream%20Classification%20Tasks.%20In%20N.%20Dziri%2C%20S.%20%28Xiang%29%20Ren%2C%20%26amp%3B%20S.%20Diao%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202025%20Conference%20of%20the%20Nations%20of%20the%20Americas%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%20295%26%23x2013%3B302%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.naacl-demo.25%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.naacl-demo.25%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22TransformerRanker%3A%20A%20Tool%20for%20Efficiently%20Finding%20the%20Best-Suited%20Language%20Models%20for%20Downstream%20Classification%20Tasks%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Lukas%22%2C%22lastName%22%3A%22Garbas%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Nouha%22%2C%22lastName%22%3A%22Dziri%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Sean%20%28Xiang%29%22%2C%22lastName%22%3A%22Ren%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Shizhe%22%2C%22lastName%22%3A%22Diao%22%7D%5D%2C%22abstractNote%22%3A%22Classification%20tasks%20in%20NLP%20are%20typically%20addressed%20by%20selecting%20a%20pre-trained%20language%20model%20%28PLM%29%20from%20a%20model%20hub%2C%20and%20fine-tuning%20it%20for%20the%20task%20at%20hand.%20However%2C%20given%20the%20very%20large%20number%20of%20PLMs%20that%20are%20currently%20available%2C%20a%20practical%20challenge%20is%20to%20determine%20which%20of%20them%20will%20perform%20best%20for%20a%20specific%20downstream%20task.%20With%20this%20paper%2C%20we%20introduce%20TransformerRanker%2C%20a%20lightweight%20library%20that%20efficiently%20ranks%20PLMs%20for%20classification%20tasks%20without%20the%20need%20for%20computationally%20costly%20fine-tuning.%20Our%20library%20implements%20current%20approaches%20for%20transferability%20estimation%20%28LogME%2C%20H-Score%2C%20kNN%29%2C%20in%20combination%20with%20layer%20aggregation%20options%2C%20which%20we%20empirically%20showed%20to%20yield%20state-of-the-art%20rankings%20of%20PLMs%20%28Garbas%20et%20al.%2C%202024%29.%20We%20designed%20the%20interface%20to%20be%20lightweight%20and%20easy%20to%20use%2C%20allowing%20users%20to%20directly%20connect%20to%20the%20HuggingFace%20Transformers%20and%20Dataset%20libraries.%20Users%20need%20only%20select%20a%20downstream%20classification%20task%20and%20a%20list%20of%20PLMs%20to%20create%20a%20ranking%20of%20likely%20best-suited%20PLMs%20for%20their%20task.%20We%20make%20TransformerRanker%20available%20as%20a%20pip-installable%20open-source%20library.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202025%20Conference%20of%20the%20Nations%20of%20the%20Americas%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222025-04%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.naacl-demo.25%22%2C%22ISBN%22%3A%22979-8-89176-191-9%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.naacl-demo.25%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A09Z%22%7D%7D%2C%7B%22key%22%3A%2245KSNRRN%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Ploner%20et%20al.%22%2C%22parsedDate%22%3A%222025-04%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BPloner%2C%20M.%2C%20Wiland%2C%20J.%2C%20Pohl%2C%20S.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20LM-Pub-Quiz%3A%20A%20Comprehensive%20Framework%20for%20Zero-Shot%20Evaluation%20of%20Relational%20Knowledge%20in%20Language%20Models.%20In%20N.%20Dziri%2C%20S.%20%28Xiang%29%20Ren%2C%20%26amp%3B%20S.%20Diao%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202025%20Conference%20of%20the%20Nations%20of%20the%20Americas%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2029%26%23x2013%3B39%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.naacl-demo.4%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.naacl-demo.4%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22LM-Pub-Quiz%3A%20A%20Comprehensive%20Framework%20for%20Zero-Shot%20Evaluation%20of%20Relational%20Knowledge%20in%20Language%20Models%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Jacek%22%2C%22lastName%22%3A%22Wiland%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Sebastian%22%2C%22lastName%22%3A%22Pohl%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Nouha%22%2C%22lastName%22%3A%22Dziri%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Sean%20%28Xiang%29%22%2C%22lastName%22%3A%22Ren%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Shizhe%22%2C%22lastName%22%3A%22Diao%22%7D%5D%2C%22abstractNote%22%3A%22Knowledge%20probing%20evaluates%20to%20which%20extent%20a%20language%20model%20%28LM%29%20has%20acquired%20relational%20knowledge%20during%20its%20pre-training%20phase.%20It%20provides%20a%20cost-effective%20means%20of%20comparing%20LMs%20of%20different%20sizes%20and%20training%20setups%20and%20is%20useful%20for%20monitoring%20knowledge%20gained%20or%20lost%20during%20continual%20learning%20%28CL%29.%20In%20prior%20work%2C%20we%20presented%20an%20improved%20knowledge%20probe%20called%20BEAR%20%28Wiland%20et%20al.%2C%202024%29%2C%20which%20enables%20the%20comparison%20of%20LMs%20trained%20with%20different%20pre-training%20objectives%20%28causal%20and%20masked%20LMs%29%20and%20addresses%20issues%20of%20skewed%20distributions%20in%20previous%20probes%20to%20deliver%20a%20more%20unbiased%20reading%20of%20LM%20knowledge.%20With%20this%20paper%2C%20we%20present%20LM-Pub-Quiz%2C%20a%20Python%20framework%20and%20leaderboard%20built%20around%20the%20BEAR%20probing%20mechanism%20that%20enables%20researchers%20and%20practitioners%20to%20apply%20it%20in%20their%20work.%20It%20provides%20options%20for%20standalone%20evaluation%20and%20direct%20integration%20into%20the%20widely-used%20training%20pipeline%20of%20the%20Hugging%20Face%20transformers%20library.%20Further%2C%20it%20provides%20a%20fine-grained%20analysis%20of%20different%20knowledge%20types%20to%20assist%20users%20in%20better%20understanding%20the%20knowledge%20in%20each%20evaluated%20LM.%20We%20publicly%20release%20LM-Pub-Quiz%20as%20an%20open-source%20project.https%3A%5C%2F%5C%2Flm-pub-quiz.github.io%5C%2F%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202025%20Conference%20of%20the%20Nations%20of%20the%20Americas%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222025-04%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.naacl-demo.4%22%2C%22ISBN%22%3A%22979-8-89176-191-9%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.naacl-demo.4%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-29T10%3A40%3A07Z%22%7D%7D%2C%7B%22key%22%3A%22P2MBVNQV%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Christoph%20et%20al.%22%2C%22parsedDate%22%3A%222025-08%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BChristoph%2C%20D.%2C%20Ploner%2C%20M.%2C%20Haller%2C%20P.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20From%20Data%20to%20Knowledge%3A%20Evaluating%20How%20Efficiently%20Language%20Models%20Learn%20Facts.%20In%20R.%20Jia%2C%20E.%20Wallace%2C%20Y.%20Huang%2C%20T.%20Pimentel%2C%20P.%20Maini%2C%20V.%20Dankers%2C%20J.%20Wei%2C%20%26amp%3B%20P.%20Lesci%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%20First%20Workshop%20on%20Large%20Language%20Model%20Memorization%20%28L2M2%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2029%26%23x2013%3B46%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.l2m2-1.3%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.l2m2-1.3%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22From%20Data%20to%20Knowledge%3A%20Evaluating%20How%20Efficiently%20Language%20Models%20Learn%20Facts%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Daniel%22%2C%22lastName%22%3A%22Christoph%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Patrick%22%2C%22lastName%22%3A%22Haller%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Robin%22%2C%22lastName%22%3A%22Jia%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Eric%22%2C%22lastName%22%3A%22Wallace%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yangsibo%22%2C%22lastName%22%3A%22Huang%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Tiago%22%2C%22lastName%22%3A%22Pimentel%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Pratyush%22%2C%22lastName%22%3A%22Maini%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Verna%22%2C%22lastName%22%3A%22Dankers%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Johnny%22%2C%22lastName%22%3A%22Wei%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Pietro%22%2C%22lastName%22%3A%22Lesci%22%7D%5D%2C%22abstractNote%22%3A%22Sample%20efficiency%20is%20a%20crucial%20property%20of%20language%20models%20with%20practical%20implications%20for%20training%20efficiency.%20In%20real-world%20text%2C%20information%20follows%20a%20long-tailed%20distribution.%20Yet%2C%20we%20expect%20models%20to%20learn%20and%20recall%20frequent%20and%20infrequent%20facts.%20Sample%20efficient%20models%20are%20better%20equipped%20to%20handle%20this%20challenge%20of%20learning%20and%20retaining%20rare%20information%20without%20requiring%20excessive%20exposure.%20This%20study%20analyzes%20multiple%20models%20of%20varying%20architectures%20and%20sizes%2C%20all%20trained%20on%20the%20same%20pre-training%20data.%20By%20annotating%20relational%20facts%20with%20their%20frequencies%20in%20the%20training%20corpus%2C%20we%20examine%20how%20model%20performance%20varies%20with%20fact%20frequency.%20Our%20findings%20show%20that%20most%20models%20perform%20similarly%20on%20high-frequency%20facts%20but%20differ%20notably%20on%20low-frequency%20facts.%20This%20analysis%20provides%20new%20insights%20into%20the%20relationship%20between%20model%20architecture%2C%20size%2C%20and%20factual%20learning%20efficiency.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%20First%20Workshop%20on%20Large%20Language%20Model%20Memorization%20%28L2M2%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222025-08%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.l2m2-1.3%22%2C%22ISBN%22%3A%22979-8-89176-278-7%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.l2m2-1.3%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A05%3A34Z%22%7D%7D%2C%7B%22key%22%3A%22G2I7ZRFU%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Alies%20et%20al.%22%2C%22parsedDate%22%3A%222025-07%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BAlies%2C%20R.%2C%20Merdjanovska%2C%20E.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20Measuring%20Label%20Ambiguity%20in%20Subjective%20Tasks%20using%20Predictive%20Uncertainty%20Estimation.%20In%20S.%20Peng%20%26amp%3B%20I.%20Rehbein%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%2019th%20Linguistic%20Annotation%20Workshop%20%28LAW-XIX-2025%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2021%26%23x2013%3B34%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.law-1.2%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.law-1.2%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Measuring%20Label%20Ambiguity%20in%20Subjective%20Tasks%20using%20Predictive%20Uncertainty%20Estimation%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Richard%22%2C%22lastName%22%3A%22Alies%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Elena%22%2C%22lastName%22%3A%22Merdjanovska%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Siyao%22%2C%22lastName%22%3A%22Peng%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Ines%22%2C%22lastName%22%3A%22Rehbein%22%7D%5D%2C%22abstractNote%22%3A%22Human%20annotations%20in%20natural%20language%20corpora%20vary%20due%20to%20differing%20human%20perspectives.%20This%20is%20especially%20prevalent%20in%20subjective%20tasks.%20In%20these%20datasets%2C%20certain%20data%20samples%20are%20more%20prone%20to%20label%20variation%20and%20can%20be%20indicated%20as%20ambiguous%20samples.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%2019th%20Linguistic%20Annotation%20Workshop%20%28LAW-XIX-2025%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222025-07%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.law-1.2%22%2C%22ISBN%22%3A%22979-8-89176-262-6%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.law-1.2%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A04%3A54Z%22%7D%7D%2C%7B%22key%22%3A%22YILTT5NP%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Merdjanovska%20and%20Akbik%22%2C%22parsedDate%22%3A%222025-11%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BMerdjanovska%2C%20E.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282025%29.%20Token-Level%20Metrics%20for%20Detecting%20Incorrect%20Gold%20Annotations%20in%20Named%20Entity%20Recognition.%20In%20C.%20Christodoulopoulos%2C%20T.%20Chakraborty%2C%20C.%20Rose%2C%20%26amp%3B%20V.%20Peng%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BFindings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20EMNLP%202025%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2015292%26%23x2013%3B15304%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.findings-emnlp.827%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2025.findings-emnlp.827%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Token-Level%20Metrics%20for%20Detecting%20Incorrect%20Gold%20Annotations%20in%20Named%20Entity%20Recognition%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Elena%22%2C%22lastName%22%3A%22Merdjanovska%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Christos%22%2C%22lastName%22%3A%22Christodoulopoulos%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Tanmoy%22%2C%22lastName%22%3A%22Chakraborty%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Carolyn%22%2C%22lastName%22%3A%22Rose%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Violet%22%2C%22lastName%22%3A%22Peng%22%7D%5D%2C%22abstractNote%22%3A%22Annotated%20datasets%20for%20supervised%20learning%20tasks%20often%20contain%20incorrect%20gold%20annotations%2C%20i.e.%20label%20noise.%20To%20address%20this%20issue%2C%20many%20noisy%20label%20learning%20approaches%20incorporate%20metrics%20to%20filter%20out%20unreliable%20samples%2C%20for%20example%20using%20heuristics%20such%20as%20high%20loss%20or%20low%20confidence.%20However%2C%20when%20these%20metrics%20are%20integrated%20into%20larger%20pipelines%2C%20it%20becomes%20difficult%20to%20compare%20their%20effectiveness%2C%20and%20understand%20their%20individual%20contribution%20to%20reducing%20label%20noise.%20This%20paper%20directly%20compares%20popular%20sample%20metrics%20for%20detecting%20incorrect%20annotations%20in%20named%20entity%20recognition%20%28NER%29.%20NER%20is%20commonly%20approached%20as%20token%20classification%2C%20so%20the%20metrics%20are%20calculated%20for%20each%20training%20token%20and%20we%20flag%20the%20incorrect%20ones%20by%20defining%20metrics%20thresholds.%20We%20compare%20the%20metrics%20based%20on%20%28i%29%20their%20accuracy%20in%20detecting%20the%20incorrect%20labels%20and%20%28ii%29%20the%20test%20scores%20when%20retraining%20a%20model%20using%20the%20cleaned%20dataset.%20We%20show%20that%20training%20dynamics%20metrics%20work%20the%20best%20overall.%20The%20best%20metrics%20effectively%20reduce%20the%20label%20noise%20across%20different%20noise%20types.%20The%20errors%20that%20the%20model%20has%20not%20yet%20memorized%20are%20more%20feasible%20to%20detect%2C%20and%20relabeling%20these%20tokens%20is%20a%20more%20effective%20strategy%20than%20excluding%20them%20from%20training.%22%2C%22proceedingsTitle%22%3A%22Findings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20EMNLP%202025%22%2C%22conferenceName%22%3A%22Findings%202025%22%2C%22date%22%3A%222025-11%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2025.findings-emnlp.827%22%2C%22ISBN%22%3A%22979-8-89176-335-7%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2025.findings-emnlp.827%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A04%3A51Z%22%7D%7D%2C%7B%22key%22%3A%22C2MIIUUN%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Garbaciauskas%20et%20al.%22%2C%22parsedDate%22%3A%222024-08%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BGarbaciauskas%2C%20L.%2C%20Ploner%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20Choose%20Your%20Transformer%3A%20Improved%20Transferability%20Estimation%20of%20Transformer%20Models%20on%20Classification%20Tasks.%20In%20L.-W.%20Ku%2C%20A.%20Martins%2C%20%26amp%3B%20V.%20Srikumar%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BFindings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20ACL%202024%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2012752%26%23x2013%3B12768%29.%20Association%20for%20Computational%20Linguistics.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-acl.757%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-acl.757%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Choose%20Your%20Transformer%3A%20Improved%20Transferability%20Estimation%20of%20Transformer%20Models%20on%20Classification%20Tasks%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Lukas%22%2C%22lastName%22%3A%22Garbaciauskas%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Lun-Wei%22%2C%22lastName%22%3A%22Ku%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Andre%22%2C%22lastName%22%3A%22Martins%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Vivek%22%2C%22lastName%22%3A%22Srikumar%22%7D%5D%2C%22abstractNote%22%3A%22There%20currently%20exists%20a%20multitude%20of%20pre-trained%20transformer%20language%20models%20%28LMs%29%20that%20are%20readily%20available.%20From%20a%20practical%20perspective%2C%20this%20raises%20the%20question%20of%20which%20pre-trained%20LM%20will%20perform%20best%20if%20fine-tuned%20for%20a%20specific%20downstream%20NLP%20task.%20However%2C%20exhaustively%20fine-tuning%20all%20available%20LMs%20to%20determine%20the%20best-fitting%20model%20is%20computationally%20infeasible.%20To%20address%20this%20problem%2C%20we%20present%20an%20approach%20that%20inexpensively%20estimates%20a%20ranking%20of%20the%20expected%20performance%20of%20a%20given%20set%20of%20candidate%20LMs%20for%20a%20given%20task.%20Following%20a%20layer-wise%20representation%20analysis%2C%20we%20extend%20existing%20approaches%20such%20as%20H-score%20and%20LogME%20by%20aggregating%20representations%20across%20all%20layers%20of%20the%20transformer%20model.%20We%20present%20an%20extensive%20analysis%20of%2020%20transformer%20LMs%2C%206%20downstream%20NLP%20tasks%2C%20and%20various%20estimators%20%28linear%20probing%2C%20kNN%2C%20H-score%2C%20and%20LogME%29.%20Our%20evaluation%20finds%20that%20averaging%20the%20layer%20representations%20significantly%20improves%20the%20Pearson%20correlation%20coefficient%20between%20the%20true%20model%20ranks%20and%20the%20estimate%2C%20increasing%20from%200.58%20to%200.86%20for%20LogME%20and%20from%200.65%20to%200.88%20for%20H-score.%22%2C%22proceedingsTitle%22%3A%22Findings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20ACL%202024%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024-08%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.findings-acl.757%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.findings-acl.757%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A01%3A04Z%22%7D%7D%2C%7B%22key%22%3A%22YVLKMAPQ%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Ploner%20and%20Akbik%22%2C%22parsedDate%22%3A%222024-03%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BPloner%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20Parameter-Efficient%20Fine-Tuning%3A%20Is%20There%20An%20Optimal%20Subset%20of%20Parameters%20to%20Tune%3F%20In%20Y.%20Graham%20%26amp%3B%20M.%20Purver%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BFindings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20EACL%202024%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%201743%26%23x2013%3B1759%29.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-eacl.122%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-eacl.122%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Parameter-Efficient%20Fine-Tuning%3A%20Is%20There%20An%20Optimal%20Subset%20of%20Parameters%20to%20Tune%3F%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yvette%22%2C%22lastName%22%3A%22Graham%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Matthew%22%2C%22lastName%22%3A%22Purver%22%7D%5D%2C%22abstractNote%22%3A%22The%20ever-growing%20size%20of%20pretrained%20language%20models%20%28PLM%29%20presents%20a%20significant%20challenge%20for%20efficiently%20fine-tuning%20and%20deploying%20these%20models%20for%20diverse%20sets%20of%20tasks%20within%20memory-constrained%20environments.In%20light%20of%20this%2C%20recent%20research%20has%20illuminated%20the%20possibility%20of%20selectively%20updating%20only%20a%20small%20subset%20of%20a%20model%26%23039%3Bs%20parameters%20during%20the%20fine-tuning%20process.Since%20no%20new%20parameters%20or%20modules%20are%20added%2C%20these%20methods%20retain%20the%20inference%20speed%20of%20the%20original%20model%20and%20come%20at%20no%20additional%20computational%20cost.%20However%2C%20an%20open%20question%20pertains%20to%20which%20subset%20of%20parameters%20should%20best%20be%20tuned%20to%20maximize%20task%20performance%20and%20generalizability.%20To%20investigate%2C%20this%20paper%20presents%20comprehensive%20experiments%20covering%20a%20large%20spectrum%20of%20subset%20selection%20strategies.%20We%20comparatively%20evaluate%20their%20impact%20on%20model%20performance%20as%20well%20as%20the%20resulting%20model%26%23039%3Bs%20capability%20to%20generalize%20to%20different%20tasks.Surprisingly%2C%20we%20find%20that%20the%20gains%20achieved%20in%20performance%20by%20elaborate%20selection%20strategies%20are%2C%20at%20best%2C%20marginal%20when%20compared%20to%20the%20outcomes%20obtained%20by%20tuning%20a%20random%20selection%20of%20parameter%20subsets.%20Our%20experiments%20also%20indicate%20that%20selection-based%20tuning%20impairs%20generalizability%20to%20new%20tasks.%22%2C%22proceedingsTitle%22%3A%22Findings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20EACL%202024%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024-03%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.findings-eacl.122%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.findings-eacl.122%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A00%3A59Z%22%7D%7D%2C%7B%22key%22%3A%22UYUDH7PP%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Dallabetta%20et%20al.%22%2C%22parsedDate%22%3A%222024-08%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BDallabetta%2C%20M.%2C%20Dobberstein%2C%20C.%2C%20Breiding%2C%20A.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20Fundus%3A%20A%20Simple-to-Use%20News%20Scraper%20Optimized%20for%20High%20Quality%20Extractions.%20In%20Y.%20Cao%2C%20Y.%20Feng%2C%20%26amp%3B%20D.%20Xiong%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%2062nd%20Annual%20Meeting%20of%20the%20Association%20for%20Computational%20Linguistics%20%28ACL%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%20305%26%23x2013%3B314%29.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.acl-demos.29%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.acl-demos.29%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Fundus%3A%20A%20Simple-to-Use%20News%20Scraper%20Optimized%20for%20High%20Quality%20Extractions%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Dallabetta%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Conrad%22%2C%22lastName%22%3A%22Dobberstein%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Adrian%22%2C%22lastName%22%3A%22Breiding%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yixin%22%2C%22lastName%22%3A%22Cao%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Yang%22%2C%22lastName%22%3A%22Feng%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Deyi%22%2C%22lastName%22%3A%22Xiong%22%7D%5D%2C%22abstractNote%22%3A%22This%20paper%20introduces%20Fundus%2C%20a%20user-friendly%20news%20scraper%20that%20enables%20users%20to%20obtain%20millions%20of%20high-quality%20news%20articles%20with%20just%20a%20few%20lines%20of%20code.%20Unlike%20existing%20news%20scrapers%2C%20we%20use%20manually%20crafted%2C%20bespoke%20content%20extractors%20that%20are%20specifically%20tailored%20to%20the%20formatting%20guidelines%20of%20each%20supported%20online%20newspaper.%20This%20allows%20us%20to%20optimize%20our%20scraping%20for%20quality%20such%20that%20retrieved%20news%20articles%20are%20textually%20complete%20and%20without%20HTML%20artifacts.%20Further%2C%20our%20framework%20combines%20both%20crawling%20%28retrieving%20HTML%20from%20the%20web%20or%20large%20web%20archives%29%20and%20content%20extraction%20into%20a%20single%20pipeline.%20By%20providing%20a%20unified%20interface%20for%20a%20predefined%20collection%20of%20newspapers%2C%20we%20aim%20to%20make%20Fundus%20broadly%20usable%20even%20for%20non-technical%20users.%20This%20paper%20gives%20an%20overview%20of%20the%20framework%2C%20discusses%20our%20design%20choices%2C%20and%20presents%20a%20comparative%20evaluation%20against%20other%20popular%20news%20scrapers.%20Our%20evaluation%20shows%20that%20Fundus%20yields%20significantly%20higher%20quality%20extractions%20%28complete%20and%20artifact-free%20news%20articles%29%20than%20prior%20work.The%20framework%20is%20available%20on%20GitHub%20under%20https%3A%5C%2F%5C%2Fgithub.com%5C%2FflairNLP%5C%2Ffundus%20and%20can%20be%20simply%20installed%20using%20pip.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%2062nd%20Annual%20Meeting%20of%20the%20Association%20for%20Computational%20Linguistics%20%28ACL%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024-08%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.acl-demos.29%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.acl-demos.29%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T12%3A00%3A47Z%22%7D%7D%2C%7B%22key%22%3A%22UBWNW8LN%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Schulte%20et%20al.%22%2C%22parsedDate%22%3A%222024%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BSchulte%2C%20D.%2C%20Hamborg%2C%20F.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20Less%20is%20More%3A%20Parameter-Efficient%20Selection%20of%20Intermediate%20Tasks%20for%20Transfer%20Learning.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202024%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%26lt%3B%5C%2Fi%26gt%3B%2C%209431%26%23x2013%3B9442.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.emnlp-main.529%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.emnlp-main.529%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Less%20is%20More%3A%20Parameter-Efficient%20Selection%20of%20Intermediate%20Tasks%20for%20Transfer%20Learning%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22David%22%2C%22lastName%22%3A%22Schulte%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Felix%22%2C%22lastName%22%3A%22Hamborg%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202024%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%22%2C%22conferenceName%22%3A%22Proceedings%20of%20the%202024%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%22%2C%22date%22%3A%222024%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.emnlp-main.529%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.emnlp-main.529%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A59%3A03Z%22%7D%7D%2C%7B%22key%22%3A%22KNY9WVVK%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Golde%20et%20al.%22%2C%22parsedDate%22%3A%222023%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BGolde%2C%20J.%2C%20Haller%2C%20P.%2C%20Hamborg%2C%20F.%2C%20Risch%2C%20J.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282023%29.%20Fabricator%3A%20An%20Open%20Source%20Toolkit%20for%20Generating%20Labeled%20Training%20Data%20with%20Teacher%20LLMs.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202023%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%20%28EMNLP%29%26lt%3B%5C%2Fi%26gt%3B%2C%201%26%23x2013%3B11.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.emnlp-demo.1%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.emnlp-demo.1%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22Fabricator%3A%20An%20Open%20Source%20Toolkit%20for%20Generating%20Labeled%20Training%20Data%20with%20Teacher%20LLMs%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Jonas%22%2C%22lastName%22%3A%22Golde%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Patrick%22%2C%22lastName%22%3A%22Haller%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Felix%22%2C%22lastName%22%3A%22Hamborg%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Julian%22%2C%22lastName%22%3A%22Risch%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202023%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%20%28EMNLP%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222023%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2023.emnlp-demo.1%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2023.emnlp-demo.1%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A56%3A13Z%22%7D%7D%2C%7B%22key%22%3A%22APEPUISA%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22R%5Cu00fccker%20and%20Akbik%22%2C%22parsedDate%22%3A%222023%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BR%26%23xFC%3Bcker%2C%20S.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282023%29.%20CleanCoNLL%3A%20A%20Nearly%20Noise-Free%20Named%20Entity%20Recognition%20Dataset.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202023%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%20%28EMNLP%29%26lt%3B%5C%2Fi%26gt%3B%2C%208628%26%23x2013%3B8645.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.emnlp-main.533%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.emnlp-main.533%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22CleanCoNLL%3A%20A%20Nearly%20Noise-Free%20Named%20Entity%20Recognition%20Dataset%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Susanna%22%2C%22lastName%22%3A%22R%5Cu00fccker%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202023%20Conference%20on%20Empirical%20Methods%20in%20Natural%20Language%20Processing%20%28EMNLP%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222023%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2023.emnlp-main.533%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2023.emnlp-main.533%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A56%3A10Z%22%7D%7D%2C%7B%22key%22%3A%223AL7FFIA%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Wiland%20et%20al.%22%2C%22parsedDate%22%3A%222024%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BWiland%2C%20J.%2C%20Ploner%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20BEAR%3A%20A%20Unified%20Framework%20for%20Evaluating%20Relational%20Knowledge%20in%20Causal%20and%20Masked%20Language%20Models.%20%26lt%3Bi%26gt%3BFindings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20NAACL%202024%26lt%3B%5C%2Fi%26gt%3B%2C%202393%26%23x2013%3B2411.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-naacl.155%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.findings-naacl.155%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22BEAR%3A%20A%20Unified%20Framework%20for%20Evaluating%20Relational%20Knowledge%20in%20Causal%20and%20Masked%20Language%20Models%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Jacek%22%2C%22lastName%22%3A%22Wiland%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Max%22%2C%22lastName%22%3A%22Ploner%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Findings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20NAACL%202024%22%2C%22conferenceName%22%3A%22Findings%20of%20the%20Association%20for%20Computational%20Linguistics%3A%20NAACL%202024%22%2C%22date%22%3A%222024%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.findings-naacl.155%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.findings-naacl.155%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A56%3A07Z%22%7D%7D%2C%7B%22key%22%3A%2263QC4DVL%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Haller%20et%20al.%22%2C%22parsedDate%22%3A%222024%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BHaller%2C%20P.%2C%20Aynetdinov%2C%20A.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20OpinionGPT%3A%20Modelling%20Explicit%20Biases%20in%20Instruction-Tuned%20LLMs.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202024%20Conference%20of%20the%20North%20American%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%20%28NAACL%29%26lt%3B%5C%2Fi%26gt%3B%2C%2078%26%23x2013%3B86.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.naacl-demo.8%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2024.naacl-demo.8%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22OpinionGPT%3A%20Modelling%20Explicit%20Biases%20in%20Instruction-Tuned%20LLMs%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Patrick%22%2C%22lastName%22%3A%22Haller%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Ansar%22%2C%22lastName%22%3A%22Aynetdinov%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202024%20Conference%20of%20the%20North%20American%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%20%28NAACL%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2024.naacl-demo.8%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.naacl-demo.8%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A56%3A04Z%22%7D%7D%2C%7B%22key%22%3A%22WM7NNNJW%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Milich%20and%20Akbik%22%2C%22parsedDate%22%3A%222023%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BMilich%2C%20M.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282023%29.%20ZELDA%3A%20A%20Comprehensive%20Benchmark%20for%20Supervised%20Entity%20Disambiguation.%20%26lt%3Bi%26gt%3BProceedings%20of%20the%2017th%20Conference%20of%20the%20European%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%26lt%3B%5C%2Fi%26gt%3B%2C%202061%26%23x2013%3B2072.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.eacl-main.151%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.18653%5C%2Fv1%5C%2F2023.eacl-main.151%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22ZELDA%3A%20A%20Comprehensive%20Benchmark%20for%20Supervised%20Entity%20Disambiguation%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Marcel%22%2C%22lastName%22%3A%22Milich%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%2017th%20Conference%20of%20the%20European%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%22%2C%22conferenceName%22%3A%22Proceedings%20of%20the%2017th%20Conference%20of%20the%20European%20Chapter%20of%20the%20Association%20for%20Computational%20Linguistics%22%2C%22date%22%3A%222023%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%2210.18653%5C%2Fv1%5C%2F2023.eacl-main.151%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2023.eacl-main.151%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22en%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-22T11%3A52%3A23Z%22%7D%7D%2C%7B%22key%22%3A%223I8DQD8J%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Kissling%20et%20al.%22%2C%22parsedDate%22%3A%222026-01-26%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BKissling%2C%20C.%2C%20Merdjanovska%2C%20E.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282026%29.%20%26lt%3Bi%26gt%3BSelf-Aware%20Knowledge%20Probing%3A%20Evaluating%20Language%20Models%26%23x2019%3B%20Relational%20Knowledge%20through%20Confidence%20Calibration%26lt%3B%5C%2Fi%26gt%3B.%20arXiv.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-DOIURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.48550%5C%2FarXiv.2601.18901%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Fdoi.org%5C%2F10.48550%5C%2FarXiv.2601.18901%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22preprint%22%2C%22title%22%3A%22Self-Aware%20Knowledge%20Probing%3A%20Evaluating%20Language%20Models%27%20Relational%20Knowledge%20through%20Confidence%20Calibration%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Christopher%22%2C%22lastName%22%3A%22Kissling%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Elena%22%2C%22lastName%22%3A%22Merdjanovska%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%5D%2C%22abstractNote%22%3A%22Knowledge%20probing%20quantifies%20how%20much%20relational%20knowledge%20a%20language%20model%20%28LM%29%20has%20acquired%20during%20pre-training.%20Existing%20knowledge%20probes%20evaluate%20model%20capabilities%20through%20metrics%20like%20prediction%20accuracy%20and%20precision.%20Such%20evaluations%20fail%20to%20account%20for%20the%20model%26%23039%3Bs%20reliability%2C%20reflected%20in%20the%20calibration%20of%20its%20confidence%20scores.%20In%20this%20paper%2C%20we%20propose%20a%20novel%20calibration%20probing%20framework%20for%20relational%20knowledge%2C%20covering%20three%20modalities%20of%20model%20confidence%3A%20%281%29%20intrinsic%20confidence%2C%20%282%29%20structural%20consistency%20and%20%283%29%20semantic%20grounding.%20Our%20extensive%20analysis%20of%20ten%20causal%20and%20six%20masked%20language%20models%20reveals%20that%20most%20models%2C%20especially%20those%20pre-trained%20with%20the%20masking%20objective%2C%20are%20overconfident.%20The%20best-calibrated%20scores%20come%20from%20confidence%20estimates%20that%20account%20for%20inconsistencies%20due%20to%20statement%20rephrasing.%20Moreover%2C%20even%20the%20largest%20pre-trained%20models%20fail%20to%20encode%20the%20semantics%20of%20linguistic%20confidence%20expressions%20accurately.%22%2C%22genre%22%3A%22%22%2C%22repository%22%3A%22arXiv%22%2C%22archiveID%22%3A%22%22%2C%22date%22%3A%222026-01-26%22%2C%22DOI%22%3A%2210.48550%5C%2FarXiv.2601.18901%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22http%3A%5C%2F%5C%2Farxiv.org%5C%2Fabs%5C%2F2601.18901%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-18T09%3A10%3A16Z%22%7D%7D%2C%7B%22key%22%3A%22YJU4DFRN%22%2C%22library%22%3A%7B%22id%22%3A6984777%7D%2C%22meta%22%3A%7B%22creatorSummary%22%3A%22Haller%20et%20al.%22%2C%22parsedDate%22%3A%222024-05%22%2C%22numChildren%22%3A1%7D%2C%22bib%22%3A%22%26lt%3Bdiv%20class%3D%26quot%3Bcsl-bib-body%26quot%3B%20style%3D%26quot%3Bline-height%3A%202%3B%20padding-left%3A%201em%3B%20text-indent%3A-1em%3B%26quot%3B%26gt%3B%5Cn%20%20%26lt%3Bdiv%20class%3D%26quot%3Bcsl-entry%26quot%3B%26gt%3BHaller%2C%20P.%2C%20Golde%2C%20J.%2C%20%26amp%3B%20Akbik%2C%20A.%20%282024%29.%20PECC%3A%20Problem%20Extraction%20and%20Coding%20Challenges.%20In%20N.%20Calzolari%2C%20M.-Y.%20Kan%2C%20V.%20Hoste%2C%20A.%20Lenci%2C%20S.%20Sakti%2C%20%26amp%3B%20N.%20Xue%20%28Eds.%29%2C%20%26lt%3Bi%26gt%3BProceedings%20of%20the%202024%20Joint%20International%20Conference%20on%20Computational%20Linguistics%2C%20Language%20Resources%20and%20Evaluation%20%28LREC-COLING%202024%29%26lt%3B%5C%2Fi%26gt%3B%20%28pp.%2012690%26%23x2013%3B12699%29.%20%26lt%3Ba%20class%3D%26%23039%3Bzp-ItemURL%26%23039%3B%20href%3D%26%23039%3Bhttps%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.lrec-main.1111%5C%2F%26%23039%3B%26gt%3Bhttps%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.lrec-main.1111%5C%2F%26lt%3B%5C%2Fa%26gt%3B%26lt%3B%5C%2Fdiv%26gt%3B%5Cn%26lt%3B%5C%2Fdiv%26gt%3B%22%2C%22data%22%3A%7B%22itemType%22%3A%22conferencePaper%22%2C%22title%22%3A%22PECC%3A%20Problem%20Extraction%20and%20Coding%20Challenges%22%2C%22creators%22%3A%5B%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Patrick%22%2C%22lastName%22%3A%22Haller%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Jonas%22%2C%22lastName%22%3A%22Golde%22%7D%2C%7B%22creatorType%22%3A%22author%22%2C%22firstName%22%3A%22Alan%22%2C%22lastName%22%3A%22Akbik%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Nicoletta%22%2C%22lastName%22%3A%22Calzolari%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Min-Yen%22%2C%22lastName%22%3A%22Kan%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Veronique%22%2C%22lastName%22%3A%22Hoste%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Alessandro%22%2C%22lastName%22%3A%22Lenci%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Sakriani%22%2C%22lastName%22%3A%22Sakti%22%7D%2C%7B%22creatorType%22%3A%22editor%22%2C%22firstName%22%3A%22Nianwen%22%2C%22lastName%22%3A%22Xue%22%7D%5D%2C%22abstractNote%22%3A%22Recent%20advancements%20in%20large%20language%20models%20%28LLMs%29%20have%20showcased%20their%20exceptional%20abilities%20across%20various%20tasks%2C%20such%20as%20code%20generation%2C%20problem-solving%20and%20reasoning.%20Existing%20benchmarks%20evaluate%20tasks%20in%20isolation%2C%20yet%20the%20extent%20to%20which%20LLMs%20can%20understand%20prose-style%20tasks%2C%20identify%20the%20underlying%20problems%2C%20and%20then%20generate%20appropriate%20code%20solutions%20is%20still%20unexplored.%20Addressing%20this%20gap%2C%20we%20introduce%20PECC%2C%20a%20novel%20benchmark%20derived%20from%20Advent%20Of%20Code%20%28AoC%29%20challenges%20and%20Project%20Euler%2C%20including%202396%20problems.%20Unlike%20conventional%20benchmarks%2C%20PECC%20requires%20LLMs%20to%20interpret%20narrative-embedded%20problems%2C%20extract%20requirements%2C%20and%20generate%20executable%20code.%20A%20key%20feature%20of%20our%20dataset%20is%20the%20complexity%20added%20by%20natural%20language%20prompting%20in%20chat-based%20evaluations%2C%20mirroring%20real-world%20instruction%20ambiguities.%20Results%20show%20varying%20model%20performance%20between%20narrative%20and%20neutral%20problems%2C%20with%20specific%20challenges%20in%20the%20Euler%20math-based%20subset%20with%20GPT-3.5-Turbo%20passing%2050%25%20of%20the%20AoC%20challenges%20and%20only%208%25%20on%20the%20Euler%20problems.%20By%20probing%20the%20limits%20of%20LLMs%26%23039%3B%20capabilities%2C%20our%20benchmark%20provides%20a%20framework%20to%20monitor%20and%20assess%20the%20subsequent%20progress%20of%20LLMs%20as%20a%20universal%20problem%20solver.%22%2C%22proceedingsTitle%22%3A%22Proceedings%20of%20the%202024%20Joint%20International%20Conference%20on%20Computational%20Linguistics%2C%20Language%20Resources%20and%20Evaluation%20%28LREC-COLING%202024%29%22%2C%22conferenceName%22%3A%22%22%2C%22date%22%3A%222024-05%22%2C%22eventPlace%22%3A%22%22%2C%22DOI%22%3A%22%22%2C%22ISBN%22%3A%22%22%2C%22citationKey%22%3A%22%22%2C%22url%22%3A%22https%3A%5C%2F%5C%2Faclanthology.org%5C%2F2024.lrec-main.1111%5C%2F%22%2C%22ISSN%22%3A%22%22%2C%22language%22%3A%22%22%2C%22collections%22%3A%5B%5D%2C%22dateModified%22%3A%222026-06-16T09%3A06%3A12Z%22%7D%7D%5D%7D
Ziletti, A., Akbik, A., Berns, C., Herold, T., Legler, M., & Viell, M. (2022). Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Track, 176–187. https://doi.org/10.18653/v1/2022.naacl-industry.21
Golde, J., Hamborg, F., & Akbik, A. (2024). Large-Scale Label Interpretation Learning for Few-Shot Named Entity Recognition. In Y. Graham & M. Purver (Eds.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2024) (pp. 2915–2930). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.eacl-long.178
Merdjanovska, E., Aynetdinov, A., & Akbik, A. (2024). NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 18182–18198). https://doi.org/10.18653/v1/2024.emnlp-main.1011
Pohl, S., Ploner, M., & Akbik, A. (2025). Towards a Principled Evaluation of Knowledge Editors. In R. Jia, E. Wallace, Y. Huang, T. Pimentel, P. Maini, V. Dankers, J. Wei, & P. Lesci (Eds.), Proceedings of the First Workshop on Large Language Model Memorization (L2M2) (pp. 47–60). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.l2m2-1.4
Garbas, L., Ploner, M., & Akbik, A. (2025). TransformerRanker: A Tool for Efficiently Finding the Best-Suited Language Models for Downstream Classification Tasks. In N. Dziri, S. (Xiang) Ren, & S. Diao (Eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (pp. 295–302). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-demo.25
Ploner, M., Wiland, J., Pohl, S., & Akbik, A. (2025). LM-Pub-Quiz: A Comprehensive Framework for Zero-Shot Evaluation of Relational Knowledge in Language Models. In N. Dziri, S. (Xiang) Ren, & S. Diao (Eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (pp. 29–39). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-demo.4
Christoph, D., Ploner, M., Haller, P., & Akbik, A. (2025). From Data to Knowledge: Evaluating How Efficiently Language Models Learn Facts. In R. Jia, E. Wallace, Y. Huang, T. Pimentel, P. Maini, V. Dankers, J. Wei, & P. Lesci (Eds.), Proceedings of the First Workshop on Large Language Model Memorization (L2M2) (pp. 29–46). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.l2m2-1.3
Alies, R., Merdjanovska, E., & Akbik, A. (2025). Measuring Label Ambiguity in Subjective Tasks using Predictive Uncertainty Estimation. In S. Peng & I. Rehbein (Eds.), Proceedings of the 19th Linguistic Annotation Workshop (LAW-XIX-2025) (pp. 21–34). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.law-1.2
Merdjanovska, E., & Akbik, A. (2025). Token-Level Metrics for Detecting Incorrect Gold Annotations in Named Entity Recognition. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 15292–15304). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.827
Garbaciauskas, L., Ploner, M., & Akbik, A. (2024). Choose Your Transformer: Improved Transferability Estimation of Transformer Models on Classification Tasks. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 12752–12768). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.757
Ploner, M., & Akbik, A. (2024). Parameter-Efficient Fine-Tuning: Is There An Optimal Subset of Parameters to Tune? In Y. Graham & M. Purver (Eds.), Findings of the Association for Computational Linguistics: EACL 2024 (pp. 1743–1759). https://doi.org/10.18653/v1/2024.findings-eacl.122
Dallabetta, M., Dobberstein, C., Breiding, A., & Akbik, A. (2024). Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions. In Y. Cao, Y. Feng, & D. Xiong (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) (pp. 305–314). https://doi.org/10.18653/v1/2024.acl-demos.29
Schulte, D., Hamborg, F., & Akbik, A. (2024). Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9431–9442. https://doi.org/10.18653/v1/2024.emnlp-main.529
Golde, J., Haller, P., Hamborg, F., Risch, J., & Akbik, A. (2023). Fabricator: An Open Source Toolkit for Generating Labeled Training Data with Teacher LLMs. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1–11. https://doi.org/10.18653/v1/2023.emnlp-demo.1
Rücker, S., & Akbik, A. (2023). CleanCoNLL: A Nearly Noise-Free Named Entity Recognition Dataset. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), 8628–8645. https://doi.org/10.18653/v1/2023.emnlp-main.533
Wiland, J., Ploner, M., & Akbik, A. (2024). BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models. Findings of the Association for Computational Linguistics: NAACL 2024, 2393–2411. https://doi.org/10.18653/v1/2024.findings-naacl.155
Haller, P., Aynetdinov, A., & Akbik, A. (2024). OpinionGPT: Modelling Explicit Biases in Instruction-Tuned LLMs. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), 78–86. https://doi.org/10.18653/v1/2024.naacl-demo.8
Milich, M., & Akbik, A. (2023). ZELDA: A Comprehensive Benchmark for Supervised Entity Disambiguation. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, 2061–2072. https://doi.org/10.18653/v1/2023.eacl-main.151
Kissling, C., Merdjanovska, E., & Akbik, A. (2026). Self-Aware Knowledge Probing: Evaluating Language Models’ Relational Knowledge through Confidence Calibration. arXiv. https://doi.org/10.48550/arXiv.2601.18901
Haller, P., Golde, J., & Akbik, A. (2024). PECC: Problem Extraction and Coding Challenges. In N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, & N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 12690–12699). https://aclanthology.org/2024.lrec-main.1111/

EACL’s Outstanding Paper Award (2023)

Emmy Noether Grant (2021)

Research

An overview of our scientific work

See our Research Projects