Torna ai risultati
Scheda bibliografica · Consultazione e accesso
Artículo

Isolating LLM Lexical Bias

Xiaoyang Ming et al · LibraryPress@UF · 2026

Materiale supplementare disponibile
Lettura rapida. Controlla i dati essenziali della risorsa e accedi al contenuto con il pulsante principale. La scheda mostra solo le informazioni necessarie per identificare, citare e aprire l’opera.

Accesso alla risorsa

Apri il contenuto dall’opzione principale o scegli un’altra fonte disponibile.

DOAJ DOAJ Articles
Entrar por DOAJ
Accesso principale

Materiale supplementare disponibile

El enlace apunta a material asociado, anexos, tablas, datos o página complementaria. No se marca como libro/texto completo.
Apri materiale

Riepilogo

Descripción general del contenido del recurso.

Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignments are thought to partly originate in the preference-learning stage, e.g., Reinforcement Learning from Human Feedback, which generally makes models more useful but simultaneously may introduce systematic lexical bias. In terms of lexical behavior, this is visible in a model’s preference for certain formats or the overuse of words (delve, furthermore), even when such patterns are not present in base model outputs. Research on lexical misalignment induced during preference training is constrained by reliance on manual curation. We address this, by introducing the Triangulated Preference Shift score, a metric that triangulates between human gold standards, base models, and instruct variants to isolate shifts induced specifically by preference learning, without manual curation. We provide data across six model families, anchor the results in the literature, and illustrate the general approach’s utility by analyzing whether preference learning shifts models toward what could be interpreted as a “language of prestige”. The metric provides an automated method to quantify behavioral shifts attributable to preference tuning, and thus, supports model alignment and development of trustworthy AI.

Come citare

Elegí el formato que necesitás y copiá la referencia al portapapeles.

APA 7

al, X. M. E. (2026). Isolating LLM Lexical Bias. https://journals.flvc.org/FLAIRS/article/view/141843

MLA

al, Xiaoyang Ming et. "Isolating LLM Lexical Bias." 2026. https://journals.flvc.org/FLAIRS/article/view/141843.

Chicago

al, Xiaoyang Ming et. 2026. "Isolating LLM Lexical Bias.". https://journals.flvc.org/FLAIRS/article/view/141843.

Harvard

al, X. M. E. 2026, Isolating LLM Lexical Bias, LibraryPress@UF, available at: https://journals.flvc.org/FLAIRS/article/view/141843 [Accessed 10 Aug. 2026].

Condividi e stampa

Salva la scheda, copia il link permanente o stampala in PDF.

Esporta riferimento

Esporta il record nei formati più comuni per usarlo con un gestore bibliografico.

Dettagli della risorsa

Informazioni bibliografiche utili per verificare che sia il materiale corretto.

Titolo
Isolating LLM Lexical Bias
Autore / collaboratori
Xiaoyang Ming et al
Editore
LibraryPress@UF
Anno di pubblicazione
2026
ISSN
2334-0754
ISSN
2334-0754
Lingua
Inglés

Soggetti

Esplora risorse correlate a partire da questi soggetti.

Copiato