Impact of Lexical Ambiguity and Underspecification on LLM Training
October 1, 2026
Training on artificial homonyms and hypernyms reveals that ambiguity and underspecification increase model performance in correlation with type-token ratios. However, models show decreased accuracy when generating sequences containing these ambiguous words. Mechanistic analysis shows internal representations successfully disambiguate pseudo-homonyms.
HOW THIS AFFECTS YOU
●
researcherUse these findings to evaluate how linguistic complexity and ambiguity influence model weight updates.