Character-Level LSTM Models for Sinhala Sandhi Splitting
September 30, 2026
A bidirectional LSTM encoder-decoder architecture achieves 68.40% exact-match accuracy on complex Sinhala Sandhi splitting tasks. Performance drops significantly on lexicalized and etymological subsets compared to the 94.00% accuracy achieved on regular affixational subsets.
HOW THIS AFFECTS YOU
●
researcherThe results highlight the specific difficulty of handling etymological morphological boundaries in low-resource language NLP.